The categorical fix for the recurring "assistant produced a flow I can't
save" problem. Until now, POST/PUT /flows rejected (400) on ANY
@ValidFlowStructure violation, conflating two very different things:
genuinely corrupt/inconsistent data, and a flow that is merely not
runnable yet. The user's model - and how workflow editors normally
behave - is that an incomplete flow must be savable as a DRAFT and only
gated at execution time.
Key realisation: FlowExecutionValidator.collectErrors already runs the
same @ValidFlowStructure bean validation, so DRAFT vs EXECUTABLE status
(toView -> isExecutable) already reflects every structural/executability
problem, and ExecutionsService.startExecution independently calls
flowExecutionValidator.validate() - so a non-executable draft can never
actually run. The hard save-gate was therefore redundant for the
executability class; only data-integrity needed to keep blocking.
FlowService.validateFlow now partitions violations by code:
- SAVE_BLOCKING_CODES (integrity: type/inputs/outputs mismatch, missing
config, duplicate/missing node ids, unknown node type, nested
containers, lane integrity, global-input integrity, bias-annotation
integrity, and non-decodable request-level constraints like a null
flow) still reject with 400, re-encoded via ValidationErrorCodec so
the structured errors[] contract is unchanged.
- everything else (dangling/absent connections, container subflow not
yet exposing its handles, exclusive-branch merges, branch-rejoin/end
gaps, dependencies, deadlocks, shared-session ordering, ...) no longer
blocks: the flow saves as DRAFT and the issue is surfaced by
GET /flows/{id}/validation.
This closes the whole class of "structurally sane but not runnable ->
can't save" failures once and for all, instead of chasing each variant.
Tests: dangling connection now saves as DRAFT (was 400); the two
exclusive-branch-merge tests updated from "rejected" to "saved as draft,
reported by the execution validator"; factory-tampering and
bias-integrity rejections still 400 unchanged. 443/443.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Answers "could the extra-braces problems be fixed statically?" - yes,
the syntactic-noise class can and now is.
The model sometimes mangles a connection endpoint name with purely
syntactic noise: a stray/unbalanced brace ("{category"), an accidental
${{...}} wrapper, or a block-qualified reference ("classify.response").
normalizeBlockReference only did trim()+toLowerCase(), so "{category"
never matched the real "category" input and the connection was silently
dropped, leaving the flow disconnected.
Added stripHandleNoise() - removes ${{ }} / {{ }} wrappers, stray
braces/$/quotes, and a leading block-name qualifier (keeps the last
dotted segment) - and wired it as a FALLBACK in findIoByName and
resolveConnectionBlock: it only runs after the exact-name match already
failed, so it can never change a currently-resolving reference, only
rescue one that would otherwise be dropped. It never invents a name, so
a genuinely-wrong reference (not just mangled) still fails and is
dropped, as before.
Scope note: this fixes the SYNTACTIC class only. Semantic/structural
problems (duplicated logic, connections to non-existent handles,
topology/deadlocks) are unaffected - those need the prompt-side and/or
soft-vs-hard-validation work, not name cleanup.
Test isolates the sanitization path by giving the target two inputs so
the pre-existing single-input shortcut cannot mask it; verified by
mutation that disabling the fallback drops the connection. 442/442.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two more issues found retesting the same live prompt after the
guardSubFlow-redaction fix (confirmed working: no more guard-scaffold
names leaking into prompts).
1. The model kept declaring the same logical step both as a top-level
block AND as a container's own inner block (e.g. "check completion"
as both b4 and c1-b1), then tried to wire the orphaned top-level
duplicate to the container via malformed connections (dotted
qualified names, null toInput) - all silently dropped as invalid,
but leaving the duplicate disconnected/deadlocked instead of fixing
the real problem. Added an explicit NO DUPLICATION rule to the plan
prompt: a step belongs in exactly one place, top-level or inside one
container, never both - only the container's own exposed I/O is
what the rest of the flow should connect to.
2. Separately, a fresh model response put a container type
("LoopContainer") as an entry in the flat "blocks" list instead of
"containers". Block assembly has no catalog descriptor for a
container type, so this threw as an unrecoverable 502 with no
retry - unlike degenerate JSON or a missing required field, which
already get a structured-repair retry. Moved the check into
validateAndNormalizePlan, inside the plan's own retry-wrapped parser
callback, so this now gets the same self-correction chance instead
of hard-failing the whole request.
Verified by mutation testing: reverting the container-type check
reproduces the exact live 502 in the new test; restoring it fixes it.
Full suite: 441/441.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Live incident: retrying the same prompt from the UI logged repeated
"Skipping invalid assistant connection draft" warnings referencing
block names like "c1-expose-feedback" and "c1-guard-evaluator" - the
deterministic guard scaffold FlowAssistantService#buildLoopGuardSubFlow
builds and the model never authors.
Root cause: summarizeFlow() serializes the entire current FlowCreateRequest
verbatim into the "Current flow" section of the PLAN/CONNECTIONS prompts,
including every LoopContainer's guardSubFlow with its real internal
block ids and names. In FIX mode the model sees this and tries to wire
connections directly to/from the guard scaffold, thinking it's an
editable part of the flow. Those connections can never resolve at the
model's scope and get silently dropped (safe, but the repair round is
wasted chasing something that was never real instead of fixing the
actual reported error).
Fix: summarizeFlow now walks the serialized flow and replaces every
container's guardSubFlow with a short backend-managed marker before
handing it to the model, so the guard mechanism - fully described to
the model via guardCondition/maxIterations/feedbackInput already -
never appears as something to reference or connect to.
Verified by mutation testing (disabling the redaction call reproduces
the leak in the new test). Full suite: 440/440.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Answers "perché il fix non riesce a risolvere il problema?" - traced via
a live repro (session log + direct save attempt) why a LoopContainer
body left in a fully closed 2-block cycle (no input left open for
guard feedback) never got fixed even after both repair rounds ran.
Root cause: when a container-level validation error isn't attributable
to any specific inner block, normalizeContainerInnerBlocksAgainstExisting
force-ADDs every inner block each round (the existing "a FIX round that
changes nothing would never fix the error" fallback) - even though the
model marks them KEEP. But resolveExistingBlock's position-based
fallback still matches the old block regardless of the forced operation,
so oldInnerIdToAssembled/oldNodeIdToAssembledNode got populated anyway,
and preserveConnections carried the OLD (still-broken) connections
forward every round on top of whatever the fresh connections-for-
container call produced. The model has no way to ask for a connection's
removal, so the stale, well-formed (non-dangling) connection survived
indefinitely - the previous dangling-reference fixes didn't catch this
because nothing here is dangling, it's a stale-but-valid connection.
Fix: only populate the old-to-new node id mapping used for connection
preservation when the block/container's operation is genuinely KEEP or
UPDATE, never ADD - regardless of why it became ADD (explicit model
choice or the force-regen fallback). Applied consistently at all three
sites (top-level blocks, top-level containers, container-inner blocks).
Verified by mutation testing: reverting the container-inner guard
reproduces the exact live failure (2 connections instead of 1) in the
new regression test; restoring it fixes it. Full suite: 439/439.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The previous fix made preserveConnections drop a stale carried-forward
connection. That closes the one reproduced cause, but the same failure
mode (a connection referencing a handle name that no longer exists,
which FlowDataValidator treats as a hard, save-blocking structural
error rather than a soft "not executable" one) could in principle recur
through a different code path we haven't hit yet.
Add dropDanglingConnections as a last-resort check applied to the fully
merged connection list (preserved + freshly generated), at both the
top-level flow and every container subflow, right before it's handed
to FlowData.builder(). It drops any connection whose source/target
node or handle name doesn't resolve against the current node graph,
and logs a warning so a future occurrence is visible instead of only
surfacing as a save failure. Dropping a connection whose handle name
was never real does not change what a node exposes, so this is safe
regardless of where the staleness originates.
This is a generic backstop, not a fix for a newly found bug - existing
tests already cover the reproduced scenario and continue to pass
unchanged (438/438), now exercising two redundant layers instead of one.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Root-caused a live "Software Development Loop" flow that could not be
saved (400, CONNECTION_TARGET_INPUT_NOT_FOUND on the LoopContainer's
subFlow). Reproduced end-to-end via the running service + assistant
response log:
1. The validation error was reported at the container level (no inner
block id attached). normalizeContainerInnerBlocksAgainstExisting
resolves every inner block to KEEP in that case (none match the
error), which trips the "a FIX round that changes nothing would
never fix the error" fallback and forces a full regen - every inner
block gets reconfigured fresh even though the model explicitly
marked them KEEP.
2. One reconfigured block got a different prompt placeholder, so its
actual input name changed. preserveConnections still carried the
OLD connection forward (remapped only by node id, not by whether
the endpoint's I/O still existed), and mergeConnections unions old
and new by exact (source,name,target,name) key, so the stale
connection survived alongside the correct new one - a dangling
reference to a since-removed input, rejected at save time.
3. Separately, an assistant-guessed feedbackInput ("previous_response")
that didn't match any actual body input was kept verbatim instead
of falling back to inference, another guaranteed validation failure
on the same flow.
Fix: preserveConnections (shared by top-level and container-subflow
reassembly) now drops a carried-forward connection whose source/target
no longer expose that output/input name. LoopContainerConfiguration's
feedbackInput is normalized against the assembled body's actual open
inputs, discarding a non-matching guess so the factory's own inference
(or a clear ambiguity error) takes over instead of a guaranteed-wrong
value.
Tests reproduce the exact scenario: an all-KEEP FIX round that forces
full regen must not leave a stale connection behind, and a mismatched
feedbackInput must not block an otherwise-valid single-input loop body.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The planning and block-config models never saw which MCP servers were
declared, so they could not reach for tool-running steps (writing to a
workspace, compiling, running shell) - the coding-agent-mcp server was
effectively invisible and the assistant fell back to HTTPServerCall.
- Inject the declared MCP catalog (id/name/description) into both the
PLAN and BLOCK_CONFIG prompts, and steer MCPAgent/MCPAgentChat toward
running declared-server tools instead of emulating them via HTTP.
- Let the model bind a concrete catalog server via mcpServers; validate
chosen serverName against the catalog and drop hallucinated ids so a
bad choice degrades to a clean validation/fallback path instead of a
runtime "Unknown MCP server" failure. The model's explicit choice is
preserved and no longer overwritten by the rag default.
Tests: catalog reaches both prompts; a chosen catalog server is kept;
an unknown server id is dropped.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Live-testing the assistant (now that the empty-plan fix lets qwen3 produce a
real plan) surfaced a 500 on a plan containing a HumanDecisionBlock. Two
causes, both fixed:
1. IODescriptor.equals/hashCode NPE'd when name was null (name.equals /
name.hashCode). FlowDataValidator.validateBlock compares block outputs via
IODescriptor.equals, so a null-named output made the @ValidFlowStructure
ConstraintValidator throw -> Hibernate HV000028 -> HTTP 500 instead of a
clean validation error. Made both null-safe with java.util.Objects. Now a
malformed config is reported as an invalid flow (and repaired), never a 500 -
a general robustness fix for any flow, not just assistant-generated ones.
2. The null-named outputs came from the model emitting HumanDecision options as
{label, value} instead of {name, label} (name is the routing key / branch
output). buildBlock now normalizes options (normalizeHumanDecisionOptions):
a missing option name is filled from its "value" (the intended routing key)
or a slug of its "label", so every branch gets a real, non-null output name
and the flow is valid.
Tests: draftNormalizesHumanDecisionOptionNamesFromValueOrLabel (options given
as {label,value} and {label} -> names "yes"/"reject", no null outputs, valid).
Full suite green (434).
Two robustness fixes found by testing the assistant live with the default
planning model (qwen3:14b), which produced a trivial single-block flow.
1. Empty-plan recovery. A reasoning model under Ollama's format=json can
return a degenerate "{}" (it cannot emit its <think> block under the JSON
grammar). invokeStructuredProvider accepted "{}" as a valid structured
response, so the plan came back empty and validateAndNormalizePlan applied
its single-LLMBlock fallback - masking the failure as a trivial flow. A
degenerate JSON body ("{}", "[]", "") is now treated like a blank response
and falls through to the non-json generate call, where such models emit the
real content (parsed by the balanced-brace extractor). The single-block
fallback stays as a genuine last resort. Fixes the trivial-flow symptom for
reasoning planning models without a config change.
2. Missing required text field. When the model omits a required free-text
config field (observed: MCPAgentChatBlockConfiguration.goalDescription),
deserialization failed with a hard 502. buildBlock now fills any required,
non-enum, non-structural string field the model left blank with a default
derived from the block's purpose (ensureRequiredTextDefaults), so an
omission degrades to a sensible default instead of a 502.
Tests: draftRecoversRealPlanWhenJsonModeReturnsEmptyObject (generateJson
returns "{}", generate returns the real 2-block plan -> flow has 2 blocks, not
the fallback); draftDefaultsMissingRequiredTextFieldInsteadOf502 (MCPAgentChat
with no goalDescription -> valid, goalDescription defaulted to the purpose).
Full assistant suite green; only the known Loop-container timing flake fails
under full-suite load (passes in isolation).
MCP shared-session ordering was enforced with a fake data connection: the
consumer's prompt carried a magic ${{state_ready}} placeholder to synthesize
an input, and the producer's output was wired into it purely to force
execution order - the consumer never actually used that data (the real
shared state is the MCP session, accessed by name). This replaces that hack
with a Dependency, the mechanism meant exactly for ordering-without-data.
No engine change: the executor already gates readiness on dependencies and
the validator already counts them for reachability; MCPAgent uses the
default activity() capabilities so it can be a dependency source/target.
- FlowAssistantService: after assembly, generate ordering dependencies
deterministically for all three cases:
- top-level chain: Dependency(producer, consumer)
- within-container chain: Dependency(producer, consumer) in the subflow
- cross-boundary (top-level producer -> consumer inside a container):
Dependency(producer, container) - the container runs after the producer
and inherits its session.
Existing dependencies are preserved on REFINE (preserveCurrentDependencies).
Removed completeRequiredSequentialConnections, validateSharedMemorySemantics
and their now-unused helpers (isReachable, summarizeMcpSessionBlocks); the
ordering no longer flows through connections.
- Prompt: removed the ${{state_ready}} instructions; the model is told the
backend orders the consumer after the producer via a dependency and must
not add a placeholder or connection for it.
- Tests: updated all MCP mocks to stop emitting state_ready (prompt and
connections) and assert the dependency instead (top-level, within-container,
cross-boundary). Zero state_ready references remain.
Full suite green (431).
The assistant tended to close every path with an EndBlock (e.g. the
Jensen flows), which hides the actual result: an EndBlock consumes its
input and produces no output - the value is kept only as an outcome
payload, not as a normal flow output. The flow's readable result is
whatever block outputs are left open (unconnected), per
ExecutionContext.completeStep (which only adds an output to the result
map when !output.isConnected()).
- BLOCK-CHOICE GUIDE: EndBlock is a labelled terminal marker used only to
record which terminal state a branched path reached (HIRED vs REJECTED)
or to close a dead branch - never for the main result, and not needed
to "finish" a linear flow (EndBlock is optional; nothing requires it).
- CONNECTIONS rules: leaving the result-producing block's output open is
intentional and correct - that open output is the flow's readable
result; don't wire it into an EndBlock just to terminate it.
Prompt-only change; full assistant suite still green (39).
The three optional assistant refinements.
FIX-path incremental container diffing:
- validateAndNormalizePlan now diffs a container's inner blocks against
the existing subflow in FIX mode too, not only REFINE - so a container-
internal fix reconfigures only the inner block(s) the model marks,
reusing the rest. Container-internal validation errors are remapped to
the container id (the inner block id is lost), so if a FIX round would
keep every inner block unchanged (the model gave no guidance) it falls
back to full regeneration, guaranteeing the fix actually happens.
Cross-boundary MCP shared sessions (top-level producer -> inner consumer):
- Verified the runtime supports it: createAndStartSubflowChild (Iterator/
Loop) and GenericContainerExecutor thread the parent's execution-variable
descriptors (the MCP session registry) into a subflow and propagate them
back, so a top-level producer's session is available to a consumer inside
a container that runs after it.
- FlowExecutionValidator.collectErrors is now external-session aware:
collectErrors(subFlow, externalSessions) threads the set of sessions
produced by a top-level block reachable before the container, and a
consumer is satisfied by such an external producer. This replaces the
separate per-scope subflow pass for main subflows (which validated them
in isolation and would reject a valid cross-boundary chain); loop guard
subflows keep a dedicated external-aware pass since the recursion doesn't
cover them. Bounded to top-level producers - inner-producer cross-boundary
chains stay in-scope-only.
- Prompt: relaxed the "never split across a boundary" rule to permit a
top-level producer wired to run before the container holding the consumer.
Multi-open-input LoopContainer bodies:
- AssistantContainerPlan gains feedbackInput; the LoopContainer config now
passes it through, so a body exposing several open inputs resolves the
otherwise-ambiguous feedback-target inference. Prompt updated.
Tests: fixReusesUnchangedInnerBlocksWhenTheModelMarksThem,
draftWiresMcpChainFromTopLevelProducerToConsumerInsideAContainer,
draftAuthorsLoopContainerWithMultiInputBodyAndExplicitFeedbackInput. Full
suite green (431). Roadmap: optional refinements marked done - only bias
annotation authoring (a separate product track) remains.
The assistant could author Generic and Iterator containers but not
LoopContainer ("repeat the subflow until a condition is met"). The
guard subflow has a rigid structural contract - it must expose a
non-multiple boolean output named "guard" and a non-multiple text
output named "feedback", produced through SwitchBlocks, reading the
body's latest result via the special ${{outputs.<name>}} reference.
Having the LLM author that freely would be fragile, so the assistant
supplies only the semantic intent and the backend builds the structure.
- AssistantContainerPlan gains guardCondition (natural-language stop/
continue rule, required for a LoopContainer) and maxIterations;
validateAndNormalizePlan accepts "LoopContainer" and requires a
guardCondition when the body is (re)built.
- assembleContainer builds the body subflow like any container, then
buildLoopGuardSubFlow deterministically assembles the guard subflow:
a Guard Evaluator LLM (reads ${{outputs.<bodyOutput>}}, answers
true=continue / false=stop per the guardCondition) -> an expose-guard
SwitchBlock (boolean guard); a Feedback Builder LLM -> an
expose-feedback SwitchBlock (text feedback) - mirroring the proven
guard structure in ExecutionTest. feedbackInput is left for the config
to infer (the body must expose one open input); the Feedback Builder
deliberately doesn't reference ${{outputs.x}} to avoid a two-block
exposed-input collision.
- Prompt: when to prefer LoopContainer, guardCondition/maxIterations,
and that the body must expose exactly one open input (receives
feedback each iteration). LoopContainer added to the plan schema.
Regression test draftAuthorsLoopContainerWithDeterministicGuardScaffold:
a one-body-block LoopContainer round-trips to a valid flow whose guard
subflow has the four scaffold blocks; passing validation proves the
scaffold satisfies the LOOP_GUARD/LOOP_BODY contracts. Full suite green
(428). Roadmap: item 2 marked done - all three container types
(Generic/Iterator/Loop) are now authorable.
Two related container-scope assistant-correctness gaps (roadmap items 5 and 6).
Item 5 - global input multiplicity and container-inner globals:
- The [] array marker now works symmetrically on globals end to end.
ExecutionTemplateResolver emits a ${{global.name[]}} substitution key
alongside ${{global.name}} (as it already did for normal inputs), so an
array-marked global renders (a list value is newline-joined).
- FlowExecutionValidator strips a trailing [] from an extracted global
reference name so ${{global.cvs[]}} matches the declared global "cvs".
- collectGlobalInputs scans top-level blocks AND every container's inner
subflow blocks; a []-marked reference declares the global multiple=true.
A global referenced only inside a container is now declared both at the
top level (so the flow collects its value) and in the container subflow's
own globalInputs (collectGlobalInputsFromBlocks) - the subflow is
validated as its own scope and, at runtime, the container executor feeds
it the parent's globals by name. This also fixes a pre-existing gap where
a global referenced only inside a container failed GLOBAL_INPUT_NOT_DECLARED.
- Prompt: "never write ${{global.name[]}}" replaced with "append [] when the
global holds a list".
Item 6 - within-container MCP shared sessions:
- Verified via GenericContainerExecutor that a container subflow inherits the
parent's execution-variable descriptors (which hold the MCP session
registry) and propagates them back, so an MCP chain inside one container
genuinely shares state at runtime.
- assembleContainer computes isSharedMemoryContext for the container's inner
plan and threads it into inner buildBlock / completeRequiredSequential-
Connections / validateSharedMemorySemantics - the same normalization the
top level uses, scoped to the subflow (was hardcoded false).
- FlowExecutionValidator.validateSubFlowSharedExecutionVariableOrdering
validates each container subflow's shared sessions as its own scope, so a
broken inner chain is caught at validation time.
- Prompt: a shared-session MCP chain must stay within one scope (all
top-level or all inside the same container), never split across a boundary.
Still out of scope (documented): MCP chains split across a container boundary.
Tests: resolvesGlobalVariableWithArrayMarker (resolver);
draftDeclaresArrayGlobalInputFromArrayMarker,
draftDeclaresGlobalReferencedOnlyInsideAContainerInnerBlock,
draftWiresMcpSharedSessionChainInsideAContainer (assistant). Full suite green
(427). Roadmap doc: items 5 and 6 marked done.
A container was previously either KEEP (reused whole) or ADD/UPDATE
(its entire inner block set regenerated from scratch) - the same waste
and churn risk targeted repair already removed for top-level blocks,
one level down. A REFINE like "add a step inside the review container"
re-authored every existing inner block.
- Extracted reusable cores from the top-level machinery so inner
assembly shares identical rules: resolveExistingBlock(blockPlan,
existingBlocks, ...) and preserveConnections(sourceConnections, ...).
- validateAndNormalizePlan resolves the existing container once, then in
REFINE + UPDATE diffs inner blocks against the existing container's
subflow (normalizeContainerInnerBlocksAgainstExisting /
normalizeInnerBlockOperation): unchanged -> KEEP, changed -> UPDATE,
new -> ADD, dropped -> REMOVE.
- assembleContainer takes the existing container: KEEP inner blocks are
reused verbatim (no BLOCK_CONFIG call) and existing inner connections
between surviving blocks are preserved and merged with generated ones.
- Prompt: the inner "operation" field is now documented
(KEEP/ADD/UPDATE/REMOVE when a container is UPDATEd; omit for
ADD/DRAFT), replacing the old "do not set operation on inner blocks".
Deliberately scoped to REFINE: FIX and ADD still fully regenerate a
container's inner blocks. FIX on purpose - a container-internal error is
most reliably fixed by rebuilding, and inner-block errors don't reliably
carry the inner block's id for targeted attribution, so incremental FIX
risked "KEEP everything and fix nothing". Left as a possible future
refinement (item 7 in the roadmap).
Along the way normalizeContainerOperation was replaced by
resolveContainerOperation (takes the already-resolved existing
container) so the same match is reused for inner diffing.
Regression test refineReusesUnchangedInnerBlocksWhenUpdatingAContainer:
REFINE a 2-block container to add a third; the mock throws if a KEEP
inner block is ever reconfigured, and the two kept blocks come back by
id with unchanged config and the pre-existing inner connection
preserved. Full suite green apart from the known Loop-container timing
flake (passes in isolation). Roadmap doc: item 3 marked done.
The assistant could group blocks into a GenericContainer but could not
express "run this subflow once per element of a list" - the per-item
iteration pattern (score each CV, process each document) that had to be
built by hand earlier this session.
- AssistantContainerPlan gains an `iterationInput` field, and
validateAndNormalizePlan now accepts containerType "IteratorContainer"
alongside "GenericContainer" (iterationInput is kept only for the
former, dropped for the latter).
- assembleGenericContainer is generalized into assembleContainer: it
builds the inner subflow exactly as before, then branches on
containerType to construct either a GenericContainerConfiguration
(GenericContainerFactory) or an IteratorContainerConfiguration with
the iterationInput (IteratorContainerFactory). The factory infers
iterationInput when the subflow has exactly one open input, and
IteratorContainerInterfaceResolver auto-promotes the iterated input
and every exposed output to multiple on the external interface - the
assistant never declares the array-ness itself.
- The connection-wiring machinery needed no changes: a container is
already a FlowNode endpoint, so an IteratorContainer slots into the
same path.
- Prompt updates: the PLAN schema shows an IteratorContainer example
with iterationInput; rules teach when to prefer it over N duplicated
blocks, and that the inner block must use a single-value placeholder
${{cv}} (never ${{cv[]}}) named by iterationInput. CONNECTIONS rules
note the iteration input takes the whole array and outputs are arrays.
- A wrong iterationInput is caught by IteratorContainerConfiguration's
bean-validation on the FlowExecutionValidator pass (wired in earlier
this session), triggering a repair round rather than shipping broken.
Bug found and fixed along the way: validateAndNormalizePlan treated
"no top-level blocks" as an empty plan and fell back to a single
synthetic LLMBlock, silently discarding every container when a plan put
all its work inside containers. The empty-check now requires both blocks
and containers to be empty before falling back, and the block loop is
null-safe for container-only plans.
Regression test draftAuthorsIteratorContainerForPerItemWork: a
container-only plan round-trips to a valid flow whose IteratorContainer
exposes cv/response as multiple while the inner cv input stays
non-multiple. Full suite green (422). Roadmap doc updated: items 1 and 4
marked done, item 3 (incremental container diffing) is now the top open
gap.
isTargetedBlockRepairEligible previously bailed out to a full replan
whenever currentFlow had any containers at all, because
buildReusedPlanForTargetedRepair only reconstructed the top-level block
list - running it on a flow with containers would have silently
dropped them from the rebuilt plan.
buildReusedPlanForTargetedRepair now also reconstructs an
AssistantContainerPlan (operation KEEP, empty inner blocks) for every
container in currentFlow.flow().getContainers(), so containers survive
the targeted-repair path unchanged instead of disappearing. The
eligibility check still requires every error to be scoped to an
existing top-level block - an error scoped to a container (or anything
else) still falls back to a full replan, since a KEEP-only container
plan can't fix a container-internal problem.
Added a regression test: a flow with a container present has a
top-level block's overlong EndBlock.outcomeLabel fixed via a single
targeted BLOCK_CONFIG call (PLAN/CONNECTIONS both skipped), and the
container comes back byte-for-byte identical (same id, same inner
subflow).
Updated docs/assistant-completion-roadmap-2026-07-25.md to mark this
item done and note what's still open (container-internal errors are
still not targeted-fixable - that's the separate, larger "incremental
container diffing" item).
The assistant was structurally blind to containers: the plan/connections
JSON schema was flat, and every connection-resolution helper was typed
to Block<?>, so a container could never be a plan entry or a connection
endpoint. This wires up a first, deliberately scoped version: only
GenericContainer (plain grouping, no iteration/loop semantics), one
level deep (nesting stays unsupported), inner blocks always fully
(re)generated rather than incrementally diffed across FIX/REFINE
rounds.
- AssistantFlowPlan gains a "containers" array alongside "blocks";
each entry (AssistantContainerPlan) carries its own nested "blocks"
list using the exact same shape as top-level blocks.
- validateAndNormalizePlan now normalizes containers too: id
uniqueness (shared namespace with block ids), containerType must be
exactly "GenericContainer", inner blocks non-empty with unique ids
and no nested container types, KEEP/UPDATE/REMOVE resolved against
currentFlow.getContainers() the same way blocks resolve against
currentFlow.getBlocks().
- The connection-resolution machinery (toValidConnections, toConnection,
resolveConnectionBlock, inferBlockByIo, resolveConnectionOutputName/
InputName, registerNodeAlias) is retyped from Block<?> to the shared
FlowNode interface (Block and Container both implement it, exposing
id/name/inputs/outputs) - a mechanical, behavior-preserving change
that lets containers slot into the exact same connection-wiring path
blocks already used, with no new logic needed there.
- assembleGenericContainer assembles a container's inner blocks (their
own BLOCK_CONFIG calls) and inner connections (a CONNECTIONS call
scoped to just that container, reusing buildConnectionsPrompt with a
synthetic inner AssistantFlowPlan), then hands the resulting subflow
to GenericContainerFactory, which auto-derives the container's
exposed inputs/outputs from whatever inner handles are left
unconnected - the assistant never has to declare them itself.
- Prompt updates: the PLAN schema/rules explain the "containers" array,
when grouping is actually useful, and that only GenericContainer is
currently authorable; the CONNECTIONS rules clarify that containers
are valid endpoints via their own id and auto-derived I/O, and that
inner blocks are never reachable from outside their container.
- isTargetedBlockRepairEligible now bails out whenever currentFlow has
any containers, since the targeted-repair plan reconstruction only
rebuilds the block list and would otherwise silently drop them.
Added a regression test covering the full path: a plan with a
top-level block plus a 2-block container, verifying the container's
subflow, its auto-derived exposed input/output names, and the single
top-level connection into it.
Deliberately out of scope for this pass: IteratorContainer/LoopContainer
(need iteration/guard semantics the assistant doesn't author),
incremental inner-block diffing across FIX/REFINE rounds, and nested
containers (still architecturally unsupported).
Every FIX round previously re-ran the whole pipeline - PLAN, every
non-KEEP block's BLOCK_CONFIG, and CONNECTIONS - even when the only
problems were plain bean-validation errors scoped to one existing
block's own configuration (e.g. an overlong EndBlock outcomeLabel, or
the placeholder-naming/BranchRejoin-size violations added earlier this
session). That's wasted LLM calls and unnecessary churn risk on parts
of the flow that were already correct.
FlowAssistantService.assembleFlow now checks, per FIX round, whether
every reported error is cleanly scoped to an existing block and none
of them is in PLAN_SENSITIVE_ERROR_CODES (anything that could imply the
flow's shape needs to change - connections, dependencies, block/
container I/O, lanes, BranchRejoin fan-in, shared-session wiring, ...).
When that holds:
- PLAN is skipped entirely; the plan is rebuilt locally from currentFlow's
existing blocks (UPDATE for the flagged ones, KEEP for the rest) with
no LLM call.
- CONNECTIONS is skipped entirely; existing connections are carried
forward unchanged.
Only BLOCK_CONFIG still runs, and only for the flagged blocks - the
same targeting that already happened via KEEP/UPDATE, now extended to
skip the two other phases too.
This is deliberately conservative: any error outside the safe set, or
any error not attributable to one existing block, falls back to the
existing full-repair behavior unchanged. If a "safe" fix unexpectedly
changes a block's derived I/O anyway (e.g. a factory that derives ports
from placeholder text), the next round's validation will surface a
sensitive code and force a normal full repair - never worse than
before, just self-correcting one round later.
Added a regression test that only mocks BLOCK_CONFIG (not PLAN/
CONNECTIONS) for a block-scoped-only error, so it fails loudly if the
skip logic regresses.
The assistant could already be told (informally) about ${{global.x}}
via FlowExecutionValidator's GLOBAL_INPUT_NOT_DECLARED check, but had
no way to satisfy it: nothing in assembleFlow ever added an entry to
FlowData.globalInputs, so any block referencing a global would always
fail validation. This is exactly why "test cv ranking" needed a manual
fix earlier - the assistant has no way to reach for a real global input
today and instead duplicates the same input across blocks.
- FlowAssistantService.assembleFlow now scans every assembled block's
specificConfiguration (serialized, then regex-scanned with the same
${{global.x}} pattern FlowExecutionValidator uses) and auto-declares
a FlowData.globalInputs entry for each distinct name referenced,
defaulting to TEXT/non-multiple. Existing declarations from
currentFlow (REFINE/FIX, or a human-edited multiplicity) are
preserved rather than overwritten.
- FlowAssistantPromptService: added rules + a worked example teaching
when to use ${{global.name}} (same value needed unchanged by 2+
blocks) versus a normal placeholder (data produced by another block,
wired via CONNECTIONS). Noted the [] array marker does not apply to
globals (ExecutionTemplateResolver only substitutes the bare
${{global.x}} key), and that a global reference has no input port to
connect to in the CONNECTIONS step.
- Added a regression test where two blocks share ${{global.cvs}}: the
flow ends up valid, with exactly one auto-declared global input and
only the genuinely inter-block connection wired.
ExecutionsService's ExecutionObject cache only calls shutdown() (which
tears down that execution's dedicated Executors.newFixedThreadPool) on
eviction, gated by app.executions.cache.max-size (default 1000) and
app.executions.cache.final-ttl-ms (default 30 minutes). The Spring test
context - and its singleton ExecutionsService - is reused across the
whole suite, and a full run finishes in well under 30 minutes with far
fewer than 1000 executions, so none of that cleanup ever fires: every
execution's thread pool stays alive simultaneously for the rest of the
run. That's a plausible contributor to the loop-container tests
occasionally missing their polling deadline under full-suite load
(they pass reliably in isolation, where far fewer executions
accumulate). Lowering both knobs for tests only keeps the live
thread-pool count small throughout a run; eviction only ever targets
already-finished executions (see evictIfNeeded/evictFinalStateByTtl),
so this cannot disrupt an in-flight test.
The assistant's validate step only ran jakarta bean-validation on
FlowCreateRequest, which cascades into FlowDataValidator (the
@ValidFlowStructure constraint on FlowData) but never into each
block's own specificConfiguration - Block.specificConfiguration has
no @Valid annotation. So per-block constraints (BranchRejoinBlockConfiguration's
2-10 input bound, LLMBlockConfiguration's placeholder-name regex, MCP
shared-session producer/consumer wiring, container subflow rules via
the execution-level checks, ...) were never actually enforced: the
assistant could emit a structurally-broken block and still report
valid=true.
- FlowAssistantService.validate() now also runs
FlowExecutionValidator.collectErrors(), merging its errors with the
existing bean-validation errors so the repair loop sees (and can fix)
what used to slip through silently.
- Default max repair attempts raised from 1 to 2 (still capped at the
existing max of 3), since a single fix round often isn't enough once
more error classes are actually being caught.
- Fixed draftDoesNotFailWhenSharedMemoryFlagsAreIncomplete: its flow was
already broken (two MCP consumer-only blocks referencing a shared
session neither produces) and is now correctly reported invalid
instead of a false-positive valid=true.
- Added two regression tests locking in the new coverage and the new
default repair budget.
The assistant's block catalog and prompts predated this session's new
conventions and had a structural gap: BranchRejoinBlock's "inputs" field
was blanket-excluded as if it were a runtime-derived IO list, so the
assistant could never declare a custom branch fan-in. Also x-ui-tip
(the field documenting the ${{name[]}} array marker) was dropped, and
the hardcoded BLOCK-CHOICE GUIDE only listed ~8 of the ~14 block types.
- BlockCatalogService: surface x-ui-tip on prompt field descriptors;
stop excluding "inputs"/"outputs" when they're marked x-ui-structural
(BranchRejoinBlockConfiguration's declared branch list) rather than
a derived IO list.
- FlowAssistantPromptService: generate the BLOCK-CHOICE GUIDE from the
live catalog instead of a stale hardcoded list; add explicit rules
for placeholder naming, the [] array marker, and BranchRejoin's
inputs contract; add small few-shot examples to the plan/config/
connections prompts.
Deliberately not teaching ${{global.name}}/global inputs yet - the
assistant's assembly pipeline has no code path to declare a flow-level
global input, so advertising the syntax would produce broken flows.
That's implementation-level work, kept separate.
Simulated HumanDecision required the LLM's "CHOICE:" line to be exactly a
configured option name. But the simulate prompt lists options as
"- name: Label", so models - including strong instruct ones (qwen2.5:14b,
etc.), not just gemma:7b - routinely echo the whole "name: Label" line, e.g.
"red: Red light" or "existing: Existing position". That was rejected with
HUMAN_DECISION_INVALID_CHOICE, breaking simulated runs of any flow with
human decisions.
resolveSimulatedChoice() now accepts, in order: an exact option name, the
token before the first ':' (the echoed "name: Label" case), or the option
label. Only ':' is treated as a separator - never '-' - so hyphenated option
names like "not-red"/"assessment-not-required" are never truncated. Unknown
choices still return null and raise the same error.
Added a direct unit test (HumanDecisionSimulatedChoiceTest) covering exact
name, echoed name:label, hyphenated names, label text and rejection. Also
narrowed the flat-flow Jensen smoke test to exclude the new container-grouped
variant (different top-level shape).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A GenericContainer whose subflow takes only one of several internal branches
left the other branches' exposed outputs unproduced. The completion path
(reconcileGenericSubflow / GenericContainerExecutor.completedResult) returned
NodeExecutionResult.completed(valueMap) with an empty notSelectedOutputs, so
ExecutionContext.completeStep marked every unproduced exposed output
UNAVAILABLE. Downstream that reads as an error signal, not a skip: a
BranchRejoin (which self-skips only when ALL inputs are NOT_SELECTED) instead
saw UNAVAILABLE, went READY and threw BRANCH_REJOIN_INPUT_UNAVAILABLE, failing
the whole execution. A plain routing block never hits this because it reports
its not-taken outputs via notSelectedOutputs.
Both container completion paths now return NodeExecutionResult.routed(values,
allExposedOutputNames): a successful subflow that didn't take an internal
branch surfaces that branch's exposed output as NOT_SELECTED, so downstream
steps and BranchRejoins skip cleanly - letting a container act as a real
branching node (early-stop outcomes exposed as outputs routed to top-level
EndBlocks). Containers with a single always-produced output are unaffected
(that output is in the value map, notSelected is empty).
Found while building the containerized Jensen recruitment flow, where phase
containers expose early-stop outputs; added a focused regression test
(genericContainerUnselectedExposedOutputsSkipDownstreamCleanlyInsteadOfErroring)
and the "Jensen Recruitment Process - Full Revised (Containerized)" seed flow
that groups the 12 swimlanes into 7 phase containers (13 top-level nodes
instead of 39).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Executed the full improvement plan for this block, found genuinely
under-baked (zero real usage, schema-only test coverage):
- Fixed a real bug: a JSON-typed output returned the raw trimmed string,
but Input.validateValue() requires JSON-typed values to be instanceof
JsonNode - configuring a JSON output was silently broken end to end.
coerce() now parses the segment with ObjectMapper.readTree(), with a
clear error on invalid JSON instead of a confusing downstream rejection.
- Replaced raw IllegalArgumentException with NodeExecutionException and
stable error codes (DELIMITED_PARSER_INPUT_MISSING,
_SEGMENT_COUNT_MISMATCH, _INVALID_BOOLEAN, _INVALID_FILE,
_INVALID_JSON), matching the convention already used by
BranchRejoinExecutor.
- Added an array-output mode: a DelimitedParserOutput (new record, kept
separate from the shared SwitchCase used by SwitchBlock) marked
multiple=true collects every split segment into a list with no
fixed-count check, instead of requiring exactly N named outputs -
complements the Iterator's "array in" with "array out" here.
- Boolean coercion now also accepts yes/no and 1/0, not just true/false.
- Added DelimitedParserExecutorTest (11 cases, was zero), two end-to-end
ExecutionTest cases (fixed mode and array mode), and a bundled "test
delimited parser" example flow (LLM -> DelimitedParserBlock -> EndBlock)
since the block had never been used in any real flow before. Verified
live against the running service with a real Ollama call.
No rename needed here - "DelimitedParserBlock" already describes what it
does, unlike ExclusiveMergeBlock.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
"Merge" was misleading: the block never combines values, it only lets
exactly one of several mutually-exclusive branches (an ExclusiveOR
reconvergence, e.g. after a HumanDecisionBlock split) pass through as the
single downstream value; two arriving values is an error, not something to
combine. BranchRejoinBlock names what it actually does. There is still no
block in the platform that genuinely merges/combines multiple parallel
values into one - this rename doesn't add one, it just stops "merge" from
implying it exists.
Renamed throughout: block/config/factory/executor/activation-policy classes
and their tests, the NodeVisualRole.MERGE -> BRANCH_REJOIN visual role and
NodeTypeCapabilities.merge() -> branchRejoin() factory method, the six
EXCLUSIVE_MERGE_* validation error codes -> BRANCH_REJOIN_*, the bundled
flows.json (11 uses across the Jensen flows), and the two docs that
reference it.
Added a legacy-id fallback in DynamicBlockConfigurationTypeResolver
(type id "ExclusiveMergeBlockConfiguration") and BlockTypes (typeName
"ExclusiveMergeBlock"), mirroring the existing ChatHumanInteraction ->
ChatInteraction precedent - found necessary the hard way: without it, any
flow already persisted under the old type id fails to deserialize and
brings the whole application down at startup, not just that one flow.
Locked this in with LegacyBlockTypeIdTest.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The CVs array needs to be read by two nodes (rank-cvs and review-ranking)
that aren't wired to each other for it. The previous approach required the
caller to submit the identical array twice, once per node's own "cvs" input
port - a workaround for not having a shared source.
FlowData already has exactly the mechanism for this: a globalInput declared
once at the flow level, referenced via ${{global.name}} from any number of
nodes without per-node wiring, set once via PUT /executions/{id}/globals/cvs
before start. Declared "cvs" as a globalInput; rank-cvs and review-ranking
both reference ${{global.cvs}} and no longer declare a "cvs" Input port at
all.
Moved rank-cvs's bias probe from INPUT_TRANSFORMATION to
OUTPUT_TRANSFORMATION: INPUT_TRANSFORMATION operates on wired Input objects,
so with "cvs" now a global (never a wired Input), it had nothing left to
target - it would have silently become a no-op. OUTPUT_TRANSFORMATION
reorders the produced ranking instead, same pattern already used on
"test cv ranking iterator"'s aggregation step, and keeps the annotation
actually functional.
Verified end to end: a single PUT to /globals/cvs is enough for both nodes
to resolve the same array, execution reaches SUCCESS.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Removes the three fixed collect-cv-1/2/3 HumanInteractionBlock nodes from
"test cv ranking": rank-cvs now declares one input, cvs, via ${{cvs[]}},
fed directly as an array before start (PUT .../input/cvs/texts) instead of
requiring a human to paste each CV into its own node. Scales to any number
of CVs instead of a fixed three.
review-ranking keeps seeing both the ranking and the original CVs (also via
${{cvs[]}} in the question) for the automation-bias check to stay
meaningful - since there's no longer a node whose output naturally fans out
to both consumers, the caller submits the same array to both rank-cvs.cvs
and review-ranking.cvs.
The order-bias probe on rank-cvs now targets the whole "cvs" array: since
INPUT_TRANSFORMATION already recurses per-element for list-valued inputs,
the instruction is applied uniformly to every CV rather than to one
specific candidate as before - a different (still meaningful) experiment,
not a broken one, but worth noting as a real tradeoff of moving from named
per-candidate inputs to a single array input.
Verified end to end against the running service: no human interaction is
needed to provide the CVs, ranking is produced correctly, and review-ranking
resolves both the ranking and the full CVs array.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A Loop/Iterator container step only ever exposes activeInnerExecutionId (the
current iteration's child), overwritten on every iteration. Past iterations'
child executions are never deleted from the DB, but there was no way to
enumerate them - only individually fetchable by id if a caller already had
it (e.g. from the event log).
Adds GET /executions/{id}/node/{stepId}/iterations, returning every child
execution created for that container step (one per iteration for
Loop/Iterator, one per run for GenericContainer), ordered by iteration
index. Backed by a new repository query
(findByParentExecutionIdAndParentStepIdOrderByParentIterationIndexAscCreationTimeAsc)
and ExecutionsService.getContainerIterationsByOwner(), which validates the
step exists and is actually a container before querying.
Verified end to end against the running service on the "test cv ranking
iterator" flow: the endpoint lists all 3 iterations in order with distinct
ids and SUCCESS status, returns 400 for a non-container step and 404 for an
unknown step id, and each listed iteration remains individually fetchable
via the existing GET /executions/{id}.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Placeholders written as ${{name[]}} now derive a multiple:true (array) input
instead of the default single value, on the three block types whose inputs
are parsed from free text: LLMBlockConfiguration.prompt,
HumanInteractiveBlockConfiguration.actionDescription and
HumanDecisionBlockConfiguration.question. No marker -> unchanged single-value
behavior, so every existing flow keeps working as-is.
TemplateInputs now returns a Placeholder(name, multiple) record instead of a
bare name: the "[]" suffix is stripped from the captured group and the
remaining name goes through the same name-validation added in the previous
step. HumanDecisionBlockFactory's mandatory forward-through "input" port is
untouched - only the additive context inputs parsed from the question can be
arrays.
ExecutionTemplateResolver.resolve() now also emits the "${{name[]}}" spelling
as a substitution key alongside "${{name}}" for every value, so a template
written with the marker actually gets the value substituted at render time -
without this, the literal marker was left unresolved in the text.
Added hasConsistentPlaceholderMultiplicity() alongside the existing
hasValidPlaceholderNames(): a name referenced both with and without "[]" in
the same text is now a validation error instead of silently picking one side.
Rebuilt the "test cv ranking iterator" flow (IteratorContainer scoring each
CV independently, then an LLM/HumanDecision consuming the accumulated array
via ${{scores[]}}) using the marker instead of the abandoned explicit-field
design, and verified it end to end against the running service: all three
iterations complete, the aggregation LLM receives the array and produces a
real ranking, and the review step shows both the ranking and the full array
of per-candidate scores.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Nothing today rejects a ${{...}} placeholder whose captured name contains
spaces, symbols, or brackets: the capture group is `.*?`, so it accepts
anything between ${{ and }}. That name flows straight into an IODescriptor
with no further checks.
Adds an @AssertTrue check (same convention as the existing
isOptionsUnique()/areSkillIdsUnique() checks) on LLMBlockConfiguration.prompt,
HumanInteractiveBlockConfiguration.actionDescription and
HumanDecisionBlockConfiguration.question, requiring every placeholder name to
start with a letter and contain only letters, digits, '-', '_' or '.'.
'.' stays allowed because it's already load-bearing: LoopContainer guard
prompts reference inputs.<name>/outputs.<name>
(ExecutionsService.buildGuardTemplateValues), and colliding container-exposed
names get qualified as nodeName.ioName
(ContainerFlowInterfaceResolver.qualifyWithNodeName). Confirmed via the full
suite - the first pass of this change broke 5 Loop container tests before
'.' was added back.
'[' and ']' stay rejected on purpose, reserving that syntax for a possible
future array-input marker on placeholder names.
TemplateInputs made public (was package-private) so the three
BlockConfiguration classes, which live in a different package, can call the
new hasValidPlaceholderNames() check.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Step.completeSuspendedContainer() (the async path that resumes a container
step once its inner subflow/iteration finishes) propagated container
outputs to downstream inputs without the same try/catch guard that
Step.run()'s synchronous path already had. When a downstream input rejected
the value (e.g. a multiplicity mismatch), the exception escaped through the
event-listener callback that drives it, leaving the container step stuck in
RUNNING forever with no recorded error.
Found while prototyping an IteratorContainer-based CV ranking flow: an
Iterator's accumulated (multiple) output wired into a plain LLM input
(always declared non-multiple) reproduced exactly this hang.
Wrapped the same body in try/catch, mirroring run()'s failure handling:
mark the step FAILED and notify the listener with a proper error message
instead of silently stalling.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Three CVs are collected separately, ranked together by one LLM call (three
named inputs feeding a single ranking prompt), and reviewed by a human
decision maker who sees both the ranking and all three original CVs side by
side via the newly added multi-input support on HumanDecisionBlock.
Carries starter design-time bias annotations and probes on the ranking step
(SELECTION_BIAS, INPUT_TRANSFORMATION on one CV) and the review step
(AUTOMATION_BIAS, ROUTING_OVERRIDE) as a base for later bias-injection
experiments on ranking tasks specifically. Verified end to end against the
running service: full run reaches SUCCESS with a real ranking produced by
the LLM and accepted by the reviewer.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Both block types previously exposed exactly one fixed input port, so a human
step could only ever be wired to a single upstream value even when the
decision or task genuinely needs more context to evaluate (e.g. seeing both
the original candidate profile and an automated screening result before
deciding).
Following the same convention LLMBlock already uses for its prompt, extra
named input ports are now derived from ${{name}} placeholders in the task's
free text: actionDescription for HumanInteractionBlock, question for
HumanDecisionBlock. No placeholders -> unchanged single "input" port, so
every existing flow keeps working as-is.
HumanDecisionBlock keeps its "input" port mandatory regardless of
placeholders, since HumanDecisionExecutor forwards that exact value as the
payload of whichever branch is chosen; extra placeholder-derived inputs are
additive context alongside it. HumanInteractionBlock has no such constraint,
so its inputs are fully derived when placeholders are present.
Extracted the placeholder-parsing regex (previously private to
LLMBlockFactory) into a shared TemplateInputs helper reused by all three
factories.
Wires the new capability into the bundled "test biased" flow: shortlist-decision
now also sees the candidate profile alongside the screening assessment, and
the container's reference-check step sees both the candidate profile and the
screening assessment. Verified end to end against the running service.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds a GenericContainer (reference-check-container) between the shortlist
decision and interview prep steps, with an interactive human step feeding an
LLM assessment that carries its own bias annotation and behavioral probe.
Exercises subflow bias propagation (includeSubflow) and interactive nodes
inside containers together on the bundled bias test flow. Verified end to
end against the running service: normal run reaches SUCCESS, and a
bias-rerun with includeSubflow on the container correctly activates and
applies the inner probe.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Closes Tappa B and Tappa C of the interactive-containers plan.
Tappa B (recovery and lifecycle):
- Generalize the GenericContainer-only watcher into a type-dispatching
coordinator (watchContainerSubflow/reconcileContainerSubflow) shared by
Generic, Loop and Iterator.
- Add reconcileContainerSubflowsOnStartup (ApplicationReadyEvent): re-arms
every persisted parent-child link at boot using only the entity's own
parentExecutionId/parentStepId columns, so a child that already finished
while the process was down is reconciled immediately, and a pending one
gets its listener re-armed. No longer relies on any in-memory listener
surviving a restart.
- cancelExecution now propagates in both directions: cancelling a parent
cancels its active container child; cancelling a child fails the
parent's container step (same outcome as an unexpected child error).
removeExecution evicts cached children too.
- Propagate simulation mode to container children: Step gains
executionSimulationEnabled (mirrors interactionSimulationDescriptor's
propagation to every step, containers included); ContainerExecutionContext
carries it plus the descriptor; startContainerChild starts the child via
startSimulationExecution when applicable. ExecutionObject.hasSimulationAvailable
now recognizes interactive nodes nested inside a container's subflow(s),
which it previously ignored entirely since a container is never itself
isUserInteractive().
Tappa C (Loop/Iterator non-blocking):
- LoopContainerExecutor and IteratorContainerExecutor become thin adapters
delegating to ExecutionsService.startLoopSubflow/startIteratorSubflow.
All advancement logic (creating/starting a child, checking whether it
finished synchronously vs. suspending) now lives in ExecutionsService, so
it is reusable both from the initial call and from reconciliation.
- Iterator advances iteration-by-iteration in a loop, chaining synchronous
completions without suspending, and persists remainingValues/
runtimeInputValues/accumulatedOutputs so it can resume exactly where it
left off.
- Loop is a two-phase (MAIN/GUARD) state machine; the phase to resume into
is derived from the completed child's own subflowRole, so it doesn't need
separate persistence. Persists currentInputs/latestOutputs.
- FlowDataValidator and FlowExecutionValidator now accept interactive nodes
in the subFlow of all three container types; LoopContainer's guardSubFlow
remains rejected (guard interactivity is out of scope).
Existing container tests that assumed the old fully-synchronous model
(assert SUCCESS right after start()) are updated: since a child's steps
always run on their own thread pool, even a fully-automatic container now
transiently visits WAITING before the coordinator resolves it, so tests
must poll through WAITING too, not just RUNNING.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A bias variant rerun can now activate a container's inner subflow: with
includeSubflow=true on a container activation, every executable bias
annotation on the subflow nodes (and a LoopContainer's guardSubFlow) is
activated as a block, and the bias context is propagated into the
container's inner execution so those probes fire at runtime.
- BiasActivation gains includeSubflow (default false, backward compatible).
- BiasExecutionContext carries subflowActivatedContainerIds.
- BiasContainerPropagation builds the inner context by scanning the
subflow for executable annotations.
- ContainerExecutor.execute receives the BiasExecutionContext; the three
container executors create the inner execution with the propagated
context (Loop applies it to subFlow and guardSubFlow, per iteration).
- validateBiasActivations validates includeSubflow (container-only, must
have an executable probe) and no longer requires annotationIds when it
is set; new error codes BIAS_ACTIVATION_ANNOTATIONS_REQUIRED,
BIAS_SUBFLOW_ON_NON_CONTAINER, BIAS_SUBFLOW_NOT_EXECUTABLE.
- compareFullFlow treats subflow-activated containers as activated nodes.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Make a bias variant rerun affect human interaction nodes, both when
answered by a real human and when simulated:
- ExecutionStepView now exposes activeBiasProbes (annotationId,
activationMode, instruction) for a step during a BIAS_VARIANT rerun,
so the frontend can surface the injected bias to the interacting user.
It is empty in normal/simulation runs, leaving the contract unchanged.
- PROMPT_DIRECTIVE is now supported on HumanInteraction/HumanDecision.
It biases the simulator prompt and is a no-op when a real human
answers (the directive is shown via activeBiasProbes instead).
- HumanInteractionExecutor.simulate and HumanDecisionExecutor.simulate
now decorate their prompt via BiasRuntimeSupport.decoratePrompt, so an
active prompt directive reaches the simulator.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
EndMode only ever had one value (PATH_END) and wasn't read anywhere,
so it was pure dead configuration. Keep JSON backward compatibility by
ignoring a legacy "mode" field on deserialization instead of failing,
and drop it from the seed flows and docs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
HumanDecisionExecutor explicitly disabled supportsSimulation(), so any
execution containing a HumanDecision node lost the aggregate
simulationAvailable flag entirely, hiding the simulate feature from
clients even though the /simulate endpoint itself was untouched.
simulate() now prompts the simulator LLM to pick one of the configured
options (and a rationale when required), validates the choice, and
returns the same output shape a real human interaction would produce.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
pom.xml declared org.flywaydb:flyway-core directly, but Spring Boot 4 moved
Flyway's autoconfiguration into a separate spring-boot-flyway module: with
only flyway-core on the classpath, Flyway never ran (no migration, not even
flyway_schema_history), so ddl-auto=validate always failed against an empty
schema. Switch to org.springframework.boot:spring-boot-starter-flyway, which
pulls in the actual autoconfiguration.
The "Jensen Recruitment Process - Mitigated (Structured)" seed description
was 269 characters, over the default varchar(255) on flow_entity.description
(no @Column(length=...) override), so FlowImportComponent failed to insert
it while every other flow (all <= 255 chars) imported fine. Shortened to the
same content in 236 characters.
Verified end-to-end against a real Postgres container: fresh schema
bootstrap (ddl-auto=update once, then default validate), all 8 seed flows
import cleanly, full test suite unaffected (362 passing, 2 skipped).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
No source in this codebase actually uses a preview language feature, but
--enable-preview combined with release=25 in pom.xml (compiler + surefire)
was tripping the IDE's language server on every file ("preview can be
enabled only at source level 26"), which read as pervasive warnings across
the whole project. Removed both flags; full suite still passes unchanged.
Also fixes a handful of genuine compiler warnings found via a manual
`javac -Xlint:all` pass: dead unused method in BlockCatalogService, two
uses of JsonNode.isTextual()/asText() (deprecated in tools.jackson 3) in
favor of isString()/asString(), and a Javadoc comment in
ExecutionsController placed after @GetMapping instead of before it (so it
was never attached to the method).
Left the remaining low-value warnings (missing serialVersionUID on a
handful of exception/resolver classes, this-escape in
ExecutionObject/ExecutionContext/Step) untouched per user's choice - purely
cosmetic and, for this-escape, riskier to touch without a concrete reason.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Audits every bullet in "Strategia di test" (unit/integration/regression/
concurrency) against the actual suite and adds coverage for what was
missing: ExclusiveMergeActivationPolicy and ExclusiveMergeExecutor direct
unit tests, dynamic configuration validation, a second-hop NOT_SELECTED
propagation assertion, implicit-merge rejection (EXCLUSIVE_BRANCH_MERGE),
a lanes/laneId JSON round-trip, per-block-type bias capability checks, and
concurrency tests for the merge (repeated concurrent split+merge, cancel
under concurrency, concurrent snapshot reads).
Fixes a latent bug found while writing the configuration-validation tests:
ExclusiveMergeBlockConfiguration.areInputsUnique() and
HumanDecisionBlockConfiguration.areOptionsUnique() never fired because Bean
Validation only recognizes isXxx/getXxx as constrained getters, so the
duplicate-name check was silently unenforced everywhere, including flow
creation. Renamed to isInputsUnique()/isOptionsUnique().
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds "Jensen Recruitment Process - Full Revised (Structured)" and
"... Mitigated (Structured)" seed flows, ported node-per-node from the
original PlantUML activity diagrams using HumanDecisionBlock gateways,
explicit ExclusiveMergeBlock rejoins, EndBlock outcomes and FlowLane
swimlanes. Preserves the 5 CONFIRMED risk annotations and all 16
FH-J1-FH-J16 MITIGATED annotations on the corresponding structured nodes.
Does not replace the existing unstructured macro-flow seeds.
Human-interactive stages are modeled as HumanInteractionBlock rather than
GenericContainer/LLMBlock so both flows can be driven to a real SUCCESS
outcome without an Ollama model. Adds structural validation tests and a
runtime test suite driving all five terminal outcomes on both flows, a
snapshot-restore check during a pending HumanDecisionBlock, and a
baseline/bias-variant comparison showing a ROUTING_OVERRIDE changing both
the branch taken and the final outcome.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>