Commit Graph

244 Commits

Author SHA1 Message Date
Lucio Lelii e823d8d357 fix(assistant): statically sanitize brace/wrapper noise in connection endpoint names
Answers "could the extra-braces problems be fixed statically?" - yes,
the syntactic-noise class can and now is.

The model sometimes mangles a connection endpoint name with purely
syntactic noise: a stray/unbalanced brace ("{category"), an accidental
${{...}} wrapper, or a block-qualified reference ("classify.response").
normalizeBlockReference only did trim()+toLowerCase(), so "{category"
never matched the real "category" input and the connection was silently
dropped, leaving the flow disconnected.

Added stripHandleNoise() - removes ${{ }} / {{ }} wrappers, stray
braces/$/quotes, and a leading block-name qualifier (keeps the last
dotted segment) - and wired it as a FALLBACK in findIoByName and
resolveConnectionBlock: it only runs after the exact-name match already
failed, so it can never change a currently-resolving reference, only
rescue one that would otherwise be dropped. It never invents a name, so
a genuinely-wrong reference (not just mangled) still fails and is
dropped, as before.

Scope note: this fixes the SYNTACTIC class only. Semantic/structural
problems (duplicated logic, connections to non-existent handles,
topology/deadlocks) are unaffected - those need the prompt-side and/or
soft-vs-hard-validation work, not name cleanup.

Test isolates the sanitization path by giving the target two inputs so
the pre-existing single-input shortcut cannot mask it; verified by
mutation that disabling the fallback drops the connection. 442/442.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 16:01:59 +02:00
Lucio Lelii fad1eaf4b4 fix(assistant): stop duplicating loop-body logic at the top level; retry a misplaced container type instead of 502ing
Two more issues found retesting the same live prompt after the
guardSubFlow-redaction fix (confirmed working: no more guard-scaffold
names leaking into prompts).

1. The model kept declaring the same logical step both as a top-level
   block AND as a container's own inner block (e.g. "check completion"
   as both b4 and c1-b1), then tried to wire the orphaned top-level
   duplicate to the container via malformed connections (dotted
   qualified names, null toInput) - all silently dropped as invalid,
   but leaving the duplicate disconnected/deadlocked instead of fixing
   the real problem. Added an explicit NO DUPLICATION rule to the plan
   prompt: a step belongs in exactly one place, top-level or inside one
   container, never both - only the container's own exposed I/O is
   what the rest of the flow should connect to.

2. Separately, a fresh model response put a container type
   ("LoopContainer") as an entry in the flat "blocks" list instead of
   "containers". Block assembly has no catalog descriptor for a
   container type, so this threw as an unrecoverable 502 with no
   retry - unlike degenerate JSON or a missing required field, which
   already get a structured-repair retry. Moved the check into
   validateAndNormalizePlan, inside the plan's own retry-wrapped parser
   callback, so this now gets the same self-correction chance instead
   of hard-failing the whole request.

Verified by mutation testing: reverting the container-type check
reproduces the exact live 502 in the new test; restoring it fixes it.
Full suite: 441/441.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 15:36:10 +02:00
Lucio Lelii f94fb883d9 fix(assistant): redact LoopContainer guardSubFlow from the model's view
Live incident: retrying the same prompt from the UI logged repeated
"Skipping invalid assistant connection draft" warnings referencing
block names like "c1-expose-feedback" and "c1-guard-evaluator" - the
deterministic guard scaffold FlowAssistantService#buildLoopGuardSubFlow
builds and the model never authors.

Root cause: summarizeFlow() serializes the entire current FlowCreateRequest
verbatim into the "Current flow" section of the PLAN/CONNECTIONS prompts,
including every LoopContainer's guardSubFlow with its real internal
block ids and names. In FIX mode the model sees this and tries to wire
connections directly to/from the guard scaffold, thinking it's an
editable part of the flow. Those connections can never resolve at the
model's scope and get silently dropped (safe, but the repair round is
wasted chasing something that was never real instead of fixing the
actual reported error).

Fix: summarizeFlow now walks the serialized flow and replaces every
container's guardSubFlow with a short backend-managed marker before
handing it to the model, so the guard mechanism - fully described to
the model via guardCondition/maxIterations/feedbackInput already -
never appears as something to reference or connect to.

Verified by mutation testing (disabling the redaction call reproduces
the leak in the new test). Full suite: 440/440.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 15:10:27 +02:00
Lucio Lelii 3d16a3b388 fix(assistant): stop resurrecting stale connections across forced-regen repair rounds
Answers "perché il fix non riesce a risolvere il problema?" - traced via
a live repro (session log + direct save attempt) why a LoopContainer
body left in a fully closed 2-block cycle (no input left open for
guard feedback) never got fixed even after both repair rounds ran.

Root cause: when a container-level validation error isn't attributable
to any specific inner block, normalizeContainerInnerBlocksAgainstExisting
force-ADDs every inner block each round (the existing "a FIX round that
changes nothing would never fix the error" fallback) - even though the
model marks them KEEP. But resolveExistingBlock's position-based
fallback still matches the old block regardless of the forced operation,
so oldInnerIdToAssembled/oldNodeIdToAssembledNode got populated anyway,
and preserveConnections carried the OLD (still-broken) connections
forward every round on top of whatever the fresh connections-for-
container call produced. The model has no way to ask for a connection's
removal, so the stale, well-formed (non-dangling) connection survived
indefinitely - the previous dangling-reference fixes didn't catch this
because nothing here is dangling, it's a stale-but-valid connection.

Fix: only populate the old-to-new node id mapping used for connection
preservation when the block/container's operation is genuinely KEEP or
UPDATE, never ADD - regardless of why it became ADD (explicit model
choice or the force-regen fallback). Applied consistently at all three
sites (top-level blocks, top-level containers, container-inner blocks).

Verified by mutation testing: reverting the container-inner guard
reproduces the exact live failure (2 connections instead of 1) in the
new regression test; restoring it fixes it. Full suite: 439/439.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 13:13:20 +02:00
Lucio Lelii 66112541a5 fix(assistant): add a final defense-in-depth filter against dangling connections
The previous fix made preserveConnections drop a stale carried-forward
connection. That closes the one reproduced cause, but the same failure
mode (a connection referencing a handle name that no longer exists,
which FlowDataValidator treats as a hard, save-blocking structural
error rather than a soft "not executable" one) could in principle recur
through a different code path we haven't hit yet.

Add dropDanglingConnections as a last-resort check applied to the fully
merged connection list (preserved + freshly generated), at both the
top-level flow and every container subflow, right before it's handed
to FlowData.builder(). It drops any connection whose source/target
node or handle name doesn't resolve against the current node graph,
and logs a warning so a future occurrence is visible instead of only
surfacing as a save failure. Dropping a connection whose handle name
was never real does not change what a node exposes, so this is safe
regardless of where the staleness originates.

This is a generic backstop, not a fix for a newly found bug - existing
tests already cover the reproduced scenario and continue to pass
unchanged (438/438), now exercising two redundant layers instead of one.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 12:46:50 +02:00
Lucio Lelii 7319443dab fix(assistant): drop stale connections and mismatched feedbackInput on container FIX
Root-caused a live "Software Development Loop" flow that could not be
saved (400, CONNECTION_TARGET_INPUT_NOT_FOUND on the LoopContainer's
subFlow). Reproduced end-to-end via the running service + assistant
response log:

1. The validation error was reported at the container level (no inner
   block id attached). normalizeContainerInnerBlocksAgainstExisting
   resolves every inner block to KEEP in that case (none match the
   error), which trips the "a FIX round that changes nothing would
   never fix the error" fallback and forces a full regen - every inner
   block gets reconfigured fresh even though the model explicitly
   marked them KEEP.
2. One reconfigured block got a different prompt placeholder, so its
   actual input name changed. preserveConnections still carried the
   OLD connection forward (remapped only by node id, not by whether
   the endpoint's I/O still existed), and mergeConnections unions old
   and new by exact (source,name,target,name) key, so the stale
   connection survived alongside the correct new one - a dangling
   reference to a since-removed input, rejected at save time.
3. Separately, an assistant-guessed feedbackInput ("previous_response")
   that didn't match any actual body input was kept verbatim instead
   of falling back to inference, another guaranteed validation failure
   on the same flow.

Fix: preserveConnections (shared by top-level and container-subflow
reassembly) now drops a carried-forward connection whose source/target
no longer expose that output/input name. LoopContainerConfiguration's
feedbackInput is normalized against the assembled body's actual open
inputs, discarding a non-matching guess so the factory's own inference
(or a clear ambiguity error) takes over instead of a guaranteed-wrong
value.

Tests reproduce the exact scenario: an all-KEEP FIX round that forces
full regen must not leave a stale connection behind, and a mismatched
feedbackInput must not block an otherwise-valid single-input loop body.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 12:31:42 +02:00
Lucio Lelii effb506572 feat(assistant): make the MCP server catalog central to block selection
The planning and block-config models never saw which MCP servers were
declared, so they could not reach for tool-running steps (writing to a
workspace, compiling, running shell) - the coding-agent-mcp server was
effectively invisible and the assistant fell back to HTTPServerCall.

- Inject the declared MCP catalog (id/name/description) into both the
  PLAN and BLOCK_CONFIG prompts, and steer MCPAgent/MCPAgentChat toward
  running declared-server tools instead of emulating them via HTTP.
- Let the model bind a concrete catalog server via mcpServers; validate
  chosen serverName against the catalog and drop hallucinated ids so a
  bad choice degrades to a clean validation/fallback path instead of a
  runtime "Unknown MCP server" failure. The model's explicit choice is
  preserved and no longer overwritten by the rag default.

Tests: catalog reaches both prompts; a chosen catalog server is kept;
an unknown server id is dropped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 11:47:15 +02:00
Lucio Lelii 97ad1776a6 fix: null-safe IODescriptor equals/hashCode + normalize HumanDecision option names
Live-testing the assistant (now that the empty-plan fix lets qwen3 produce a
real plan) surfaced a 500 on a plan containing a HumanDecisionBlock. Two
causes, both fixed:

1. IODescriptor.equals/hashCode NPE'd when name was null (name.equals /
   name.hashCode). FlowDataValidator.validateBlock compares block outputs via
   IODescriptor.equals, so a null-named output made the @ValidFlowStructure
   ConstraintValidator throw -> Hibernate HV000028 -> HTTP 500 instead of a
   clean validation error. Made both null-safe with java.util.Objects. Now a
   malformed config is reported as an invalid flow (and repaired), never a 500 -
   a general robustness fix for any flow, not just assistant-generated ones.

2. The null-named outputs came from the model emitting HumanDecision options as
   {label, value} instead of {name, label} (name is the routing key / branch
   output). buildBlock now normalizes options (normalizeHumanDecisionOptions):
   a missing option name is filled from its "value" (the intended routing key)
   or a slug of its "label", so every branch gets a real, non-null output name
   and the flow is valid.

Tests: draftNormalizesHumanDecisionOptionNamesFromValueOrLabel (options given
as {label,value} and {label} -> names "yes"/"reject", no null outputs, valid).
Full suite green (434).
2026-08-03 11:26:35 +02:00
Lucio Lelii 653184eeae fix: recover from empty-plan JSON and default missing required config fields
Two robustness fixes found by testing the assistant live with the default
planning model (qwen3:14b), which produced a trivial single-block flow.

1. Empty-plan recovery. A reasoning model under Ollama's format=json can
   return a degenerate "{}" (it cannot emit its <think> block under the JSON
   grammar). invokeStructuredProvider accepted "{}" as a valid structured
   response, so the plan came back empty and validateAndNormalizePlan applied
   its single-LLMBlock fallback - masking the failure as a trivial flow. A
   degenerate JSON body ("{}", "[]", "") is now treated like a blank response
   and falls through to the non-json generate call, where such models emit the
   real content (parsed by the balanced-brace extractor). The single-block
   fallback stays as a genuine last resort. Fixes the trivial-flow symptom for
   reasoning planning models without a config change.

2. Missing required text field. When the model omits a required free-text
   config field (observed: MCPAgentChatBlockConfiguration.goalDescription),
   deserialization failed with a hard 502. buildBlock now fills any required,
   non-enum, non-structural string field the model left blank with a default
   derived from the block's purpose (ensureRequiredTextDefaults), so an
   omission degrades to a sensible default instead of a 502.

Tests: draftRecoversRealPlanWhenJsonModeReturnsEmptyObject (generateJson
returns "{}", generate returns the real 2-block plan -> flow has 2 blocks, not
the fallback); draftDefaultsMissingRequiredTextFieldInsteadOf502 (MCPAgentChat
with no goalDescription -> valid, goalDescription defaulted to the purpose).
Full assistant suite green; only the known Loop-container timing flake fails
under full-suite load (passes in isolation).
2026-08-03 11:05:58 +02:00
Lucio Lelii 9cc110464d refactor: order MCP shared sessions with a dependency, drop the state_ready connection hack
MCP shared-session ordering was enforced with a fake data connection: the
consumer's prompt carried a magic ${{state_ready}} placeholder to synthesize
an input, and the producer's output was wired into it purely to force
execution order - the consumer never actually used that data (the real
shared state is the MCP session, accessed by name). This replaces that hack
with a Dependency, the mechanism meant exactly for ordering-without-data.
No engine change: the executor already gates readiness on dependencies and
the validator already counts them for reachability; MCPAgent uses the
default activity() capabilities so it can be a dependency source/target.

- FlowAssistantService: after assembly, generate ordering dependencies
  deterministically for all three cases:
    - top-level chain: Dependency(producer, consumer)
    - within-container chain: Dependency(producer, consumer) in the subflow
    - cross-boundary (top-level producer -> consumer inside a container):
      Dependency(producer, container) - the container runs after the producer
      and inherits its session.
  Existing dependencies are preserved on REFINE (preserveCurrentDependencies).
  Removed completeRequiredSequentialConnections, validateSharedMemorySemantics
  and their now-unused helpers (isReachable, summarizeMcpSessionBlocks); the
  ordering no longer flows through connections.
- Prompt: removed the ${{state_ready}} instructions; the model is told the
  backend orders the consumer after the producer via a dependency and must
  not add a placeholder or connection for it.
- Tests: updated all MCP mocks to stop emitting state_ready (prompt and
  connections) and assert the dependency instead (top-level, within-container,
  cross-boundary). Zero state_ready references remain.

Full suite green (431).
2026-07-26 11:37:42 +02:00
Lucio Lelii 68a86564fa feat: teach the assistant that EndBlock is a terminal label, not the flow's output
The assistant tended to close every path with an EndBlock (e.g. the
Jensen flows), which hides the actual result: an EndBlock consumes its
input and produces no output - the value is kept only as an outcome
payload, not as a normal flow output. The flow's readable result is
whatever block outputs are left open (unconnected), per
ExecutionContext.completeStep (which only adds an output to the result
map when !output.isConnected()).

- BLOCK-CHOICE GUIDE: EndBlock is a labelled terminal marker used only to
  record which terminal state a branched path reached (HIRED vs REJECTED)
  or to close a dead branch - never for the main result, and not needed
  to "finish" a linear flow (EndBlock is optional; nothing requires it).
- CONNECTIONS rules: leaving the result-producing block's output open is
  intentional and correct - that open output is the flow's readable
  result; don't wire it into an EndBlock just to terminate it.

Prompt-only change; full assistant suite still green (39).
2026-07-26 00:45:08 +02:00
Lucio Lelii 5d5e94f315 feat: FIX-path container diffing, cross-boundary MCP sessions, multi-input loop bodies
The three optional assistant refinements.

FIX-path incremental container diffing:
- validateAndNormalizePlan now diffs a container's inner blocks against
  the existing subflow in FIX mode too, not only REFINE - so a container-
  internal fix reconfigures only the inner block(s) the model marks,
  reusing the rest. Container-internal validation errors are remapped to
  the container id (the inner block id is lost), so if a FIX round would
  keep every inner block unchanged (the model gave no guidance) it falls
  back to full regeneration, guaranteeing the fix actually happens.

Cross-boundary MCP shared sessions (top-level producer -> inner consumer):
- Verified the runtime supports it: createAndStartSubflowChild (Iterator/
  Loop) and GenericContainerExecutor thread the parent's execution-variable
  descriptors (the MCP session registry) into a subflow and propagate them
  back, so a top-level producer's session is available to a consumer inside
  a container that runs after it.
- FlowExecutionValidator.collectErrors is now external-session aware:
  collectErrors(subFlow, externalSessions) threads the set of sessions
  produced by a top-level block reachable before the container, and a
  consumer is satisfied by such an external producer. This replaces the
  separate per-scope subflow pass for main subflows (which validated them
  in isolation and would reject a valid cross-boundary chain); loop guard
  subflows keep a dedicated external-aware pass since the recursion doesn't
  cover them. Bounded to top-level producers - inner-producer cross-boundary
  chains stay in-scope-only.
- Prompt: relaxed the "never split across a boundary" rule to permit a
  top-level producer wired to run before the container holding the consumer.

Multi-open-input LoopContainer bodies:
- AssistantContainerPlan gains feedbackInput; the LoopContainer config now
  passes it through, so a body exposing several open inputs resolves the
  otherwise-ambiguous feedback-target inference. Prompt updated.

Tests: fixReusesUnchangedInnerBlocksWhenTheModelMarksThem,
draftWiresMcpChainFromTopLevelProducerToConsumerInsideAContainer,
draftAuthorsLoopContainerWithMultiInputBodyAndExplicitFeedbackInput. Full
suite green (431). Roadmap: optional refinements marked done - only bias
annotation authoring (a separate product track) remains.
2026-07-25 20:11:51 +02:00
Lucio Lelii b0c9ef06f9 feat: teach the flow assistant to author LoopContainer via deterministic guard scaffolding
The assistant could author Generic and Iterator containers but not
LoopContainer ("repeat the subflow until a condition is met"). The
guard subflow has a rigid structural contract - it must expose a
non-multiple boolean output named "guard" and a non-multiple text
output named "feedback", produced through SwitchBlocks, reading the
body's latest result via the special ${{outputs.<name>}} reference.
Having the LLM author that freely would be fragile, so the assistant
supplies only the semantic intent and the backend builds the structure.

- AssistantContainerPlan gains guardCondition (natural-language stop/
  continue rule, required for a LoopContainer) and maxIterations;
  validateAndNormalizePlan accepts "LoopContainer" and requires a
  guardCondition when the body is (re)built.
- assembleContainer builds the body subflow like any container, then
  buildLoopGuardSubFlow deterministically assembles the guard subflow:
  a Guard Evaluator LLM (reads ${{outputs.<bodyOutput>}}, answers
  true=continue / false=stop per the guardCondition) -> an expose-guard
  SwitchBlock (boolean guard); a Feedback Builder LLM -> an
  expose-feedback SwitchBlock (text feedback) - mirroring the proven
  guard structure in ExecutionTest. feedbackInput is left for the config
  to infer (the body must expose one open input); the Feedback Builder
  deliberately doesn't reference ${{outputs.x}} to avoid a two-block
  exposed-input collision.
- Prompt: when to prefer LoopContainer, guardCondition/maxIterations,
  and that the body must expose exactly one open input (receives
  feedback each iteration). LoopContainer added to the plan schema.

Regression test draftAuthorsLoopContainerWithDeterministicGuardScaffold:
a one-body-block LoopContainer round-trips to a valid flow whose guard
subflow has the four scaffold blocks; passing validation proves the
scaffold satisfies the LOOP_GUARD/LOOP_BODY contracts. Full suite green
(428). Roadmap: item 2 marked done - all three container types
(Generic/Iterator/Loop) are now authorable.
2026-07-25 19:52:06 +02:00
Lucio Lelii 6142f3fc24 feat: array-valued globals + container-inner globals + within-container MCP sessions
Two related container-scope assistant-correctness gaps (roadmap items 5 and 6).

Item 5 - global input multiplicity and container-inner globals:
- The [] array marker now works symmetrically on globals end to end.
  ExecutionTemplateResolver emits a ${{global.name[]}} substitution key
  alongside ${{global.name}} (as it already did for normal inputs), so an
  array-marked global renders (a list value is newline-joined).
- FlowExecutionValidator strips a trailing [] from an extracted global
  reference name so ${{global.cvs[]}} matches the declared global "cvs".
- collectGlobalInputs scans top-level blocks AND every container's inner
  subflow blocks; a []-marked reference declares the global multiple=true.
  A global referenced only inside a container is now declared both at the
  top level (so the flow collects its value) and in the container subflow's
  own globalInputs (collectGlobalInputsFromBlocks) - the subflow is
  validated as its own scope and, at runtime, the container executor feeds
  it the parent's globals by name. This also fixes a pre-existing gap where
  a global referenced only inside a container failed GLOBAL_INPUT_NOT_DECLARED.
- Prompt: "never write ${{global.name[]}}" replaced with "append [] when the
  global holds a list".

Item 6 - within-container MCP shared sessions:
- Verified via GenericContainerExecutor that a container subflow inherits the
  parent's execution-variable descriptors (which hold the MCP session
  registry) and propagates them back, so an MCP chain inside one container
  genuinely shares state at runtime.
- assembleContainer computes isSharedMemoryContext for the container's inner
  plan and threads it into inner buildBlock / completeRequiredSequential-
  Connections / validateSharedMemorySemantics - the same normalization the
  top level uses, scoped to the subflow (was hardcoded false).
- FlowExecutionValidator.validateSubFlowSharedExecutionVariableOrdering
  validates each container subflow's shared sessions as its own scope, so a
  broken inner chain is caught at validation time.
- Prompt: a shared-session MCP chain must stay within one scope (all
  top-level or all inside the same container), never split across a boundary.

Still out of scope (documented): MCP chains split across a container boundary.

Tests: resolvesGlobalVariableWithArrayMarker (resolver);
draftDeclaresArrayGlobalInputFromArrayMarker,
draftDeclaresGlobalReferencedOnlyInsideAContainerInnerBlock,
draftWiresMcpSharedSessionChainInsideAContainer (assistant). Full suite green
(427). Roadmap doc: items 5 and 6 marked done.
2026-07-25 19:23:53 +02:00
Lucio Lelii 065256fb6c feat: incremental container diffing on REFINE (reuse unchanged inner blocks)
A container was previously either KEEP (reused whole) or ADD/UPDATE
(its entire inner block set regenerated from scratch) - the same waste
and churn risk targeted repair already removed for top-level blocks,
one level down. A REFINE like "add a step inside the review container"
re-authored every existing inner block.

- Extracted reusable cores from the top-level machinery so inner
  assembly shares identical rules: resolveExistingBlock(blockPlan,
  existingBlocks, ...) and preserveConnections(sourceConnections, ...).
- validateAndNormalizePlan resolves the existing container once, then in
  REFINE + UPDATE diffs inner blocks against the existing container's
  subflow (normalizeContainerInnerBlocksAgainstExisting /
  normalizeInnerBlockOperation): unchanged -> KEEP, changed -> UPDATE,
  new -> ADD, dropped -> REMOVE.
- assembleContainer takes the existing container: KEEP inner blocks are
  reused verbatim (no BLOCK_CONFIG call) and existing inner connections
  between surviving blocks are preserved and merged with generated ones.
- Prompt: the inner "operation" field is now documented
  (KEEP/ADD/UPDATE/REMOVE when a container is UPDATEd; omit for
  ADD/DRAFT), replacing the old "do not set operation on inner blocks".

Deliberately scoped to REFINE: FIX and ADD still fully regenerate a
container's inner blocks. FIX on purpose - a container-internal error is
most reliably fixed by rebuilding, and inner-block errors don't reliably
carry the inner block's id for targeted attribution, so incremental FIX
risked "KEEP everything and fix nothing". Left as a possible future
refinement (item 7 in the roadmap).

Along the way normalizeContainerOperation was replaced by
resolveContainerOperation (takes the already-resolved existing
container) so the same match is reused for inner diffing.

Regression test refineReusesUnchangedInnerBlocksWhenUpdatingAContainer:
REFINE a 2-block container to add a third; the mock throws if a KEEP
inner block is ever reconfigured, and the two kept blocks come back by
id with unchanged config and the pre-existing inner connection
preserved. Full suite green apart from the known Loop-container timing
flake (passes in isolation). Roadmap doc: item 3 marked done.
2026-07-25 18:56:30 +02:00
Lucio Lelii 85dcc11039 feat: teach the flow assistant to author IteratorContainer groupings
The assistant could group blocks into a GenericContainer but could not
express "run this subflow once per element of a list" - the per-item
iteration pattern (score each CV, process each document) that had to be
built by hand earlier this session.

- AssistantContainerPlan gains an `iterationInput` field, and
  validateAndNormalizePlan now accepts containerType "IteratorContainer"
  alongside "GenericContainer" (iterationInput is kept only for the
  former, dropped for the latter).
- assembleGenericContainer is generalized into assembleContainer: it
  builds the inner subflow exactly as before, then branches on
  containerType to construct either a GenericContainerConfiguration
  (GenericContainerFactory) or an IteratorContainerConfiguration with
  the iterationInput (IteratorContainerFactory). The factory infers
  iterationInput when the subflow has exactly one open input, and
  IteratorContainerInterfaceResolver auto-promotes the iterated input
  and every exposed output to multiple on the external interface - the
  assistant never declares the array-ness itself.
- The connection-wiring machinery needed no changes: a container is
  already a FlowNode endpoint, so an IteratorContainer slots into the
  same path.
- Prompt updates: the PLAN schema shows an IteratorContainer example
  with iterationInput; rules teach when to prefer it over N duplicated
  blocks, and that the inner block must use a single-value placeholder
  ${{cv}} (never ${{cv[]}}) named by iterationInput. CONNECTIONS rules
  note the iteration input takes the whole array and outputs are arrays.
- A wrong iterationInput is caught by IteratorContainerConfiguration's
  bean-validation on the FlowExecutionValidator pass (wired in earlier
  this session), triggering a repair round rather than shipping broken.

Bug found and fixed along the way: validateAndNormalizePlan treated
"no top-level blocks" as an empty plan and fell back to a single
synthetic LLMBlock, silently discarding every container when a plan put
all its work inside containers. The empty-check now requires both blocks
and containers to be empty before falling back, and the block loop is
null-safe for container-only plans.

Regression test draftAuthorsIteratorContainerForPerItemWork: a
container-only plan round-trips to a valid flow whose IteratorContainer
exposes cv/response as multiple while the inner cv input stays
non-multiple. Full suite green (422). Roadmap doc updated: items 1 and 4
marked done, item 3 (incremental container diffing) is now the top open
gap.
2026-07-25 17:33:34 +02:00
Lucio Lelii ab1def92f1 feat: allow targeted repair on flows that contain containers
isTargetedBlockRepairEligible previously bailed out to a full replan
whenever currentFlow had any containers at all, because
buildReusedPlanForTargetedRepair only reconstructed the top-level block
list - running it on a flow with containers would have silently
dropped them from the rebuilt plan.

buildReusedPlanForTargetedRepair now also reconstructs an
AssistantContainerPlan (operation KEEP, empty inner blocks) for every
container in currentFlow.flow().getContainers(), so containers survive
the targeted-repair path unchanged instead of disappearing. The
eligibility check still requires every error to be scoped to an
existing top-level block - an error scoped to a container (or anything
else) still falls back to a full replan, since a KEEP-only container
plan can't fix a container-internal problem.

Added a regression test: a flow with a container present has a
top-level block's overlong EndBlock.outcomeLabel fixed via a single
targeted BLOCK_CONFIG call (PLAN/CONNECTIONS both skipped), and the
container comes back byte-for-byte identical (same id, same inner
subflow).

Updated docs/assistant-completion-roadmap-2026-07-25.md to mark this
item done and note what's still open (container-internal errors are
still not targeted-fixable - that's the separate, larger "incremental
container diffing" item).
2026-07-25 16:59:15 +02:00
Lucio Lelii e482716068 feat: teach the flow assistant to author GenericContainer groupings
The assistant was structurally blind to containers: the plan/connections
JSON schema was flat, and every connection-resolution helper was typed
to Block<?>, so a container could never be a plan entry or a connection
endpoint. This wires up a first, deliberately scoped version: only
GenericContainer (plain grouping, no iteration/loop semantics), one
level deep (nesting stays unsupported), inner blocks always fully
(re)generated rather than incrementally diffed across FIX/REFINE
rounds.

- AssistantFlowPlan gains a "containers" array alongside "blocks";
  each entry (AssistantContainerPlan) carries its own nested "blocks"
  list using the exact same shape as top-level blocks.
- validateAndNormalizePlan now normalizes containers too: id
  uniqueness (shared namespace with block ids), containerType must be
  exactly "GenericContainer", inner blocks non-empty with unique ids
  and no nested container types, KEEP/UPDATE/REMOVE resolved against
  currentFlow.getContainers() the same way blocks resolve against
  currentFlow.getBlocks().
- The connection-resolution machinery (toValidConnections, toConnection,
  resolveConnectionBlock, inferBlockByIo, resolveConnectionOutputName/
  InputName, registerNodeAlias) is retyped from Block<?> to the shared
  FlowNode interface (Block and Container both implement it, exposing
  id/name/inputs/outputs) - a mechanical, behavior-preserving change
  that lets containers slot into the exact same connection-wiring path
  blocks already used, with no new logic needed there.
- assembleGenericContainer assembles a container's inner blocks (their
  own BLOCK_CONFIG calls) and inner connections (a CONNECTIONS call
  scoped to just that container, reusing buildConnectionsPrompt with a
  synthetic inner AssistantFlowPlan), then hands the resulting subflow
  to GenericContainerFactory, which auto-derives the container's
  exposed inputs/outputs from whatever inner handles are left
  unconnected - the assistant never has to declare them itself.
- Prompt updates: the PLAN schema/rules explain the "containers" array,
  when grouping is actually useful, and that only GenericContainer is
  currently authorable; the CONNECTIONS rules clarify that containers
  are valid endpoints via their own id and auto-derived I/O, and that
  inner blocks are never reachable from outside their container.
- isTargetedBlockRepairEligible now bails out whenever currentFlow has
  any containers, since the targeted-repair plan reconstruction only
  rebuilds the block list and would otherwise silently drop them.

Added a regression test covering the full path: a plan with a
top-level block plus a 2-block container, verifying the container's
subflow, its auto-derived exposed input/output names, and the single
top-level connection into it.

Deliberately out of scope for this pass: IteratorContainer/LoopContainer
(need iteration/guard semantics the assistant doesn't author),
incremental inner-block diffing across FIX/REFINE rounds, and nested
containers (still architecturally unsupported).
2026-07-25 14:43:03 +02:00
Lucio Lelii 94ed6b245d feat: targeted per-error repair instead of full replan on FIX rounds
Every FIX round previously re-ran the whole pipeline - PLAN, every
non-KEEP block's BLOCK_CONFIG, and CONNECTIONS - even when the only
problems were plain bean-validation errors scoped to one existing
block's own configuration (e.g. an overlong EndBlock outcomeLabel, or
the placeholder-naming/BranchRejoin-size violations added earlier this
session). That's wasted LLM calls and unnecessary churn risk on parts
of the flow that were already correct.

FlowAssistantService.assembleFlow now checks, per FIX round, whether
every reported error is cleanly scoped to an existing block and none
of them is in PLAN_SENSITIVE_ERROR_CODES (anything that could imply the
flow's shape needs to change - connections, dependencies, block/
container I/O, lanes, BranchRejoin fan-in, shared-session wiring, ...).
When that holds:
- PLAN is skipped entirely; the plan is rebuilt locally from currentFlow's
  existing blocks (UPDATE for the flagged ones, KEEP for the rest) with
  no LLM call.
- CONNECTIONS is skipped entirely; existing connections are carried
  forward unchanged.
Only BLOCK_CONFIG still runs, and only for the flagged blocks - the
same targeting that already happened via KEEP/UPDATE, now extended to
skip the two other phases too.

This is deliberately conservative: any error outside the safe set, or
any error not attributable to one existing block, falls back to the
existing full-repair behavior unchanged. If a "safe" fix unexpectedly
changes a block's derived I/O anyway (e.g. a factory that derives ports
from placeholder text), the next round's validation will surface a
sensitive code and force a normal full repair - never worse than
before, just self-correcting one round later.

Added a regression test that only mocks BLOCK_CONFIG (not PLAN/
CONNECTIONS) for a block-scoped-only error, so it fails loudly if the
skip logic regresses.
2026-07-25 14:24:34 +02:00
Lucio Lelii 8299c50c65 feat: auto-declare global inputs referenced by the assistant's blocks
The assistant could already be told (informally) about ${{global.x}}
via FlowExecutionValidator's GLOBAL_INPUT_NOT_DECLARED check, but had
no way to satisfy it: nothing in assembleFlow ever added an entry to
FlowData.globalInputs, so any block referencing a global would always
fail validation. This is exactly why "test cv ranking" needed a manual
fix earlier - the assistant has no way to reach for a real global input
today and instead duplicates the same input across blocks.

- FlowAssistantService.assembleFlow now scans every assembled block's
  specificConfiguration (serialized, then regex-scanned with the same
  ${{global.x}} pattern FlowExecutionValidator uses) and auto-declares
  a FlowData.globalInputs entry for each distinct name referenced,
  defaulting to TEXT/non-multiple. Existing declarations from
  currentFlow (REFINE/FIX, or a human-edited multiplicity) are
  preserved rather than overwritten.
- FlowAssistantPromptService: added rules + a worked example teaching
  when to use ${{global.name}} (same value needed unchanged by 2+
  blocks) versus a normal placeholder (data produced by another block,
  wired via CONNECTIONS). Noted the [] array marker does not apply to
  globals (ExecutionTemplateResolver only substitutes the bare
  ${{global.x}} key), and that a global reference has no input port to
  connect to in the CONNECTIONS step.
- Added a regression test where two blocks share ${{global.cvs}}: the
  flow ends up valid, with exactly one auto-declared global input and
  only the genuinely inter-block connection wired.
2026-07-25 14:06:52 +02:00
Lucio Lelii d08fb4af4d test: evict finished executions from the in-memory cache more aggressively
ExecutionsService's ExecutionObject cache only calls shutdown() (which
tears down that execution's dedicated Executors.newFixedThreadPool) on
eviction, gated by app.executions.cache.max-size (default 1000) and
app.executions.cache.final-ttl-ms (default 30 minutes). The Spring test
context - and its singleton ExecutionsService - is reused across the
whole suite, and a full run finishes in well under 30 minutes with far
fewer than 1000 executions, so none of that cleanup ever fires: every
execution's thread pool stays alive simultaneously for the rest of the
run. That's a plausible contributor to the loop-container tests
occasionally missing their polling deadline under full-suite load
(they pass reliably in isolation, where far fewer executions
accumulate). Lowering both knobs for tests only keeps the live
thread-pool count small throughout a run; eviction only ever targets
already-finished executions (see evictIfNeeded/evictFinalStateByTtl),
so this cannot disrupt an in-flight test.
2026-07-25 14:06:38 +02:00
Lucio Lelii 7cabf6aa08 feat: validate assistant flows with FlowExecutionValidator, raise repair budget to 2
The assistant's validate step only ran jakarta bean-validation on
FlowCreateRequest, which cascades into FlowDataValidator (the
@ValidFlowStructure constraint on FlowData) but never into each
block's own specificConfiguration - Block.specificConfiguration has
no @Valid annotation. So per-block constraints (BranchRejoinBlockConfiguration's
2-10 input bound, LLMBlockConfiguration's placeholder-name regex, MCP
shared-session producer/consumer wiring, container subflow rules via
the execution-level checks, ...) were never actually enforced: the
assistant could emit a structurally-broken block and still report
valid=true.

- FlowAssistantService.validate() now also runs
  FlowExecutionValidator.collectErrors(), merging its errors with the
  existing bean-validation errors so the repair loop sees (and can fix)
  what used to slip through silently.
- Default max repair attempts raised from 1 to 2 (still capped at the
  existing max of 3), since a single fix round often isn't enough once
  more error classes are actually being caught.
- Fixed draftDoesNotFailWhenSharedMemoryFlagsAreIncomplete: its flow was
  already broken (two MCP consumer-only blocks referencing a shared
  session neither produces) and is now correctly reported invalid
  instead of a false-positive valid=true.
- Added two regression tests locking in the new coverage and the new
  default repair budget.
2026-07-25 12:31:07 +02:00
Lucio Lelii 9556779241 feat: teach the flow assistant the new placeholder/BranchRejoin rules
The assistant's block catalog and prompts predated this session's new
conventions and had a structural gap: BranchRejoinBlock's "inputs" field
was blanket-excluded as if it were a runtime-derived IO list, so the
assistant could never declare a custom branch fan-in. Also x-ui-tip
(the field documenting the ${{name[]}} array marker) was dropped, and
the hardcoded BLOCK-CHOICE GUIDE only listed ~8 of the ~14 block types.

- BlockCatalogService: surface x-ui-tip on prompt field descriptors;
  stop excluding "inputs"/"outputs" when they're marked x-ui-structural
  (BranchRejoinBlockConfiguration's declared branch list) rather than
  a derived IO list.
- FlowAssistantPromptService: generate the BLOCK-CHOICE GUIDE from the
  live catalog instead of a stale hardcoded list; add explicit rules
  for placeholder naming, the [] array marker, and BranchRejoin's
  inputs contract; add small few-shot examples to the plan/config/
  connections prompts.

Deliberately not teaching ${{global.name}}/global inputs yet - the
assistant's assembly pipeline has no code path to declare a flow-level
global input, so advertising the syntax would produce broken flows.
That's implementation-level work, kept separate.
2026-07-25 09:42:32 +02:00
Lucio Lelii a5d94a2e55 fix: tolerant parsing of the simulator's HumanDecision choice
Simulated HumanDecision required the LLM's "CHOICE:" line to be exactly a
configured option name. But the simulate prompt lists options as
"- name: Label", so models - including strong instruct ones (qwen2.5:14b,
etc.), not just gemma:7b - routinely echo the whole "name: Label" line, e.g.
"red: Red light" or "existing: Existing position". That was rejected with
HUMAN_DECISION_INVALID_CHOICE, breaking simulated runs of any flow with
human decisions.

resolveSimulatedChoice() now accepts, in order: an exact option name, the
token before the first ':' (the echoed "name: Label" case), or the option
label. Only ':' is treated as a separator - never '-' - so hyphenated option
names like "not-red"/"assessment-not-required" are never truncated. Unknown
choices still return null and raise the same error.

Added a direct unit test (HumanDecisionSimulatedChoiceTest) covering exact
name, echoed name:label, hyphenated names, label text and rejection. Also
narrowed the flat-flow Jensen smoke test to exclude the new container-grouped
variant (different top-level shape).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 22:04:15 +02:00
Lucio Lelii 4f0f52c4f1 fix: container acting as a branch marks unselected exposed outputs NOT_SELECTED
A GenericContainer whose subflow takes only one of several internal branches
left the other branches' exposed outputs unproduced. The completion path
(reconcileGenericSubflow / GenericContainerExecutor.completedResult) returned
NodeExecutionResult.completed(valueMap) with an empty notSelectedOutputs, so
ExecutionContext.completeStep marked every unproduced exposed output
UNAVAILABLE. Downstream that reads as an error signal, not a skip: a
BranchRejoin (which self-skips only when ALL inputs are NOT_SELECTED) instead
saw UNAVAILABLE, went READY and threw BRANCH_REJOIN_INPUT_UNAVAILABLE, failing
the whole execution. A plain routing block never hits this because it reports
its not-taken outputs via notSelectedOutputs.

Both container completion paths now return NodeExecutionResult.routed(values,
allExposedOutputNames): a successful subflow that didn't take an internal
branch surfaces that branch's exposed output as NOT_SELECTED, so downstream
steps and BranchRejoins skip cleanly - letting a container act as a real
branching node (early-stop outcomes exposed as outputs routed to top-level
EndBlocks). Containers with a single always-produced output are unaffected
(that output is in the value map, notSelected is empty).

Found while building the containerized Jensen recruitment flow, where phase
containers expose early-stop outputs; added a focused regression test
(genericContainerUnselectedExposedOutputsSkipDownstreamCleanlyInsteadOfErroring)
and the "Jensen Recruitment Process - Full Revised (Containerized)" seed flow
that groups the 12 swimlanes into 7 phase containers (13 top-level nodes
instead of 39).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 21:41:52 +02:00
Lucio Lelii 1c94cd48d6 feat: harden DelimitedParserBlock (JSON fix, error codes, array output, tests)
Executed the full improvement plan for this block, found genuinely
under-baked (zero real usage, schema-only test coverage):

- Fixed a real bug: a JSON-typed output returned the raw trimmed string,
  but Input.validateValue() requires JSON-typed values to be instanceof
  JsonNode - configuring a JSON output was silently broken end to end.
  coerce() now parses the segment with ObjectMapper.readTree(), with a
  clear error on invalid JSON instead of a confusing downstream rejection.
- Replaced raw IllegalArgumentException with NodeExecutionException and
  stable error codes (DELIMITED_PARSER_INPUT_MISSING,
  _SEGMENT_COUNT_MISMATCH, _INVALID_BOOLEAN, _INVALID_FILE,
  _INVALID_JSON), matching the convention already used by
  BranchRejoinExecutor.
- Added an array-output mode: a DelimitedParserOutput (new record, kept
  separate from the shared SwitchCase used by SwitchBlock) marked
  multiple=true collects every split segment into a list with no
  fixed-count check, instead of requiring exactly N named outputs -
  complements the Iterator's "array in" with "array out" here.
- Boolean coercion now also accepts yes/no and 1/0, not just true/false.
- Added DelimitedParserExecutorTest (11 cases, was zero), two end-to-end
  ExecutionTest cases (fixed mode and array mode), and a bundled "test
  delimited parser" example flow (LLM -> DelimitedParserBlock -> EndBlock)
  since the block had never been used in any real flow before. Verified
  live against the running service with a real Ollama call.

No rename needed here - "DelimitedParserBlock" already describes what it
does, unlike ExclusiveMergeBlock.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 20:41:03 +02:00
Lucio Lelii 1fff0d3dc8 refactor: rename ExclusiveMergeBlock to BranchRejoinBlock
"Merge" was misleading: the block never combines values, it only lets
exactly one of several mutually-exclusive branches (an ExclusiveOR
reconvergence, e.g. after a HumanDecisionBlock split) pass through as the
single downstream value; two arriving values is an error, not something to
combine. BranchRejoinBlock names what it actually does. There is still no
block in the platform that genuinely merges/combines multiple parallel
values into one - this rename doesn't add one, it just stops "merge" from
implying it exists.

Renamed throughout: block/config/factory/executor/activation-policy classes
and their tests, the NodeVisualRole.MERGE -> BRANCH_REJOIN visual role and
NodeTypeCapabilities.merge() -> branchRejoin() factory method, the six
EXCLUSIVE_MERGE_* validation error codes -> BRANCH_REJOIN_*, the bundled
flows.json (11 uses across the Jensen flows), and the two docs that
reference it.

Added a legacy-id fallback in DynamicBlockConfigurationTypeResolver
(type id "ExclusiveMergeBlockConfiguration") and BlockTypes (typeName
"ExclusiveMergeBlock"), mirroring the existing ChatHumanInteraction ->
ChatInteraction precedent - found necessary the hard way: without it, any
flow already persisted under the old type id fails to deserialize and
brings the whole application down at startup, not just that one flow.
Locked this in with LegacyBlockTypeIdTest.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 14:17:28 +02:00
Lucio Lelii 98cf28bd9f fix: use a global input for the shared CVs array instead of duplicating it
The CVs array needs to be read by two nodes (rank-cvs and review-ranking)
that aren't wired to each other for it. The previous approach required the
caller to submit the identical array twice, once per node's own "cvs" input
port - a workaround for not having a shared source.

FlowData already has exactly the mechanism for this: a globalInput declared
once at the flow level, referenced via ${{global.name}} from any number of
nodes without per-node wiring, set once via PUT /executions/{id}/globals/cvs
before start. Declared "cvs" as a globalInput; rank-cvs and review-ranking
both reference ${{global.cvs}} and no longer declare a "cvs" Input port at
all.

Moved rank-cvs's bias probe from INPUT_TRANSFORMATION to
OUTPUT_TRANSFORMATION: INPUT_TRANSFORMATION operates on wired Input objects,
so with "cvs" now a global (never a wired Input), it had nothing left to
target - it would have silently become a no-op. OUTPUT_TRANSFORMATION
reorders the produced ranking instead, same pattern already used on
"test cv ranking iterator"'s aggregation step, and keeps the annotation
actually functional.

Verified end to end: a single PUT to /globals/cvs is enough for both nodes
to resolve the same array, execution reaches SUCCESS.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 13:23:25 +02:00
Lucio Lelii ca349209b6 feat: rank-cvs takes an array of CVs as a single initial input
Removes the three fixed collect-cv-1/2/3 HumanInteractionBlock nodes from
"test cv ranking": rank-cvs now declares one input, cvs, via ${{cvs[]}},
fed directly as an array before start (PUT .../input/cvs/texts) instead of
requiring a human to paste each CV into its own node. Scales to any number
of CVs instead of a fixed three.

review-ranking keeps seeing both the ranking and the original CVs (also via
${{cvs[]}} in the question) for the automation-bias check to stay
meaningful - since there's no longer a node whose output naturally fans out
to both consumers, the caller submits the same array to both rank-cvs.cvs
and review-ranking.cvs.

The order-bias probe on rank-cvs now targets the whole "cvs" array: since
INPUT_TRANSFORMATION already recurses per-element for list-valued inputs,
the instruction is applied uniformly to every CV rather than to one
specific candidate as before - a different (still meaningful) experiment,
not a broken one, but worth noting as a real tradeoff of moving from named
per-candidate inputs to a single array input.

Verified end to end against the running service: no human interaction is
needed to provide the CVs, ranking is produced correctly, and review-ranking
resolves both the ranking and the full CVs array.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 12:47:13 +02:00
Lucio Lelii f28fe0312e feat: add endpoint to list all iterations of a container step
A Loop/Iterator container step only ever exposes activeInnerExecutionId (the
current iteration's child), overwritten on every iteration. Past iterations'
child executions are never deleted from the DB, but there was no way to
enumerate them - only individually fetchable by id if a caller already had
it (e.g. from the event log).

Adds GET /executions/{id}/node/{stepId}/iterations, returning every child
execution created for that container step (one per iteration for
Loop/Iterator, one per run for GenericContainer), ordered by iteration
index. Backed by a new repository query
(findByParentExecutionIdAndParentStepIdOrderByParentIterationIndexAscCreationTimeAsc)
and ExecutionsService.getContainerIterationsByOwner(), which validates the
step exists and is actually a container before querying.

Verified end to end against the running service on the "test cv ranking
iterator" flow: the endpoint lists all 3 iterations in order with distinct
ids and SUCCESS status, returns 400 for a non-container step and 404 for an
unknown step id, and each listed iteration remains individually fetchable
via the existing GET /executions/{id}.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 10:58:22 +02:00
Lucio Lelii c6efa32c79 feat: array-input marker (${{name[]}}) for LLM/HumanInteraction/HumanDecision
Placeholders written as ${{name[]}} now derive a multiple:true (array) input
instead of the default single value, on the three block types whose inputs
are parsed from free text: LLMBlockConfiguration.prompt,
HumanInteractiveBlockConfiguration.actionDescription and
HumanDecisionBlockConfiguration.question. No marker -> unchanged single-value
behavior, so every existing flow keeps working as-is.

TemplateInputs now returns a Placeholder(name, multiple) record instead of a
bare name: the "[]" suffix is stripped from the captured group and the
remaining name goes through the same name-validation added in the previous
step. HumanDecisionBlockFactory's mandatory forward-through "input" port is
untouched - only the additive context inputs parsed from the question can be
arrays.

ExecutionTemplateResolver.resolve() now also emits the "${{name[]}}" spelling
as a substitution key alongside "${{name}}" for every value, so a template
written with the marker actually gets the value substituted at render time -
without this, the literal marker was left unresolved in the text.

Added hasConsistentPlaceholderMultiplicity() alongside the existing
hasValidPlaceholderNames(): a name referenced both with and without "[]" in
the same text is now a validation error instead of silently picking one side.

Rebuilt the "test cv ranking iterator" flow (IteratorContainer scoring each
CV independently, then an LLM/HumanDecision consuming the accumulated array
via ${{scores[]}}) using the marker instead of the abandoned explicit-field
design, and verified it end to end against the running service: all three
iterations complete, the aggregation LLM receives the array and produces a
real ranking, and the review step shows both the ranking and the full array
of per-candidate scores.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 10:11:35 +02:00
Lucio Lelii 6ed8fc748f feat: validate placeholder-derived input names on LLM/HumanInteraction/HumanDecision
Nothing today rejects a ${{...}} placeholder whose captured name contains
spaces, symbols, or brackets: the capture group is `.*?`, so it accepts
anything between ${{ and }}. That name flows straight into an IODescriptor
with no further checks.

Adds an @AssertTrue check (same convention as the existing
isOptionsUnique()/areSkillIdsUnique() checks) on LLMBlockConfiguration.prompt,
HumanInteractiveBlockConfiguration.actionDescription and
HumanDecisionBlockConfiguration.question, requiring every placeholder name to
start with a letter and contain only letters, digits, '-', '_' or '.'.

'.' stays allowed because it's already load-bearing: LoopContainer guard
prompts reference inputs.<name>/outputs.<name>
(ExecutionsService.buildGuardTemplateValues), and colliding container-exposed
names get qualified as nodeName.ioName
(ContainerFlowInterfaceResolver.qualifyWithNodeName). Confirmed via the full
suite - the first pass of this change broke 5 Loop container tests before
'.' was added back.

'[' and ']' stay rejected on purpose, reserving that syntax for a possible
future array-input marker on placeholder names.

TemplateInputs made public (was package-private) so the three
BlockConfiguration classes, which live in a different package, can call the
new hasValidPlaceholderNames() check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 09:55:24 +02:00
Lucio Lelii e11ef08537 fix: container completion no longer hangs in RUNNING on a downstream output error
Step.completeSuspendedContainer() (the async path that resumes a container
step once its inner subflow/iteration finishes) propagated container
outputs to downstream inputs without the same try/catch guard that
Step.run()'s synchronous path already had. When a downstream input rejected
the value (e.g. a multiplicity mismatch), the exception escaped through the
event-listener callback that drives it, leaving the container step stuck in
RUNNING forever with no recorded error.

Found while prototyping an IteratorContainer-based CV ranking flow: an
Iterator's accumulated (multiple) output wired into a plain LLM input
(always declared non-multiple) reproduced exactly this hang.

Wrapped the same body in try/catch, mirroring run()'s failure handling:
mark the step FAILED and notify the listener with a proper error message
instead of silently stalling.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 09:05:37 +02:00
Lucio Lelii ba4f05ba0a feat: add "test cv ranking" example flow with multi-CV input and ranking output
Three CVs are collected separately, ranked together by one LLM call (three
named inputs feeding a single ranking prompt), and reviewed by a human
decision maker who sees both the ranking and all three original CVs side by
side via the newly added multi-input support on HumanDecisionBlock.

Carries starter design-time bias annotations and probes on the ranking step
(SELECTION_BIAS, INPUT_TRANSFORMATION on one CV) and the review step
(AUTOMATION_BIAS, ROUTING_OVERRIDE) as a base for later bias-injection
experiments on ranking tasks specifically. Verified end to end against the
running service: full run reaches SUCCESS with a real ranking produced by
the LLM and accepted by the reviewer.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 08:43:47 +02:00
Lucio Lelii 9f4dc5303b feat: allow HumanInteractionBlock and HumanDecisionBlock to declare multiple inputs
Both block types previously exposed exactly one fixed input port, so a human
step could only ever be wired to a single upstream value even when the
decision or task genuinely needs more context to evaluate (e.g. seeing both
the original candidate profile and an automated screening result before
deciding).

Following the same convention LLMBlock already uses for its prompt, extra
named input ports are now derived from ${{name}} placeholders in the task's
free text: actionDescription for HumanInteractionBlock, question for
HumanDecisionBlock. No placeholders -> unchanged single "input" port, so
every existing flow keeps working as-is.

HumanDecisionBlock keeps its "input" port mandatory regardless of
placeholders, since HumanDecisionExecutor forwards that exact value as the
payload of whichever branch is chosen; extra placeholder-derived inputs are
additive context alongside it. HumanInteractionBlock has no such constraint,
so its inputs are fully derived when placeholders are present.

Extracted the placeholder-parsing regex (previously private to
LLMBlockFactory) into a shared TemplateInputs helper reused by all three
factories.

Wires the new capability into the bundled "test biased" flow: shortlist-decision
now also sees the candidate profile alongside the screening assessment, and
the container's reference-check step sees both the candidate profile and the
screening assessment. Verified end to end against the running service.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 08:34:11 +02:00
Lucio Lelii 7f92e1a5bf test: add interactive container to the bias test flow
Adds a GenericContainer (reference-check-container) between the shortlist
decision and interview prep steps, with an interactive human step feeding an
LLM assessment that carries its own bias annotation and behavioral probe.
Exercises subflow bias propagation (includeSubflow) and interactive nodes
inside containers together on the bundled bias test flow. Verified end to
end against the running service: normal run reaches SUCCESS, and a
bias-rerun with includeSubflow on the container correctly activates and
applies the inner probe.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 08:08:51 +02:00
Lucio Lelii a31e1692f6 feat: durable container coordinator and non-blocking Loop/Iterator subflows
Closes Tappa B and Tappa C of the interactive-containers plan.

Tappa B (recovery and lifecycle):
- Generalize the GenericContainer-only watcher into a type-dispatching
  coordinator (watchContainerSubflow/reconcileContainerSubflow) shared by
  Generic, Loop and Iterator.
- Add reconcileContainerSubflowsOnStartup (ApplicationReadyEvent): re-arms
  every persisted parent-child link at boot using only the entity's own
  parentExecutionId/parentStepId columns, so a child that already finished
  while the process was down is reconciled immediately, and a pending one
  gets its listener re-armed. No longer relies on any in-memory listener
  surviving a restart.
- cancelExecution now propagates in both directions: cancelling a parent
  cancels its active container child; cancelling a child fails the
  parent's container step (same outcome as an unexpected child error).
  removeExecution evicts cached children too.
- Propagate simulation mode to container children: Step gains
  executionSimulationEnabled (mirrors interactionSimulationDescriptor's
  propagation to every step, containers included); ContainerExecutionContext
  carries it plus the descriptor; startContainerChild starts the child via
  startSimulationExecution when applicable. ExecutionObject.hasSimulationAvailable
  now recognizes interactive nodes nested inside a container's subflow(s),
  which it previously ignored entirely since a container is never itself
  isUserInteractive().

Tappa C (Loop/Iterator non-blocking):
- LoopContainerExecutor and IteratorContainerExecutor become thin adapters
  delegating to ExecutionsService.startLoopSubflow/startIteratorSubflow.
  All advancement logic (creating/starting a child, checking whether it
  finished synchronously vs. suspending) now lives in ExecutionsService, so
  it is reusable both from the initial call and from reconciliation.
- Iterator advances iteration-by-iteration in a loop, chaining synchronous
  completions without suspending, and persists remainingValues/
  runtimeInputValues/accumulatedOutputs so it can resume exactly where it
  left off.
- Loop is a two-phase (MAIN/GUARD) state machine; the phase to resume into
  is derived from the completed child's own subflowRole, so it doesn't need
  separate persistence. Persists currentInputs/latestOutputs.
- FlowDataValidator and FlowExecutionValidator now accept interactive nodes
  in the subFlow of all three container types; LoopContainer's guardSubFlow
  remains rejected (guard interactivity is out of scope).

Existing container tests that assumed the old fully-synchronous model
(assert SUCCESS right after start()) are updated: since a child's steps
always run on their own thread pool, even a fully-automatic container now
transiently visits WAITING before the coordinator resolves it, so tests
must poll through WAITING too, not just RUNNING.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-23 20:28:06 +02:00
Lucio Lelii 0e5831289f feat: suspend generic containers for subflows 2026-07-23 19:25:21 +02:00
Lucio Lelii a64642923f feat: persist container subflow relationships 2026-07-23 19:11:41 +02:00
Lucio Lelii c2ae39f292 refactor: move container bias handling into subflows 2026-07-23 18:50:43 +02:00
Lucio Lelii 47dc3b3c3b feat: propagate bias context into container subflows (coarse model)
A bias variant rerun can now activate a container's inner subflow: with
includeSubflow=true on a container activation, every executable bias
annotation on the subflow nodes (and a LoopContainer's guardSubFlow) is
activated as a block, and the bias context is propagated into the
container's inner execution so those probes fire at runtime.

- BiasActivation gains includeSubflow (default false, backward compatible).
- BiasExecutionContext carries subflowActivatedContainerIds.
- BiasContainerPropagation builds the inner context by scanning the
  subflow for executable annotations.
- ContainerExecutor.execute receives the BiasExecutionContext; the three
  container executors create the inner execution with the propagated
  context (Loop applies it to subFlow and guardSubFlow, per iteration).
- validateBiasActivations validates includeSubflow (container-only, must
  have an executable probe) and no longer requires annotationIds when it
  is set; new error codes BIAS_ACTIVATION_ANNOTATIONS_REQUIRED,
  BIAS_SUBFLOW_ON_NON_CONTAINER, BIAS_SUBFLOW_NOT_EXECUTABLE.
- compareFullFlow treats subflow-activated containers as activated nodes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-23 18:17:37 +02:00
Lucio Lelii 1752b4ed15 feat: enable bias injection on human interaction nodes
Make a bias variant rerun affect human interaction nodes, both when
answered by a real human and when simulated:

- ExecutionStepView now exposes activeBiasProbes (annotationId,
  activationMode, instruction) for a step during a BIAS_VARIANT rerun,
  so the frontend can surface the injected bias to the interacting user.
  It is empty in normal/simulation runs, leaving the contract unchanged.
- PROMPT_DIRECTIVE is now supported on HumanInteraction/HumanDecision.
  It biases the simulator prompt and is a no-op when a real human
  answers (the directive is shown via activeBiasProbes instead).
- HumanInteractionExecutor.simulate and HumanDecisionExecutor.simulate
  now decorate their prompt via BiasRuntimeSupport.decoratePrompt, so an
  active prompt directive reaches the simulator.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-23 16:50:16 +02:00
Lucio Lelii 10420ab6b7 refactor: drop the unused EndMode enum from EndBlockConfiguration
EndMode only ever had one value (PATH_END) and wasn't read anywhere,
so it was pure dead configuration. Keep JSON backward compatibility by
ignoring a legacy "mode" field on deserialization instead of failing,
and drop it from the seed flows and docs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-23 14:55:35 +02:00
Lucio Lelii 2a349befa5 fix: let HumanDecision blocks support execution simulation
HumanDecisionExecutor explicitly disabled supportsSimulation(), so any
execution containing a HumanDecision node lost the aggregate
simulationAvailable flag entirely, hiding the simulate feature from
clients even though the /simulate endpoint itself was untouched.
simulate() now prompts the simulator LLM to pick one of the configured
options (and a rationale when required), validates the choice, and
returns the same output shape a real human interaction would produce.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-23 14:54:59 +02:00
Lucio Lelii a0ed4874d8 feat: add node capabilities and biased seed flow 2026-07-22 17:38:42 +02:00
Lucio Lelii 007bca14fb fix: add missing Flyway autoconfiguration and shorten an overlong flow description
pom.xml declared org.flywaydb:flyway-core directly, but Spring Boot 4 moved
Flyway's autoconfiguration into a separate spring-boot-flyway module: with
only flyway-core on the classpath, Flyway never ran (no migration, not even
flyway_schema_history), so ddl-auto=validate always failed against an empty
schema. Switch to org.springframework.boot:spring-boot-starter-flyway, which
pulls in the actual autoconfiguration.

The "Jensen Recruitment Process - Mitigated (Structured)" seed description
was 269 characters, over the default varchar(255) on flow_entity.description
(no @Column(length=...) override), so FlowImportComponent failed to insert
it while every other flow (all <= 255 chars) imported fine. Shortened to the
same content in 236 characters.

Verified end-to-end against a real Postgres container: fresh schema
bootstrap (ddl-auto=update once, then default validate), all 8 seed flows
import cleanly, full test suite unaffected (362 passing, 2 skipped).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-22 15:24:51 +02:00
Lucio Lelii 26ab50f06c fix: remove unused --enable-preview and clean up compiler warnings
No source in this codebase actually uses a preview language feature, but
--enable-preview combined with release=25 in pom.xml (compiler + surefire)
was tripping the IDE's language server on every file ("preview can be
enabled only at source level 26"), which read as pervasive warnings across
the whole project. Removed both flags; full suite still passes unchanged.

Also fixes a handful of genuine compiler warnings found via a manual
`javac -Xlint:all` pass: dead unused method in BlockCatalogService, two
uses of JsonNode.isTextual()/asText() (deprecated in tools.jackson 3) in
favor of isString()/asString(), and a Javadoc comment in
ExecutionsController placed after @GetMapping instead of before it (so it
was never attached to the method).

Left the remaining low-value warnings (missing serialVersionUID on a
handful of exception/resolver classes, this-escape in
ExecutionObject/ExecutionContext/Step) untouched per user's choice - purely
cosmetic and, for this-escape, riskier to touch without a concrete reason.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-22 14:21:09 +02:00
Lucio Lelii 28670d2436 test: close gaps against the control-flow doc's test checklist
Audits every bullet in "Strategia di test" (unit/integration/regression/
concurrency) against the actual suite and adds coverage for what was
missing: ExclusiveMergeActivationPolicy and ExclusiveMergeExecutor direct
unit tests, dynamic configuration validation, a second-hop NOT_SELECTED
propagation assertion, implicit-merge rejection (EXCLUSIVE_BRANCH_MERGE),
a lanes/laneId JSON round-trip, per-block-type bias capability checks, and
concurrency tests for the merge (repeated concurrent split+merge, cancel
under concurrency, concurrent snapshot reads).

Fixes a latent bug found while writing the configuration-validation tests:
ExclusiveMergeBlockConfiguration.areInputsUnique() and
HumanDecisionBlockConfiguration.areOptionsUnique() never fired because Bean
Validation only recognizes isXxx/getXxx as constrained getters, so the
duplicate-name check was silently unenforced everywhere, including flow
creation. Renamed to isInputsUnique()/isOptionsUnique().

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-22 12:35:24 +02:00
Lucio Lelii 2d6f12a019 feat: add structured Jensen recruitment flows (phase 6)
Adds "Jensen Recruitment Process - Full Revised (Structured)" and
"... Mitigated (Structured)" seed flows, ported node-per-node from the
original PlantUML activity diagrams using HumanDecisionBlock gateways,
explicit ExclusiveMergeBlock rejoins, EndBlock outcomes and FlowLane
swimlanes. Preserves the 5 CONFIRMED risk annotations and all 16
FH-J1-FH-J16 MITIGATED annotations on the corresponding structured nodes.
Does not replace the existing unstructured macro-flow seeds.

Human-interactive stages are modeled as HumanInteractionBlock rather than
GenericContainer/LLMBlock so both flows can be driven to a real SUCCESS
outcome without an Ollama model. Adds structural validation tests and a
runtime test suite driving all five terminal outcomes on both flows, a
snapshot-restore check during a pending HumanDecisionBlock, and a
baseline/bias-variant comparison showing a ROUTING_OVERRIDE changing both
the branch taken and the final outcome.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-22 12:08:58 +02:00
Lucio Lelii 1cf780966b feat: add swimlanes and JSON dossier IO type (phases 4-5)
Implements Phase 4 (FlowLane/laneId on FlowData, Block and Container with
dedicated validation) and Phase 5 (IOType.JSON, optional value schema on
IODescriptor, JSON conversion/validation in Input, template rendering) from
the control-flow engine implementation plan. Updates the plan document's
progress table and phase sections accordingly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-22 11:32:25 +02:00