Commit Graph

271 Commits

Author SHA1 Message Date
Lucio Lelii ddc792ff28 Let a provider declare it needs an endpoint, and accept one as a credential
Two additions to LLMProvider, both additive defaults so no existing
provider or test stub changes behaviour.

requiresEndpoint() names the one thing every provider until now has had
in common without anyone needing to say so: a base URL it already knows,
whether server-configured or a constant of the service it talks to. A
provider whose endpoint is not known until a credential names it -
coming next - is the first that needs to say otherwise.

ProviderCredential carries that endpoint alongside the secret value a
provider has always received. The three new generate/generateJson/chat
overloads that take one default to unwrapping .value() and calling the
String-authorization overload above them, so a provider that only
overrides the old ones - which today is every one of them, including
every anonymous test stub across the suite - keeps behaving exactly as it
did. Only a provider that overrides the new overloads directly gets to
read .endpoint() at all.

Adding an abstract method instead would have broken every one of those
stubs, since none of them implement anything beyond the three methods the
interface already requires.

One ambiguity fell out of this at the call site InternalOllamaLLMProvider
used to have: chat(model, messages, null, null) no longer resolves
unambiguously, since a bare null now fits both the String and the
ProviderCredential overload equally. Not visible in this diff - that call
site was rewritten away in the Ollama extraction - but worth naming since
it is the shape of thing this kind of overload addition can trigger
elsewhere too.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 12:27:19 +02:00
Lucio Lelii 68c3a1e4a5 Add a shared HTTP transport base for LLM providers
Every provider that speaks HTTP wrote its own version of the same handful
of things - a client to reuse, an error mapping, a blocking call with
logging - and each ended up with a different half of it. Gemini has a
retry policy and no logging at all; Ollama has logging but no retry; both
map only 5xx responses and drop the body of a 4xx, which is exactly the
detail that would say whether a model name is wrong, a key lacks access,
or an endpoint is misconfigured.

AbstractHttpLLMProvider consolidates that: a WebClient cache keyed by base
URL (needed once an endpoint can vary per call, which a user-supplied one
will), 4xx and 5xx both mapped through LLMProviderHttpException carrying
the response body, and failure logging that distinguishes an HTTP error
from a connection failure from a timeout.

It is deliberately transport only - no generate/chat/generateJson, no
request body, no parsing. Two hooks, timeout() and retryPolicy(), are
overridable rather than fixed, because Gemini authenticates through a
query parameter rather than a header and any future provider might too;
a base class that assumed otherwise would not be a base class Gemini could
actually sit on.

Not wired into any real provider yet - AbstractHttpLLMProviderTest proves
the plumbing generically, through a minimal test-only subclass and a real
JDK HttpServer, before anything depends on it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 12:27:00 +02:00
Lucio Lelii e7f3ee5474 Add the SPDX licence header to every source file
Mechanical: three comment lines above the package declaration of every
.java file under src, main and test alike, and nothing else. Its own
commit because it moves the blame line on 501 files and would otherwise
bury the licence change it belongs to.

The short SPDX form rather than the full GNU notice - machine-readable
under REUSE, sufficient to keep the licence notice intact, and it defers
the attribution term to LICENSE-ADDENDUM rather than repeating it five
hundred times.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 12:10:34 +02:00
Lucio Lelii 9b354b4286 Release under the AGPL with an attribution term
Same licence and the same additional term as the web repository: the two
halves are one product, and licensing them differently would leave the
question of what a derivative of the whole owes unanswerable.

The AGPL rather than the GPL because this service is meant to be hosted,
and section 13 is what obliges whoever hosts a modified copy to offer its
source to the people using it. The section 7(b) term in LICENSE-ADDENDUM
requires the attribution to be preserved, including in the Appropriate
Legal Notices a derivative displays.

The pom's licence, developer and scm blocks were the empty placeholders
Spring Initializr generates, so the published artifact declared no licence
at all - now they say what the LICENSE file says.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 12:10:26 +02:00
Lucio Lelii a13092dcea Log the whole completed operation, not just the field we read
Only result.result is modelled here, so anything else the bridge reports
is dropped at deserialization - which makes a bridge that returned its
answer under another name indistinguishable from one that produced
nothing. "No output generated" is exactly the case where that distinction
decides what to fix, and the bridge's schema is not in this codebase, so
the payload itself is the only way to tell.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 16:51:32 +02:00
Lucio Lelii 337b2939ac Stop reporting a client that went away as a server failure
The editor polls a running execution and drops the request when it
navigates, refreshes or supersedes it - routine, and more likely the
longer the response takes to write, which an execution view carrying a
large global input does. Each one was logged as "Unhandled request
failure" with a hundred-line stack trace, for something nobody can act on
and where the reply goes to a connection that is already gone.

Recognised through the cause chain, because Jackson wraps the broken pipe
several times over before it surfaces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 16:28:49 +02:00
Lucio Lelii 97f3ab0e66 Give a query with attachments the same step budget, within the documented ceiling
max_steps went only on the JSON query operation, so a query carrying an
attachment still ran on the bridge's default and could abort the same way.
The multipart operation takes it too.

The default drops from 200 to 100, which is the ceiling the bridge
documents: asking for more is at best ignored and at worst refused.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 15:10:56 +02:00
Lucio Lelii 4ce74e9863 Ask the bridge for more steps by the name it actually takes
The query kept aborting with "Recursion limit of 60 reached". That
message names LangGraph's own recursion_limit, which is what the bridge
sets internally - so that is the key that was sent, on the session and on
every query, and it changed nothing. The bridge's API takes the budget as
max_steps, on the query operation, which is worth more than any amount of
reasoning about its error text.

Sent there and nowhere else now: the budget belongs to a query, not to
the session that may run several.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 15:04:54 +02:00
Lucio Lelii 03db2dee5c Ask an upload row which source it is on, then show only that branch
The row offered both ways of giving a file at once and greyed out the one
not in use, which still leaves the reader working out which half is live.
It reads better as what it is: one choice, then the fields that choice
needs - an input name, or which global input holds the file.

Greying was all the schema could express, so this adds the annotation for
showing a field only in the state it belongs to. The distinction earns
the second annotation: a field that still tells the reader something
while unavailable should stay and grey, but the branch nobody picked is
not unavailable, it is irrelevant.

A row saved before the choice existed says which branch it is on by what
it carries, so it is stamped on read rather than left reading as the
default and demanding an input name it never had.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 14:46:01 +02:00
Lucio Lelii fed7141c06 Offer an upload row's two sources as the alternatives they are
Picking a global and nothing else still failed the block update: the kind
select posts "" when nobody chose from it, and an enum cannot be coerced
from that, so the request came back naming a field that is optional and
that the person had deliberately left alone. Blank now reads as unset.

The row also left both sources on offer at once, with nothing saying which
one wins. Filling either now greys out the other, and the input name is
required exactly when no global is chosen - which needed a way to say
"while this other field is empty", since a primitive boolean cannot tell
an unspecified present() from a false one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 14:37:39 +02:00
Lucio Lelii 665a05899d Ask an upload input only for what it alone can answer
Choosing a global input and nothing else broke the block update outright:
"multiple" was a primitive boolean, so a row posted before that box had
been touched failed to deserialize, and the editor showed a bad request
naming a field nobody had filled in. Absent now means single, which is
what a half-filled row means.

The other two fields were being asked for without earning it. A name is
the port's name, so an attachment taken from a global has none to give -
and declaring one grew a port that asked for the same document a second
time, once as a global and once as a step input. A kind is written before
any file exists, can be wrong by accident and can be wrong on purpose, so
what a file is now comes from its own bytes when the query is built; the
declaration only filters the picker, and only where there is one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 14:18:50 +02:00
Lucio Lelii 1d5d1eb78b Keep isFromGlobalInput out of the payload it is not part of
Jackson reads an is-prefixed no-arg boolean as a property, so the helper
added with the field put a "fromGlobalInput" into every block payload and
every persisted flow - a key the generated schema never declares, sitting
next to the one it is derived from. ModelParameters.isEmpty carries the
same guard for the same reason.

Also cover the two steps between this field and the code that reads it:
the generated schema has to carry it, and a configuration posted back by
the editor has to survive deserialization into a block whose ports reflect
the choice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 14:11:54 +02:00
Lucio Lelii f4693ac6ab Attach a global input's file, and name a file by its name in a prompt
Two halves of the same trap. An MCP agent handed ${{global.document}} got
this server's temp path - "/tmp/document1190411166785638981plan.pdf" -
and spent its turns trying to open a file it cannot reach, reporting it
could not access the plan. Meanwhile the only way to attach a file for
real was a port on the block, so a document needed by four agents had to
be uploaded four times.

An upload input can now name a global input to take its file from, chosen
in the editor from the flow's file-typed globals (the retriever pattern
sharedSessionRef already uses); named that way it grows no port, since a
global reaches a block by being named and a port for it would sit
unsatisfiable. And a file interpolated into a prompt now renders as its
name, which is also the name the bridge is told the attachment has - for
which the upload had to stop mangling it, so each one now lands in a temp
directory of its own under the name it arrived with.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 12:53:01 +02:00
Lucio Lelii 818f9c28e1 Stop re-deleting MCP sessions the bridge has already dropped
The bridge expires sessions on its own, so the idle sweep regularly asks
it to delete one that is already gone. That 404 was treated as a failure,
and the cost was not just the warning and stack trace every minute: the
removal from activeSessions sat after the call that threw, so the dead
session was never untracked and the sweep retried it forever - and each
attempt counted against the circuit breaker that guards real MCP calls,
where five of them open it.

A session the bridge no longer has is the outcome this method wants, so
404 now completes it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 12:34:42 +02:00
Lucio Lelii 4516105d74 Give a global input of a file kind somewhere to be uploaded to
The only endpoints for a global input took JSON, so an upload aimed at one
was refused before it reached any handler: "Content-Type
'multipart/form-data' is not supported". A flow whose global input is a
file - the plan document of the orchestrator flows, for one - could
therefore never be given its file at all.

Add the multipart pair the node inputs already had, named the same way
because it is the same operation on the other scope.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 12:31:43 +02:00
Lucio Lelii 01d739ee7d Stop a short input name from failing a file upload
The uploaded file's temp name used the input's own name as the prefix
handed to File.createTempFile, which rejects anything shorter than three
characters. Uploading to an input called "dc" therefore threw
IllegalArgumentException - past the IOException catch, so the client saw
only a 500 that reads as "Failed to upload file" in the UI, with nothing
naming the real cause. A filename carrying a path separator failed the
same way.

Sanitise and pad both parts: neither is the uploader's mistake to pay
for, and nothing downstream reads meaning out of the temp name beyond the
extension, which is preserved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 12:19:20 +02:00
Lucio Lelii cc820554ac Close shared MCP sessions when their execution is cancelled
ExecutionContext.cancel() cleared executionVariableDescriptors before
transitioning to CANCELLED, so ExecutionsService.cleanupManagedResourcesIfFinal
(triggered from the same state-change notification) always found an empty
map and closed nothing. The CLOSE_RESOURCE mechanism already worked
correctly on SUCCESS and ERROR - only cancel/stop silently leaked any
shared MCP bridge session still open at that point, until the bridge's
own idle timeout, eventually hitting its session cap ("Reached the
maximum limit of 5 sessions").

Defer clearing executionVariableDescriptors until after the state-change
notification runs, so the cleanup sees the still-registered CLOSE_RESOURCE
session descriptor and closes it before the map is scrubbed. Added a
regression test that reproduces the leak (confirmed red without this
change) and asserts closeSessionQuietly runs on cancel.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 11:39:01 +02:00
Lucio Lelii 44991f4a79 Raise the MCP bridge's agent-loop recursion limit for complex queries
A production run hit "Recursion limit of 60 reached without hitting a
stop condition" on initialize-persistent-orchestrator, which needs many
tool round-trips (inspect workspace, write/verify two files) in one
query - the bridge's own LangGraph agent loop aborted before finishing,
even though nothing on our side errored.

Send recursion_limit (configurable via app.mcp.bridge.recursion-limit,
default 200) on both the session-open request and each query operation,
same best-effort spirit as think: the bridge's request schema isn't in
this codebase, so this is sent on faith it's honoured somewhere - if it
isn't, nothing changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 10:55:54 +02:00
Lucio Lelii 9ad2394090 Keep a thinking model's reasoning trace out of block/query outputs
A reasoning model (e.g. qwen3) mixed its <think> preamble into the same
text this code treats as the final answer, since neither the direct Ollama
calls nor the MCP bridge session ever asked for reasoning to be kept
separate. That let a Conditional's SpEL condition (or any other consumer)
silently see reasoning prose instead of the expected value - in a
LoopContainer this meant looping through iterations without ever taking the
intended branch, with nothing logged to show why.

- Ask Ollama to think explicitly (think: true) on both the direct provider
  and the MCP bridge's llm_provider options, so reasoning is returned
  separately instead of folded into response/content - full reasoning
  quality kept, unlike think:false which would ask the model to reason
  less.
- Defensively strip a closed <think>/<thinking> block from an MCP query
  result, and fail loudly instead of returning garbage when the block is
  unterminated or leaves nothing behind - the bridge is an external service
  we don't control, so this is the fallback for whatever it sends anyway.
- Log every MCP query result's raw text to its own file
  (mcp-agent-responses.log), mirroring the existing assistant-responses
  log, since there was previously no way to see what a given model/bridge
  combination actually returns.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-12 13:27:12 +02:00
Lucio Lelii 3ccdeca637 Key coding-agent-mcp workspace by root execution id, not per-iteration id
A LoopContainer iteration spawns a fresh child ExecutionObject each time, so
context.executionId (used to key the coding-agent-mcp workspace subpath) was
never stable across iterations once MCPAgent nodes stopped sharing one MCP
session. Add context.rootExecutionId (the top-level execution an iteration
belongs to) and use it for the workspace subpath instead, so independently
sessioned MCPAgent nodes still land in the same workspace across iterations.

Also fixes a narrow, real race in JensenStructuredFlowsExecutionTest: a
step's in-memory status flip and its listener-triggered persistence happen
on the same background thread but aren't atomic, so evicting the execution
from cache and reloading it could occasionally observe a stale snapshot.
Retry the evict-and-reload instead of asserting on a single attempt.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 17:27:58 +02:00
Lucio Lelii 6f9a0d1929 Add MCP multipart file uploads 2026-09-11 14:50:16 +02:00
Lucio Lelii aa308fd70f Use asynchronous MCP query operations 2026-09-11 10:48:11 +02:00
Lucio Lelii 89126a2b9e Avoid retrying MCP session creation 2026-09-10 12:04:11 +02:00
Lucio Lelii 18f2ab2fa0 Let a secure retriever answer the questions the editor asks it
The editor asks every retriever-backed field whether its list is open,
and whether the field is required, on endpoints suffixed onto the
field's own URL. Only the unsecured retriever had them, so opening an
MCP agent's shared session picker asked /secure-retriever/.../open and
got a NoResourceFoundException stack in the log. The editor swallowed
the failure and fell back to a closed list, which is why nothing looked
wrong.

Both questions now exist on the secure side too, defaulting the same
way, so the answer comes from the retriever rather than from a 404.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 11:35:19 +02:00
Lucio Lelii 90c97f9aad Harden Maven downloads in Docker build 2026-09-10 11:34:39 +02:00
Lucio Lelii 760554c2d7 Fix shared MCP session ownership in subflows 2026-09-10 11:34:37 +02:00
Lucio Lelii a2743c6c66 Validate inherited subflow globals 2026-09-09 17:17:04 +02:00
Lucio Lelii ec009411e1 Fix shared MCP sessions in container subflows 2026-09-09 16:34:05 +02:00
Lucio Lelii 289bcfb1db Note how to run the service locally and build its image
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-09 12:24:51 +02:00
Lucio Lelii 44f722e102 Publish which of an enumerated field's values is the default
A field with a fixed set of values had no way to say which one it opens on, so the MCP server
dialog started with no source type chosen: nothing was selected, nothing said it had to be, and the
server could be saved that way. `@SchemaAllowedValues` now takes a `defaultValue`, emitted as the
schema's `default`, and both MCP configurations declare CATALOG - which is also the choice whose
own required fields the dialog can then enforce.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-09 12:23:42 +02:00
Lucio Lelii 8d6ddfecfb Replace Flyway schema setup with JPA mappings 2026-09-09 12:20:36 +02:00
Lucio Lelii e71ba9c290 Stop a second LLM assessment from erasing the one still running
Two assessments of the same report, started minutes apart, left only one:
observed in a real run as two REPORT_JUDGE jobs both COMPLETED (14:33-14:37
and 14:36-14:37) against a report that ended up holding a single judgement -
the one that finished last.

judgeReport was one @Transactional method that read the report, spent minutes
in one model call per compared pair, then wrote. The second assessment read
the report while the first was still calling models, saw no history, and
saved its own judgement as the only one there. Nothing warned, because from
each writer's side the write succeeded.

The model calls now happen outside any transaction - holding one open across
minutes also pins a connection for no reason - and the write moved to
BiasJudgementStore, a bean of its own so the transaction starts there and a
retry re-enters through the proxy rather than inside the transaction that just
failed. It re-reads the report as it is at that moment, puts the new verdicts
on those pairs and the summary on that history, and the report entity carries
a @Version so a writer working from a stale read is refused instead of
overwriting: the next attempt reads the assessment that landed meanwhile and
appends after it.

Optimistic rather than SELECT ... FOR UPDATE because the tests said so:
Hibernate renders PESSIMISTIC_WRITE as "for no key update", which H2 - what
the suite runs on - cannot parse. A fix that only holds on one dialect is not
one.

The regression test runs the two assessments on two threads, with a stub
provider that blocks inside the model call until both have reached it, and
asserts the report keeps both. Reinstating the old shape fails it.

This prevents further losses; it does not recover an assessment already lost.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 15:14:55 +02:00
Lucio Lelii 02b284d368 Keep every LLM assessment of a bias report, and fall back to text when JSON comes back empty
A report held one BiasJudgeSummary, so asking a second model overwrote the
first: there was no way to compare two opinions on the same comparison, and
a report survived only until someone reevaluated it. BiasImpactReport now
keeps a capped history (10) of assessments, newest first, each carrying its
own verdicts per compared pair rather than a single shared set - the point
of keeping several is being able to trust each one's own reasoning, not just
its headline. Recomputing an outdated report now carries the history forward
instead of discarding it.

schemaVersion had to become Integer rather than int in the same change:
Jackson reads an absent field as null, and a primitive rejected every report
persisted before today - a bug this history change would otherwise have
inherited silently, caught by a new compatibility test that reads an old
report's JSON.

Separately, the judge asked for its verdict in JSON mode only, and Ollama's
format=json makes some models - reasoning models especially - answer with an
empty body or a degenerate {}, since they have nowhere to put their thinking
under that flag. Every compared pair failed with "the model returned an
empty answer". The flow assistant has ripped this exact seam out before and
falls back to a text call; the judge now does the same, extracting the
verdict object from whatever prose it comes wrapped in.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 11:40:21 +02:00
Lucio Lelii 7d3f6b088c Resolve every configurable-as-input field through one rule
Three fields are declared @ConfigurableAsInput - the model of an LLM
descriptor and the two on the MCP blocks - and the editor offers the same
choices on all of them: leave it blank and feed the port, or write a template
such as ${{global.modelName}}. It was resolved three times and differently.
The MCP executors each carried an identical private copy that fell back to the
configured value raw, so a placeholder written there reached the MCP service
verbatim, and a global input could not decide the model of an MCP block at all.

The rule now lives in ConfigurableInputBinding: a value bound to the port
wins, otherwise the configured value is resolved as a template against the
same inputs and variables a prompt is. LLMDescriptorInputBinding keeps only
what is specific to it - rebuilding the record around the resolved model, and
resolving the model alone, since the provider decides the credential and the
sampling parameters and cannot arrive mid-execution.

The MCP model fields carry @AcceptsVariablePlaceholder to match, so the editor
declares what the runtime has always been asked to accept.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 10:52:52 +02:00
Lucio Lelii 8f9017891c Compare a bias variant subject by subject, and let an LLM assess it
A full-flow comparison measured one number: an edit distance between the
outputs of every activated node stringified into one map. On a flow that
evaluates one subject per iteration - the multi-CV case - that number was
dominated by the map's punctuation and could not say which subject moved,
by how much, or whether anything of substance had changed.

The comparison is now field by field, and list values are paired element by
element, so an iterated node reports per subject. Where both sides label a
number the same way, its movement is reported in the flow's own units
(Score 8 -> 5), which is the only figure here a reviewer can act on. Values
that only one side produced are marked rather than guessed at, and both the
element count and the edit distance are capped, the latter because it was
quadratic with no bound on exactly the long model outputs it runs on.

Iterator and loop containers are compared iteration by iteration by joining
the child executions on parentIterationIndex, which they already record.
Guard subflows are excluded: a loop creates one per main iteration, and
pairing a guard with a main run reported every loop as rewritten. This is
what turns "the accumulated list changed" into "iteration 3, on this inner
node".

None of that says whether the change is the intervention doing what its
probe described or the same model answering differently, so a report can now
be handed to a model: one call per aligned pair, answering on two separate
axes - how far the meaning moved, and whether the change carries the
intervention's fingerprint. Provider, model and sampling arrive as an
LLMDescriptor and are resolved the way the interaction simulator's are. The
level and the attribution shown are rolled up from the per-pair verdicts in
code; only the narrative comes from the model, so re-reading a report cannot
show a different headline than the pairs it is made of. A model that answers
with something unusable leaves an error on its pair and nothing else.

Finally, a rerun now carries the simulator of the run it repeats - the
descriptor only, never the enabled flag - and the report records what
answered the interactive steps on each side. Two runs answered by different
simulators differ for a reason the intervention had no part in, and until now
nothing said so.

Reports are persisted and served from storage, so they carry a schema version
and an older one is recomputed instead of serving its new sections empty.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 10:50:15 +02:00
Lucio Lelii d024f80239 Separate "takes a ${{...}} placeholder" from being drawn as a textarea
LongText.acceptVariableAsPlaceholder welded a fact about the value to a
rendering choice: only a field also drawn as a textarea could say it is
interpolated. That left LLMDescriptor.model unable to declare it - the
capability worked, since the executors resolve the field as a template,
but nothing in the schema said so and nothing in the editor showed it.

@AcceptsVariablePlaceholder is that fact on its own. LongText keeps its
flag as the shorthand for the many prompt fields that are both, and the
new metadata is applied after it so that LongText writing the same key as
false cannot win: either annotation saying yes is enough.

The model field also gains a description, so the editor can say what an
empty value and a placeholder each do rather than leaving both to be
guessed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 14:31:12 +02:00
Lucio Lelii 3c5b44e108 Let a node's LLM model come from an input instead of the configuration
In the four blocks that hold an LLMDescriptor - LLM, ChatInteraction,
Conditional and Switch - the model can now be decided at run time rather
than only at design time. Leaving the field blank turns it into a "model"
input port; writing ${{global.modelName}} into it resolves the value from
a global input, which is the only way one can reach a block since globals
travel through template interpolation and never feed a port.

The mechanism already existed: @ConfigurableAsInput, until now used only
on the flat model field of the MCP blocks. What was missing was reaching a
field one level down. configurableInputDescriptors now descends into the
objects a configuration holds, and the editor needed nothing at all - it
already reads the binding at a dotted path and exempts a bound field from
its required marker.

Three states, and only the middle one is new. No descriptor at all offers
no port: a freshly dropped node is a scaffold, and a port there would ask
for an input before a provider had even been chosen. A descriptor with the
model set needs no port. A descriptor whose model was left blank is asking
for one.

@NotBlank had to go from LLMDescriptor.model, because blank is the state
the binding is made of - the MCP equivalent does not carry it either. The
consequence is deliberate and changes a tested behaviour: a blank model
used to make the flow DRAFT, and now makes it EXECUTABLE. That is right,
because an unconnected port is simply one of the values the run asks for
before starting, exactly like the ${{...}} placeholders a prompt declares.
The field stays required in the published schema so the editor still marks
it.

The port is named "model", which a ${{model}} placeholder in the same
block's template could also claim. Deduplicating the two would be worse
than failing - the value wired for the prompt would silently become the
model - so that collision is refused with an error naming it.

LLMDescriptorInputBinding is the single place the rule lives: the port
wins over the template, a configured literal is returned unchanged rather
than copied, and an empty port never blanks out a configured model. The
provider is deliberately not bindable: it decides the credential, which
AuthorizationRequirementResolver has to resolve before the run starts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 14:24:01 +02:00
Lucio Lelii a1163adb46 Let hosted providers take a typed model, and cap temperature at 1.0
A provider whose catalogue cannot be listed without a credential now says
so, and the editor offers a free text field instead of a select. Gemini is
the first: its four hardcoded model names went stale as fast as Google
renamed them, and the whitelist guard rejected models that do exist.

canListModels() is a declared capability rather than an inference from an
empty list, because for our own Ollama an empty answer means it is
unreachable - not that anything goes. isOpen() answers it per field on a
sibling endpoint, the way isRequired() already does, so nothing needed a
new annotation or a new schema key.

Temperature now stops at 1.0, below what Gemini's API accepts: past that
the output is noise, and one range that holds for every provider beats a
per-provider ceiling nobody can remember.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 11:53:28 +02:00
Lucio Lelii 19f9fcdbe0 Mark a nested object as a group of optional settings
@UiOptionalGroup says what a field is - a group of settings that are all optional
- rather than how to draw it. The editor will render it as one control that opens
a dialog instead of unfolding five empty chips inline; the read-only execution
view will keep showing only what was actually set. Two renderings, one
declaration.

A field annotation, not a type one, and that is forced rather than preferred.
Shared definitions are hoisted into sharedDefinitions only when their JSON is
identical in every schema containing them, so a label that varies by owner would
un-share ModelParameters silently. The web merges a property's x-ui-* keys over
the definition it $refs, so the renderer sees the marker on the object node
anyway - the constraint costs nothing.

The test asserts the marker is on the property, that the $ref survives beside it,
that it is absent from the definition, and that ModelParameters is still shared.
Making the label vary by owner fails it.

Applied to LLMDescriptor.parameters and to both MCP configurations. Nothing reads
it yet; the editor comes next.

534 backend tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-04 12:24:34 +02:00
Lucio Lelii 871acc4423 Stop the helper methods leaking into every stored flow, and pin the MCP wiring
Found by saving a real block against the running service: the persisted
llmDescriptor.parameters carried a sixth field, "empty". Jackson reads an
is-prefixed no-arg boolean as a property, so isEmpty() was being serialised into
every stored flow, every execution snapshot and every API payload - in a shape
the published schema does not declare. Both helpers are @JsonIgnore now, and a
test asserts the serialised object has exactly the five parameters.

The MCP bridge cannot be exercised from here - it answers 403 Forbidden even on
/health - so what its live behaviour does with these values is still unverified.
What is checkable is that the block's parameters reach the call at all, which is
the half that lives in this repo: MCPAgentChatExecutor now has a test capturing
the argument passed to openSession. Passing null instead makes it fail. A field
that exists on a configuration and is never passed on is precisely the silent
hole this codebase has produced twice already today.

534 backend tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-04 11:44:06 +02:00
Lucio Lelii acd46dd566 Pass the sampling parameters through every model call
The seven executors that call a provider now hand it the descriptor's
parameters, and the MCP blocks carry their own alongside their bare model name.
Where a provider does not honour one, the executor says so on the run through
ModelParameterReporting - the provider knows what it supports but has no event
logger, the executor has the logger but not the knowledge, and this is the one
place they meet so the seven cannot drift.

The thirteen duplicated `authorization == null ? twoArg : threeArg` ternaries are
gone: the new call shape tolerates a null authorization, so each is one call.

MCPAgentService.openSession lost its three-argument overload rather than gaining
a fourth. With the overload, the executors moved to a signature the test mocks
did not stub and the calls returned null - a NullPointerException at runtime
where the compiler could have said it instead. One signature, and the eight stubs
had to be updated deliberately.

Fixed while here, because it is what the change surfaced:
ExecutionContext.addEvent iterated a plain ArrayList of listeners while a
container child registered one from another thread. The resulting
ConcurrentModificationException failed the whole subflow and reported itself only
as "ended with status ERROR". It went from unseen in three baseline runs to two
failures in four once these calls shifted the timing. Both lists are
CopyOnWriteArrayList now; five runs since, no recurrence.

532 backend tests green. A separate, pre-existing flake remains - container runs
occasionally time out waiting on a status - which predates this work; it appeared
once in five runs earlier today on untouched code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-04 11:10:24 +02:00
Lucio Lelii f22fdf9161 Add optional sampling parameters, and teach the providers to apply them
temperature, topP, topK, maxTokens and seed, all nullable and independent: unset
means "leave it to the provider", so an absent object behaves exactly as before.

Seed is here for a reason specific to this product. Comparing two runs and
reading a bias report both currently end with a warning that model output varies
on its own; with a seed that warning becomes a thing you can rule out.

A type of its own rather than fields on LLMDescriptor, because the MCP blocks
have no descriptor - they carry a bare model name - and need the same parameters.

Three new default methods on LLMProvider carry them, each delegating to the
existing one, so the two real providers can override while the six test stubs
keep behaving as they did. They also tolerate a null authorization, which is what
will let the thirteen duplicated `auth == null ? twoArg : threeArg` ternaries at
the call sites collapse to one call each.

Two things guarded by tests rather than by care:

LLMDescriptor.provider and .model gain @JsonProperty(required = true). The editor
marks every descendant of a required object as required unless that object
declares its own required list, so without this the new optional parameters would
come out mandatory and every existing LLM node would be reported invalid.

Ollama's JSON path keeps forcing temperature 0.1 and num_predict 4096. Those
predate this change and the flow assistant's output depends on them; a set value
overrides its own field and nothing else. Removing them fails two assertions.

Numeric bounds are mirrored into the schema by hand, the way @Size already is.
Registering the jakarta-validation victools module would have retroactively
rewritten every annotated field in the product.

Gemini declares four of the five as supported: seed varies by API version, and a
knob that silently does nothing is worse than one reported as ignored.

527 backend tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-04 11:02:33 +02:00
Lucio Lelii 2653cd974f Stop a bulk global input save from wiping the other values
Saving the inputs of a run could lose the values the user had just typed. The
client sends a multi-value global through PUT /executions/{id}/globals and a
single-value one through PUT /executions/{id}/globals/{key}. The bulk endpoint
delegated to setGlobalInputDescriptors, which *replaces* the whole set: saving
the list therefore cleared every global not named in that one payload, and
registerGlobalInputs then re-registered them with null values. With the panel's
single Save firing all the requests together, the list save landed after the
others and took two typed fields down with it.

"Set these values" is a merge. Replacing the set is a different operation and
keeps its own method, still used where a replace is what is meant - carrying
descriptors into a container child, and copying them onto a rerun.

Project context seeding also passes a partial map here. It was unharmed only by
accident: everything it wiped was still null at that point.

The test sets one global, then sets another in bulk, and asserts the first
survives. It fails on the old delegation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-03 15:56:55 +02:00
Lucio Lelii 2ea8e33987 Say why a subflow cannot start, and propagate its globals from one place
Two follow-ups to the iterator/loop global input fix.

**One propagation.** Every container path carried the parent's global inputs
into its child with its own copy of the same code, and the copies drifted: the
one in ExecutionsService read the unprefixed variables map through a "global."
view and handed every iterator and loop child nulls, while the copy in
GenericContainerExecutor read the runtime map and kept working. Both now call
SubflowGlobalInputs.descriptorsFor, so there is nothing left to diverge. It also
sets `multiple` on the descriptor, which neither copy did.

The IODescriptor-to-kind switch had grown four identical copies for the same
reason, one per descriptor-building site. It now lives on ExecutionVariableKind
as forDescriptor.

**A message that says something.** Starting a non-READY execution reported only
"is not in READY status (CURRENT STATUS is CREATED)". For a container subflow
that named an execution the user never sees and gave nothing to act on. What
holds an execution back is one of three things, so notStartableReason names it:
missing global inputs, missing credentials, or inputs still to provide and on
which step. startContainerChild is the single point every container path starts
its child from, so it frames that with the container's own name:

  The subflow of container 'Candidate loop' cannot start: no value for the
  global input 'who'

and the error is filed against the container step, so the diagram points at it.

Also fixed: BranchRejoinConcurrencyTest did not compile on its own (a capture
conversion the Eclipse compiler rejects), which the full suite had been hiding
through incremental compilation.

516 tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-03 15:34:21 +02:00
Lucio Lelii 4fa87ae3a2 Carry the parent's global inputs into an iterator or loop subflow
A container subflow is its own execution, so the global inputs typed on the
parent have to be handed to it. createAndStartSubflowChild took a "global."
view of getExecutionVariables(), but that is the *unprefixed* variables map -
the prefixed keys live in runtimeExecutionVariables, which is deliberately not
exposed. The view therefore always matched nothing, and every child was created
with null values for the globals its blocks reference.

The child then never left CREATED, and the parent failed with

  Execution with id <child> is not in READY status (CURRENT STATUS is CREATED)

naming a child execution the user never sees and saying nothing about the cause.
GenericContainer was unaffected: its executor has its own copy of this
propagation that reads the runtime map, which is why only iterator and loop
subflows were broken.

The fix reads the parent's global inputs from the map that holds them.

The test runs an iterator whose subflow prompt references ${{global.who}}, and
fails with exactly the error above when the line is reverted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-03 14:53:09 +02:00
Lucio Lelii 7e69e21197 Keep a global input's multiplicity at execution time
A global input declared multiple in the editor was offered a single-value editor
when the flow ran. The client picks the editor from the execution's variable
descriptor, and that descriptor had no multiplicity to pick from: registering a
global input mapped the IODescriptor through ExecutionVariableKind, which encodes
only the type, so `multiple` was dropped on the way in. The frontend already read
`multiple` off the descriptor - it had simply never been sent.

The descriptor now carries it, set from the flow when global inputs are
registered. It is also carried explicitly through ExecutionVariableRegistry
.normalize, which rebuilds descriptors field by field: anything not named there
is silently dropped on every register, set and restore, which is how a field like
this disappears without a single error.

Covered end to end - declaration, supplying one value, supplying them all at
once, and a snapshot round-trip - because each of those rebuilds the descriptor
by a different path. Removing the one-line fix fails all four.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-03 14:04:57 +02:00
Lucio Lelii f9b94d823e Add projects: grouping, shared context, and sequential project runs
A project groups 1..N flows. A flow's project is optional, so flows without one
stay fully valid, and the association is a plain String column - this codebase
has no JPA relations, and the hottest entity is not the place to introduce the
first one.

Ownership and visibility
- Projects are strictly owner-scoped and report another user's project as 404,
  not 403: unlike a flow, whose existence is not secret, a project is private
  workspace structure.
- projectId and projectName are disclosed only to the flow's owner. A published
  flow inside a project stays readable by everyone - membership is organizational
  structure, not access control - but a project name can itself be sensitive, so
  a non-owner sees the flow as unassigned.
- Assignment gets its own PUT /flows/{id}/project rather than a field on the flow
  body: that body is the full-replace PUT the editor issues on every save, so a
  project carried there would be silently dropped on each save. A finalized flow
  stays reassignable, since finalizing is irreversible and must not freeze a flow
  out of reorganization forever.

Deleting a project deletes its flows, finalized ones included, and requires
confirm=true. Refusing the cascade was not an option: finalizing cannot be
undone, so a project holding one finalized flow could never be deleted by any
API. Executions are kept - each snapshots its own copy of the flow graph.

Shared context
Values a project shares with its flows, readable as ${{project.x}} and as
#project['x'] in conditional expressions. Three changes make that work, each
silent if missed: the project. prefix is preserved by ExecutionTemplateResolver
(otherwise keys publish as ${{vars.project.x}} and never resolve), admitted by
PlaceholderInputs (otherwise the placeholder becomes a dangling block input), and
bound in the SpEL context. Values are frozen into the execution snapshot at
creation, so editing a project never rewrites a run that already happened, and
they are applied only when the person running owns the project.

Project runs
POST /projects/{id}/execute creates one execution per flow, sharing a
projectRunId, and refuses the lot if any flow is not executable - a half-created
run group is worse than a clear refusal. POST .../runs/{runId}/start then runs
them one at a time in the project's order, each step starting only when the
previous succeeded. The run keeps no state of its own: its position is derived
from the executions, so there is no second status machine to drift out of sync.
A failed step stops the run and leaves the rest untouched, so it can be resumed.
Project values pre-fill matching global inputs - same name and type, never
overwriting a user's own value - which is what lets a run start without a
per-flow round trip.

Also fixes an ordering hazard in AuthService: account deletion deliberately
preserves finalized flows, so their project assignment is now cleared before the
projects are removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-03 07:53:25 +02:00
Lucio Lelii 61bbbe50d4 Sharpen block descriptions and add placeholder tips
Rewrites the MCPAgent and MCPAgentChat descriptions so the catalog says what the
blocks are actually for - calling a declared MCP server's tools - and when to
prefer them over HTTPServerCall or LLMBlock.

Adds a tip and acceptVariableAsPlaceholder to the two human-facing block
configurations, spelling out how ${{}} placeholders behave: in the interactive
block each one becomes a real named input, while in the decision block they are
additional read-only context alongside the forwarded input port.
2026-09-03 07:52:52 +02:00
Lucio Lelii ee695b2f5f Finalize async-only assistant flow endpoints, sync client flow edits into session
Removes the now-unused synchronous /flows/draft, /flows/refine and /flows/fix
mappings in favor of the session-based submitMessage + polling flow, and lets
submitMessage accept an optional flow snapshot so manual canvas edits made
outside the chat are reflected before the assistant acts on the next message.
2026-09-02 13:32:06 +02:00
Lucio Lelii 7f57b8683a refactor(executions): extract authorization requirement resolution into AuthorizationRequirementResolver
Cluster M from the structural analysis: scans a FlowData (recursively
through container subflows and loop guard subflows) for blocks/steps
that need an authorization value - LLM provider credentials and HTTP
server-call auth - and aggregates them per requirement key. Only field
dependency is llmProviders, now passed as a parameter.

- New AuthorizationRequirementResolver holds resolveRequiredAuthorizations,
  collectRequirements, resolveDescriptors, listOfDescriptors,
  collectRequirement, collectHttpRequirement, resolveProvider, and the
  private RequirementAccumulator helper class
- Both call sites (execution creation, execution rebuild) and the
  cross-cluster call from Cluster L's validateCredentialReference
  updated to pass llmProviders explicitly

1620 -> 1479 lines. Behavior-preserving: pure extraction plus explicit
parameter-passing for the one field this cluster touches, no logic
changes.

Verified with `mvn test`: 462 tests, 0 failures, 0 errors (note: this
suite has pre-existing intermittent flakiness under parallel load in
Loop/Iterator container tests, unrelated to these changes - confirmed
by re-running the full suite multiple times with a different failing
test each time, then a clean pass).

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
2026-09-02 10:57:56 +02:00