Execution views now carry the round each step is on and the steps of
every earlier round, which is what a page needs to show a loop's history.
A node in a loop asks the same question once per round. The interaction
and evaluation endpoints take an optional iteration: an answer for a
round the loop has already left is refused with 409 instead of being
taken as the answer to a draft its author never saw.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When a loop's guard chooses the output that leads back, the execution
resets the loop's steps and starts the next round from the entry with
the value the guard produced. Every step of the loop leads to the guard,
so all of them have finished by then; the reset and the delivery happen
under one lock.
The back edge is not wired like other connections, which push their
value the moment it is produced, before the loop has been reset. Left
unwired, the entry's input still counts as open when nothing else feeds
it, so the first round takes it from the person starting the run. Inputs
from outside the loop keep their values from round to round. The guard's
other outputs are not marked as branches not taken while it goes round,
since a later round may leave through them.
Each step knows its round, events of loop steps carry it, and each
round's steps are kept in the execution's step history, so earlier
rounds and earlier verdicts stay readable after a reload. Choosing to go
round with no rounds left fails the execution with
LOOP_ITERATION_LIMIT_REACHED rather than leaving the loop.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An execution waiting for a person is read back from storage whenever it
leaves the in-memory cache or the service restarts. Its steps were rebuilt
with no executor, and resuming a waiting execution never gave them one. A
person's answer then made the next step ready, the step had nowhere to
run, and the execution sat in RUNNING for good.
Resuming a waiting execution now attaches the executor to its steps,
scheduling only a step that was already ready to run.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The execution-order check refused every cycle as a deadlock, and the
branch check refused a router whose branches meet again - which, going
round a loop, they do. Both now set a loop's back edge aside.
A cycle that cannot run gets the reason in its own code instead of a
generic deadlock: nothing to decide when it stops, a side way out, no
way out, loops inside loops, or a limit out of range. The assistant
treats them like the deadlock it already re-planned for.
The engine does not run loops yet; that is the next change on this
branch.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A loop can run when it goes round through a block whose type routes
exclusively: that block decides each time whether to go again or leave,
and the connection it goes round by is the back edge. Nothing marks that
connection; FlowLoops finds it from the graph, so validation and the
engine agree on it and imported or generated flows need no marking.
Only shapes the engine can run safely are accepted: one back edge per
loop, and nothing leaving the loop except through the guard's other
outputs. The author's iteration limit rides on the back edge, from 1 to
100, defaulting to the Loop container's 10.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Which blocks pick one output and leave the rest as branches not taken was
written out twice by name: in branch validation and in the engine's
not-selected marking. It is now a capability the type declares, so a
routing block added later is treated as one without either learning its
name - and loops can use the same capability to find their guard.
The loop container's default of ten iterations moves to a shared
constant, for connections that lead back to reuse.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Rebuilding dependency state on reload re-asked every step's activation
policy. A step that had already run has all its inputs, so the policy
answered READY, and a completed step came back from the database as one
waiting to start. Its dependents were then never told it had completed.
Activation now only decides for steps that have not run yet.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The catalog hides the IO lists the runtime derives, but keeps structural
configuration that happens to be called inputs or outputs. An LLM node's
declared output fields are that kind: the author writes them and each one
becomes a port. The test predated them and counted them as hidden IO.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An agent builds something, runs it and looks at it in a browser; a person then
opens the running preview and judges it, with the agent's own report hidden
until they have. That last part is the reason the evaluation node exists, and
until now there was nothing running for anyone to judge.
Checked here rather than only by being imported, because the two halves are
wired to each other by name. The agent declares previewUrl and agentReport; the
evaluation node names those same two in its target and reads one of them as its
blind reference. A typo in either imports perfectly cleanly and then hands the
evaluator an empty address at the moment they are asked to open something.
The agent node must prove it used start_service, browser_navigate and
browser_snapshot before its report is accepted. That is not caution in the
abstract: this node has already reported building and testing something without
making a single tool call, and nothing in the report said so.
It reports previewUrl and not publicUrl, and the prompt says why they differ -
one is reachable from inside the container network and the other by a person.
Handing over the wrong one produces an address that works for everybody who
tests it and for nobody who is asked to use it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A node could only ever produce one port: its whole answer, as text. That is
enough until it has two things to say to two different readers - the address of
something it started, and the report about it - at which point whoever is
downstream has to pick them apart by guesswork, or a parser block has to be
wired in to split on a delimiter the model may or may not honour.
Declaring fields turns the answer into ports the editor can wire. The mechanism
is only asking: the prompt gains an instruction to close with a JSON object, and
that object is read back out of the reply. No provider feature is involved, so
it works the same for a single call and for a whole tool loop, on every provider
in the catalogue.
It refuses rather than guesses. A missing object, or a declared field the model
left out, fails and says which - the same choice the loop already makes for a
tool call typed as prose. A port that silently arrives empty is worse, because
everything downstream reads the emptiness as an answer. The whole reply stays on
the response port either way, so a node whose structure disappointed is still
readable.
The closing object is found by scanning forwards, tracking strings: the last
balanced one wins, because an agent that has just written a JSON file tends to
quote it before answering. Scanning backwards was the first attempt and cannot
be made correct - a quote cannot be told from an escaped one without counting
the backslashes before it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A turn with no tool calls is how the loop knows it has the final answer. A model
that writes "<function=list_services>" into its content satisfies that test
exactly, so the step completed and handed the flow a call that had never run -
a success carrying nothing, which is worse than a failure because everything
downstream believes it. Seen with qwen3-coder on Ollama, which advertises tool
support and then answers this way when it is left to decide about reasoning for
itself: the trace comes back mixed into the content rather than as tool calls.
The guard reads only markers a chat template emits and prose never does, so an
answer that merely talks about tool calling still comes back untouched. The
message says what to try, because the cause is the model and not the flow.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A flow could write software but never run it: the coding agent edits files and
that is where it ended. These two entries close the loop - start the code the
execution wrote, then open it in a browser and read what the page says.
Both are streamable-http, so an LLM node binds them with no new Java: the
retriever already filtered the catalogue by transport and the draft normalizer
already dropped unknown servers. The development server's preview is pinned to
the execution with x-preview-key, the way the coding agent's workspace already
is, so the scope travels in a header the service fills in rather than in an
argument a model could choose for itself.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Nothing could answer it. Authentication is a stateless JWT, so there is no
session table to count; nothing streams to the browser, so there is no open
connection to count either. LoginEntity.currentSessionStartedAt looks like the
answer and is not - it only clears on an explicit logout, so anyone who closed
a tab stays "in session" indefinitely. Fair input to an average session length,
useless as a list of who is here.
ActiveUserTracker records a sighting from the authentication filter, the one
place every authenticated request passes. In memory rather than a column: a
write per request would put the busiest path in the application on the database
to record something nobody reads between restarts, and after a restart "nobody
is connected" is not lost state but the correct answer. Per instance therefore,
which is what this is deployed as; behind replicas each would report its own
callers, and the fix then is to ask each of them.
GET /stats/activity returns that alongside the runs a restart would interrupt -
RUNNING, SUSPENDED and WAITING, every state that is neither initial nor final.
WAITING belongs there precisely because it looks idle from the outside: someone
is mid-task on the other side of it, and it is the most expensive thing to lose.
The two are reported separately because they fail differently - an active user
loses an unsaved edit and comes back, an in-flight run is gone.
POST /auth/heartbeat exists only to give the editor something to call: it
answers nothing, the filter having already recorded the sighting. Without it
anyone reading a flow or typing a prompt makes no request for minutes and would
be invisible, which is the person a restart interrupts worst.
Logout was excluded from the authentication filter, so its @AuthenticationPrincipal
was always null and recordLogout never ran - sessions only ever closed on the
same user's next login. Removed from the exclusion list: SecurityConfig still
permits it, so it keeps working with a bad or missing token, and it now both
records the logout and drops the person from the presence map immediately
rather than letting them age out.
Window is five minutes, PRESENCE_WINDOW_SECONDS, wide enough to survive a
couple of missed heartbeats.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every Ollama call carried think:true unconditionally, on the belief - stated
in a comment - that a model without the capability would ignore it. It does
not: Ollama answers with 400 "gemma:7b" does not support thinking, so every
flow pointed at gemma, llama3 or any other non-reasoning model failed outright.
Resolved per model rather than per flow, in two steps so that neither an old
Ollama nor one behind a proxy is left broken:
- /api/show reports a capability list, and "thinking" in it is the answer.
Cached per endpoint and model, so it is asked once, not per call.
- When it cannot be asked - no capabilities field, /show unreachable - the
flag is sent anyway and the rejection itself is the answer: the identical
call is repeated without it and the result remembered. One wasted request
per model per process, and none after that. Matched on the server's own
wording, never the bare 400, so a misspelled model stays the error it is.
Off means the key is absent, not think:false - absence is the only state every
version of Ollama reads as "no opinion".
ModelParameters gains "reasoning" to overrule that per-model answer: ON sends
it whatever the model is and lets an incapable one fail rather than quietly
answering without it, OFF never sends it - the case detection cannot guess,
a model that does reason but where the tokens are not worth it. Empty, which
is what every existing flow has, is the per-model default.
An enum and not a boolean because the editor collapses a boolean to true or
false as soon as its group is opened (schema-driven-fields, next[key] =
rawValue === true), which would have turned reasoning off for the models that
have it. It follows MCPAgentUploadKind, lenient @JsonCreator included, since
a select nobody chose from posts "" and an enum cannot be coerced from that.
THINKING is deliberately outside LLMProvider.supportedParameters()'s default,
which is no longer allOf: only the Ollama protocol has the switch, and a
provider inheriting a claim to it would leave the run log silent about a
setting that reached nothing.
The MCP agent path carries the same setting but cannot detect anything - the
bridge is not Ollama and exposes no capability endpoint - so there it is
whatever the flow says, defaulting to on as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The attribution term of section 7(b) names the source, so the URL it names has to
be the one that actually serves it: anyone redistributing this must be able to
reach the original. Eight references moved together - the licence addendum, the
NOTICE, the README attribution, the citation metadata and the four in the POM,
including the SSH developer connection.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three things that belong to one machine rather than to the project: the README
named the build host and the account on it, the Ollama default URL pointed at an
internal deployment, and docs/ held thirty working notes written for this team.
The Ollama default is now the address Ollama listens on out of the box, so a
clone runs against a local one without editing anything.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
local.env carries the database password and the provider key of whoever runs the
service: it is one person's machine, not the project. .claude/settings.json is the
same kind of thing, a local tool's permission list. Both stay on disk and are now
ignored.
This removes them from future commits only. Their contents remain in the history
that is already on origin, so the credentials they held must be treated as known
to everyone with access to the repository.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The budget is in characters and a model's window is in tokens, and the
ratio between them is not a constant: prose runs about four characters
per token, a conversation of JSON, file paths and UUIDs closer to two and
a half. 60000 was sized for the first. An orchestrator reading its own
registry reached 66332 characters - roughly 26k tokens with the answer
still to generate - and a 32k window dropped the opening message, which
the provider reports as "no user query found in messages".
Halving the per-result cap attacks the same failure from the other end:
the current iteration's results are the ones pruning can never shrink, so
what one turn reads is what decides whether the next call fits. A model
that reads the same file twice in a turn spent 24000 characters of a
60000 budget on one file.
Every call now records what was sent and what was allowed back. Without
promptChars and maxTokens, a context failure could only be reconstructed
by reading the node's configuration beside a warning about characters -
which is how this one was diagnosed, slowly.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A verdict off the criterion's scale, or naming a criterion the node does
not have, is the caller getting it wrong. Reported as a 500 it reads as
an outage, and the message saying exactly which criterion is at fault -
the one thing that makes it fixable - gets buried in a stack trace.
Found by submitting one against a running service, not by reading the
code: the executor was right to refuse, the endpoint was wrong about
whose mistake it was.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The human nodes could ask a question or collect a paragraph, so a flow
that wanted a judgement about a running piece of software got back prose:
not comparable between two runs, and nothing a later node could branch on.
HumanEvaluationBlock asks instead for a verdict per named criterion, with
an optional script so two testers exercise the same thing, and evidence
files. It points at the target rather than hosting it - provisioning
environments or driving browsers belongs outside the engine, and the flow
already knows the URL.
blindUntilSubmitted keeps a reference verdict, usually an agent's, out of
sight until the person commits to their own, and the execution records
which came first. That turns the node from a place to rubber-stamp an
agent into something that measures the automation bias the catalogue
already names.
The engine needed only a way to submit several fields as one act, which
Step.interact already accepted; and the interaction contract now comes
from the block type, so the editor and the assistant stop keeping their
own table of block names.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A skill on a node is only an id, and the editor grew a rule of its own to
show the instructions behind it: it looked for the literal retriever name
"Skills" and built the endpoint by hand. That is a special case inside
machinery that is otherwise entirely schema-driven, and the next binding
that wanted the same view would have needed another one.
@FieldRetriever now takes a definitionUrl, emitted as
x-retriever-definition-url, and the editor offers the view wherever it is
declared. The MCP server binding deliberately declares none: its
definitions endpoint answers with a JSON schema rather than readable text.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The bridge coerces a tool argument from string to object whenever the string
happens to be valid JSON, even though the model emits it correctly and the
tool's own schema declares it as a string. The coercion happens inside the
bridge, between the model's response and the downstream MCP call: verified by
calling the model directly (arguments.content stays a string, byte for byte)
and by comparing a JSON-valid value (rejected) against a malformed or plain
text one (passed through untouched). The failing call never reaches the MCP
server, and the agent's retries under a different encoding are what turn an
unfinished operation into one that returns status: completed with the model's
last preamble as if it were the answer.
There is no fix available on our side for the bridge itself, so this removes
it from the path instead. An LLMBlock can now bind MCP servers straight from
the catalog and run its own tool-calling loop in this service:
- LLMProvider gains chatWithTools/supportsTools; only OllamaProtocolProvider
implements it for now. Tool arguments stay JsonNode end to end - never a
string, never re-parsed - which is the one change that actually closes the
bridge's bug rather than working around it.
- A native streamable-http MCP client (mcp/client/) talks to a server without
the bridge: initialize, tools/list, tools/call, session header handling,
both response shapes the spec allows.
- MCPToolServerBinding is a narrower binding than MCPAgent's, restricted to
catalog servers reachable over streamable-http - the ones this service can
call directly, not the stdio ones the bridge still hosts a process for.
- LLMToolLoop runs the model/tool/model cycle with real budgets: a wall-clock
deadline and iteration cap that fail the block explicitly rather than
return a partial answer, and a character-based context budget that replaces
older tool results with a placeholder once the conversation - plus the tool
schemas sent on every call, which do not appear in the conversation but are
not free either - grows past it. The iteration just completed is never
pruned, and a result under ~500 characters is left alone: shrinking it would
cost about as much as it saves.
- Ollama's done_reason now travels back as ToolChatResult.finishReason, so a
turn that answers nothing can say whether the model chose silence or
num_predict cut it off mid-thought - two different problems with two
different fixes, previously indistinguishable from the error alone.
- A new skill, mcp-context-economy, carries the operating rules a real run
against a 24-task plan exposed the hard way: write_file to create a file,
apply_patch only to edit one that exists, and never read a file straight
back after writing it or re-pull an already-inline document into the
conversation - each halves the context a node needs for the same work.
MCPAgent and the bridge are untouched: this is a second path, not a
replacement, for the one transport (streamable-http) this service can reach
without it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The editor derived the control from the field being optional, which put it on
nearly every field in every dialog. "You may leave this blank" and "leaving this
blank means something specific" are different claims, and only the second is
worth a control.
The claim is now made per field with @DefaultsWhenEmpty, published as
x-ui-defaults-when-empty. It goes on the five sampling parameters - where empty
means the provider decides, and no typed number gets that state back - and on the
three fields that declare a concrete default, which the editor already names
alongside. The value itself still comes from JSON Schema's own `default`: a
parameter has no value to name, only an absence to return to, so declaring
`default: null` would have said something false to every other reader of the
schema.
Providers also now report which sampling parameters they actually apply. All five
were offered to every provider and the unsupported ones were dropped at run time,
reported in a warning on an execution that had already happened - Gemini applies
no seed, the OpenAI-protocol providers no top_k.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Nobody checked that descriptor.provider actually named a registered
LLMProvider bean: a removed or misspelled provider surfaced only at
runtime, with a bare "Provider not found" and no indication of which node
named it.
FlowDataValidator.validateLlmDescriptorProviders runs from validateBlock,
so the existing subflow recursion in validateContainerSubFlow already
covers every container and Loop guard subflow for free. Only the provider
is checked, never the model (that's step 12, and a hosted provider's
catalogue isn't known here anyway), and a templated provider name is
skipped defensively even though nothing in the codebase ever writes one.
Since this constraint backs ValidFlowStructure, the new
LLM_PROVIDER_NOT_FOUND error surfaces two ways: saving a flow with one
still succeeds, as DRAFT, with the error in its validation list - the same
treatment every other not-yet-executable state already gets - while
creating an execution from one is rejected outright, since there would be
nothing such an execution could ever do.
This surfaced a pre-existing, widespread test convention: seven structural
validation tests used a placeholder provider name ("testProvider") that
was never a real bean, only ever exercised through flow save/execution
creation, never through an actual provider call. Renamed to "InternalOllama"
in all seven, the one provider name always registered in a full Spring
context.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A wrong model name only ever failed on the first call to it, which can be
minutes into a run for a step deep in a flow - by which point the endpoint
and credential that would have let it fail immediately were already known.
AuthorizationRequirementResolver.resolveAllDescriptors mirrors the existing
requirement-collecting walk (blocks, containers, Loop guard subflows) to
list every LLMDescriptor in a flow instead. ExecutionsService verifies each
one against its provider's own catalogue, for a provider whose
canListModels() is true, before starting - today that is only our own
Ollama, since every hosted provider declares canListModels() false
precisely because it cannot be asked without a credential the check does
not have. A model that is empty or still a template placeholder is
skipped, since its real value is only known at call time; a catalogue
that cannot be listed just now does not block the run either - this is a
defense in depth, not a gate a transient network failure should be able
to close.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
BiasJudgeRequest carried no way to supply a credential: a judge whose
provider required one only ever worked by coincidence, when the baseline
execution it judges happened to already carry a saved credential for that
same provider. There was nowhere to pick a credential for the judging
itself, unlike the interaction simulator and the assistant.
BiasJudgeRequest now carries an optional credentialId, mirroring
AssistantLlmSelection and the simulator's own field from the previous
commit. It flows through the whole asynchronous path -
BiasExperimentsController, BiasImpactJobService.createJudgeJob (a new
judge_credential_id column on BiasImpactJobEntity, so a job recovered after
a restart keeps it), BiasImpactService.judgeReport - down to
BiasImpactJudge.resolveAuthorization, which resolves it via
UserSecretService.resolveCredential when present and falls back to the
existing baseline-authorizations lookup otherwise, unchanged.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
startSimulationExecution resolved the simulator's provider authorization
lazily, inside whichever step happened to reach it first - a credential-
requiring simulator would start a RUNNING execution that then failed deep
inside a step, instead of being refused outright. ExecutionSimulationRequest
now carries an optional credentialId (mirroring AssistantLlmSelection), and
the service validates or falls back to an already-provided credential for
the same provider before starting simulation at all.
Fixing this surfaced a second, previously silent gap: a simulated
container's child never inherited the simulator's credential, since it is
never part of any execution's requiredAuthorizations and so the ordinary
per-container authorization propagation loop never touched it. Every
simulated container subflow would have started failing the same upfront
check once it went in, so propagateSimulatorAuthorizationToChild copies the
already-validated credential down to the child before it starts simulating.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
resolveProviderAuthorization used to name "InternalOllama" as the only
provider that could be used from the assistant without a credential, and
reject every other one with 409 - including a provider that plainly
declares requiresAuthorization() false, such as a credential-free remote
Ollama. The rule is now exactly that capability: !requiresAuthorization()
means no credential is asked for, whatever the provider is called.
The now-dead INTERNAL_PROVIDER_NAME constant goes with it - nothing else
referenced it.
AssistantSelectionResolverTest is new: this method had never been tested
in isolation, only indirectly through AssistantControllerTest, which
never exercised a credential-free provider under any name but the one
that used to be hardcoded.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Gemini is the provider that verifies AbstractHttpLLMProvider is really
transport only: it authenticates through a "key" query parameter, not a
header, so if the base assumed a bearer header this migration would have
had nowhere to go. It did not need to - the base has no opinion on
authentication mechanics at all, so query-string auth needed no hook,
just its own call site, same as before.
GeminiLLMProviderBodyTest was written first, against the
still-unmigrated provider, specifically to survive this move: it pins the
retry policy (ten attempts, 30s backoff), the two-minute timeout, the
role mapping (system and user both become "user", only assistant becomes
"model"), and the deliberate exclusion of seed from supportedParameters -
every one of which the interface's own defaults or the base's own
defaults could have silently replaced if an override were dropped by
accident. All thirteen assertions pass unchanged after the migration.
The retry policy and timeout are now the explicit overrides
AbstractHttpLLMProvider expects (retryPolicy(), timeout()) rather than
being built inline in the one method that used them - same values, same
pinned constants, just named as what the base already knows how to ask
for.
Gained for free, the same way Ollama did: 4xx responses now carry their
body, and every failure is logged - Gemini previously had no logging of
its own at all.
A live end-to-end HTTP test was attempted and dropped: Gemini's base URL
is a private constant with no way to redirect it to a local test server
without either changing production code or wrapping WebClient.Builder in
a test double fragile enough to break on the next Spring release. Given
that AbstractHttpLLMProviderTest already proves the shared HTTP mechanics
generically, repeating that proof through Gemini's specific, unreachable
URL would not have added real coverage.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
InternalOllamaLLMProvider carried its own client, its own error mapping
and logging, and its own request bodies and parsing, all mixed together.
OllamaProtocolProvider pulls out everything that is genuinely the
protocol - nested options, think:true, top_k, the JSON-path defaults the
flow assistant depends on, response parsing - onto the transport base,
leaving InternalOllamaLLMProvider about fifty lines: its own URL, its own
key, and nothing else. getName() still returns "InternalOllama" - that
string is persisted in seventeen places in workflow-editor-init/flows.json,
in every existing flow, and in the vault's provider column, so it could
not change even in a refactor this size.
Not extended from OpenAIProtocolProvider, even though Ollama also exposes
an OpenAI-compatible endpoint: the native shape differs enough - nested
options, think, top_k, none of which OpenAI has - that a subclass would
override every method the parent provides, which is not a subclass, it is
a different implementation wearing one.
The base ended up with two hooks instead of the OpenAI family's one,
because there is a real asymmetry here that family does not have:
resolveApiKey() lets InternalOllamaLLMProvider ignore whatever credential
a caller passes and always use its own server-configured key, while
RemoteOllamaProvider - the new provider, for an Ollama instance other than
our own - requires the caller's. baseUrl() has the same shape as
OpenAICompatibleProvider's: a constant for the internal instance, read
from the credential (and validated through OutboundEndpointGuard) for the
remote one. RemoteOllamaProvider cannot list its models either, for the
same reason OpenAICompatibleProvider cannot: listing would run from the
editor, with no credential and therefore no endpoint to ask.
InternalOllamaLLMProviderBodyTest - the existing test pinning every
request body byte for byte - passes unchanged, which is what "extraction"
is supposed to mean here. InternalOllamaLLMProviderHttpTest is new: no
test before this one exercised the actual HTTP round trip, only the
bodies, so there was no way to confirm the 4xx-carries-its-body upgrade
(the whole reason for building the shared base) actually reached Ollama
until now.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
OpenAIProtocolProvider is the shared request/response shape - flat
sampling parameters, native system/user/assistant roles, no top_k (OpenAI
has no equivalent, so it is excluded from supportedParameters rather than
silently ignored) - built on the transport base added earlier. Three
concrete providers sit on it:
- OpenAIProvider and OpenRouterProvider have a constant endpoint, the
way any client of either service does; requiresEndpoint() stays false
and their baseUrl() ignores whatever the credential carries.
- OpenAICompatibleProvider is the one whose endpoint the user supplies -
a self-hosted vLLM, LM Studio, a company gateway - so
requiresEndpoint() is true and baseUrl() reads the credential's
endpoint, validated through OutboundEndpointGuard before every call.
Registering "OpenAI" as a real provider bean was checked against the
existing tests that used that exact name as a stand-in for an
*unregistered* provider (VaultCredentialGateTest, UserSecretControllerTest)
- none of them break, and the ones that specifically assert the
unregistered-name behaviour now exercise the real bean instead, which is
closer to what they were meant to prove.
The 4xx-carries-its-body improvement from the transport base applies here
from day one: with a free-typed model name, that body is often the only
thing that says whether the model does not exist, the key lacks access to
it, or the endpoint is wrong - OpenAI's own error responses say so
directly ("The model 'x' does not exist or you do not have access to it").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
LLMCredentialResolver.resolve now returns ProviderCredential instead of a
bare String, so the endpoint travels with the value from the vault all
the way to the provider that needs it. Every one of the eight call sites
had to change to compile - there was no way to touch only one - so all of
them now pass the resolved credential straight through instead of
unwrapping it first.
That turned out to be the right amount of change, not more than
necessary. Where an endpoint-aware provider is not actually reachable yet
(the interaction simulator and the bias judge choose their descriptor
after the execution already exists, and never had their authorization
requirement computed up front to begin with - a separate, pre-existing
gap this does not close), a missing credential fails exactly as it always
did: LLMCredentialResolver still throws "Missing saved credential" when
the authorizations map has no entry for the provider's key, whether the
caller then unwraps .value() or keeps the whole ProviderCredential makes
no difference to that failure. The only place behaviour actually changes
is the success case, and only for a provider that reads the endpoint at
all - every existing provider still only reads .value() through the
interface's own default unwrapping, so Gemini, InternalOllama and every
test stub keep behaving exactly as before.
LLMCredentialResolverTest is new: this resolver was previously exercised
only indirectly, through a full execution.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A vault secret has always been one opaque value. A provider whose
endpoint is not known in advance - coming in the next commits - has to
get it from somewhere, and the credential the user already picks is the
one place that does not mean typing a URL on every run.
UserSecretEntity gains endpoint: nullable, unencrypted (it identifies
where to connect, not a secret, and needs to be readable to show in a
credential picker), populated only for a provider whose requiresEndpoint()
is true. UserSecretService enforces that at both create and update -
present when required, absent otherwise - and, when present, validates it
before it is ever stored.
That validation is OutboundEndpointGuard: http/https only, no credentials
embedded in the URL, and a rejection of loopback, link-local (including
169.254.169.254, the instance-metadata endpoint on every major cloud and
the single most valuable SSRF target there is), private and multicast
addresses, with an operator override for a legitimate private-network
endpoint. Deliberately not a reuse of HTTPServerCallService's existing
guard: that one resolves a hostname once and never again, which is
exactly the DNS-rebinding gap. This one is meant to be called immediately
before use as well as at save time, so the resolution it checks is the one
about to be connected to - though even then it does not pin the resolved
address for the connection that follows, so it is a baseline, not a
complete defence.
Saving is a courtesy: a comprehensible 400 instead of a mysterious
failure at execution time. The check that actually protects runs where
the endpoint is used, not here - that call site is not in this commit yet.
The provider catalog (LLMProviderMetadata, over the wire at
/llm/providers) grows the matching requiresEndpoint flag, which is what
will let the "Add credential" dialog show the field only where it applies.
Test literals worth a note: OutboundEndpointGuardTest resolves only
literal IP addresses and loopback names, never a real hostname, so it
needs no network access and cannot be flaky because of one. The new vault
tests use public IP literals (8.8.8.8 and 8.8.4.4) for the same reason -
subdomains of example.com mostly do not resolve at all, and a real DNS
name would make these tests depend on network access they should not need.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two additions to LLMProvider, both additive defaults so no existing
provider or test stub changes behaviour.
requiresEndpoint() names the one thing every provider until now has had
in common without anyone needing to say so: a base URL it already knows,
whether server-configured or a constant of the service it talks to. A
provider whose endpoint is not known until a credential names it -
coming next - is the first that needs to say otherwise.
ProviderCredential carries that endpoint alongside the secret value a
provider has always received. The three new generate/generateJson/chat
overloads that take one default to unwrapping .value() and calling the
String-authorization overload above them, so a provider that only
overrides the old ones - which today is every one of them, including
every anonymous test stub across the suite - keeps behaving exactly as it
did. Only a provider that overrides the new overloads directly gets to
read .endpoint() at all.
Adding an abstract method instead would have broken every one of those
stubs, since none of them implement anything beyond the three methods the
interface already requires.
One ambiguity fell out of this at the call site InternalOllamaLLMProvider
used to have: chat(model, messages, null, null) no longer resolves
unambiguously, since a bare null now fits both the String and the
ProviderCredential overload equally. Not visible in this diff - that call
site was rewritten away in the Ollama extraction - but worth naming since
it is the shape of thing this kind of overload addition can trigger
elsewhere too.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Every provider that speaks HTTP wrote its own version of the same handful
of things - a client to reuse, an error mapping, a blocking call with
logging - and each ended up with a different half of it. Gemini has a
retry policy and no logging at all; Ollama has logging but no retry; both
map only 5xx responses and drop the body of a 4xx, which is exactly the
detail that would say whether a model name is wrong, a key lacks access,
or an endpoint is misconfigured.
AbstractHttpLLMProvider consolidates that: a WebClient cache keyed by base
URL (needed once an endpoint can vary per call, which a user-supplied one
will), 4xx and 5xx both mapped through LLMProviderHttpException carrying
the response body, and failure logging that distinguishes an HTTP error
from a connection failure from a timeout.
It is deliberately transport only - no generate/chat/generateJson, no
request body, no parsing. Two hooks, timeout() and retryPolicy(), are
overridable rather than fixed, because Gemini authenticates through a
query parameter rather than a header and any future provider might too;
a base class that assumed otherwise would not be a base class Gemini could
actually sit on.
Not wired into any real provider yet - AbstractHttpLLMProviderTest proves
the plumbing generically, through a minimal test-only subclass and a real
JDK HttpServer, before anything depends on it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Mechanical: three comment lines above the package declaration of every
.java file under src, main and test alike, and nothing else. Its own
commit because it moves the blame line on 501 files and would otherwise
bury the licence change it belongs to.
The short SPDX form rather than the full GNU notice - machine-readable
under REUSE, sufficient to keep the licence notice intact, and it defers
the attribution term to LICENSE-ADDENDUM rather than repeating it five
hundred times.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Same licence and the same additional term as the web repository: the two
halves are one product, and licensing them differently would leave the
question of what a derivative of the whole owes unanswerable.
The AGPL rather than the GPL because this service is meant to be hosted,
and section 13 is what obliges whoever hosts a modified copy to offer its
source to the people using it. The section 7(b) term in LICENSE-ADDENDUM
requires the attribution to be preserved, including in the Appropriate
Legal Notices a derivative displays.
The pom's licence, developer and scm blocks were the empty placeholders
Spring Initializr generates, so the published artifact declared no licence
at all - now they say what the LICENSE file says.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Only result.result is modelled here, so anything else the bridge reports
is dropped at deserialization - which makes a bridge that returned its
answer under another name indistinguishable from one that produced
nothing. "No output generated" is exactly the case where that distinction
decides what to fix, and the bridge's schema is not in this codebase, so
the payload itself is the only way to tell.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The editor polls a running execution and drops the request when it
navigates, refreshes or supersedes it - routine, and more likely the
longer the response takes to write, which an execution view carrying a
large global input does. Each one was logged as "Unhandled request
failure" with a hundred-line stack trace, for something nobody can act on
and where the reply goes to a connection that is already gone.
Recognised through the cause chain, because Jackson wraps the broken pipe
several times over before it surfaces.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
max_steps went only on the JSON query operation, so a query carrying an
attachment still ran on the bridge's default and could abort the same way.
The multipart operation takes it too.
The default drops from 200 to 100, which is the ceiling the bridge
documents: asking for more is at best ignored and at worst refused.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The query kept aborting with "Recursion limit of 60 reached". That
message names LangGraph's own recursion_limit, which is what the bridge
sets internally - so that is the key that was sent, on the session and on
every query, and it changed nothing. The bridge's API takes the budget as
max_steps, on the query operation, which is worth more than any amount of
reasoning about its error text.
Sent there and nowhere else now: the budget belongs to a query, not to
the session that may run several.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The row offered both ways of giving a file at once and greyed out the one
not in use, which still leaves the reader working out which half is live.
It reads better as what it is: one choice, then the fields that choice
needs - an input name, or which global input holds the file.
Greying was all the schema could express, so this adds the annotation for
showing a field only in the state it belongs to. The distinction earns
the second annotation: a field that still tells the reader something
while unavailable should stay and grey, but the branch nobody picked is
not unavailable, it is irrelevant.
A row saved before the choice existed says which branch it is on by what
it carries, so it is stamped on read rather than left reading as the
default and demanding an input name it never had.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Picking a global and nothing else still failed the block update: the kind
select posts "" when nobody chose from it, and an enum cannot be coerced
from that, so the request came back naming a field that is optional and
that the person had deliberately left alone. Blank now reads as unset.
The row also left both sources on offer at once, with nothing saying which
one wins. Filling either now greys out the other, and the input name is
required exactly when no global is chosen - which needed a way to say
"while this other field is empty", since a primitive boolean cannot tell
an unspecified present() from a false one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>