Commit Graph

26 Commits

Author SHA1 Message Date
Lucio Lelii 90a04dc480 Let a value be copied from its full view and from the outcome
The full view of an output or input gains a Copy button - Copy JSON for
a JSON value - and the outcome payload a copy button in its corner, which
stays put while a long payload scrolls. JSON is copied indented, whether
it came as an object or as a model's text in a fence, so it pastes as
clean JSON elsewhere; text is copied as shown.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 14:32:19 +02:00
Lucio Lelii 6134590bd8 Keep the JSON full view still while its tree is opened and closed
Sized to its content and centred on the page, the modal grew and shrank
with every node opened or closed, its top edge jumping each time. Showing
a JSON tree it now has a fixed height, and only its content scrolls.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 14:29:31 +02:00
Lucio Lelii 3ffaf6e596 Show JSON values as trees, and keep a read-only graph's connections unselectable
Outputs, intermediate inputs, loop rounds and outcome payloads that are
JSON - objects, arrays, or JSON a model wrote as text, fenced or not -
are shown with the JSON tree instead of their text. A preview opens the
first level and a few entries, summing up the rest; the full view opens
the whole tree. The JSON viewer gains the depth and entry limits this
takes.

In a read-only graph, such as an execution's, connections can no longer
be picked or shown as selected: there is nothing to do with one there.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 14:14:20 +02:00
Lucio Lelii 022df1909a Add a Rounds tab showing each loop of an execution round by round
For an execution with a loop, the side panel gains a Rounds tab: each
loop, which round it is on out of how many, and every round with what
its steps were given - the draft sent for review, the verdict handed on
- and how it ended: went round again, left the loop, or stopped at the
limit. Earlier rounds come from the execution's step history, the last
from its steps as they stand.

From the second round of a loop on, a step's inputs are shown as they
are now rather than as prepared at the start, which is only the first
round's value. This also corrects the Intermediate tab and the inputs
shown on a node in a loop.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 14:03:28 +02:00
Lucio Lelii ed4dae1a39 Let the log be read backwards while it is still running
The panel scrolled itself to the newest line on every poll, so a reader
who scrolled up to follow something got yanked back five seconds later -
exactly when a log is worth reading, and exactly the moment it became
unusable.

Following now stops as soon as the reader leaves the end, and resumes on
its own when they come back to it, so the common case needs no button.
The button is there to say which of the two is happening and to override
it: a log that stops moving on its own otherwise looks like a log that
stopped.

The threshold lives in the utils with the rule, not in the component: not
zero, because a list that grew by a line between the scroll and the
handler would otherwise read as the reader having walked away.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 11:55:51 +02:00
Lucio Lelii 8c3d9c6ed3 Let an event's own numbers be opened from the log
The log showed each event's sentence and threw away everything behind it.
A pruning warning saying "6332 characters over budget" could not tell you
what the total actually was, a tool call could not say which tool: those
numbers reached the browser and were dropped, leaving the API as the only
way to read them.

Rows that carry details now offer an eye, which opens them read-only.
Rows that carry none offer nothing - a button opening an empty box is
worse than no button - so the decision of what is worth opening lives in
the view model, where it can be tested, rather than in the template.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 11:07:38 +02:00
Lucio Lelii 5787a6fe44 Add the SPDX licence header to every source file
Mechanical: three comment lines at the top of every .ts, .html and .css
under src, and nothing else. Split out from the licence commit so the
files that carry an actual change stay readable in the history, and kept
to its own commit because it moves the blame line on 390 files.

The short SPDX form rather than the full GNU notice - it is
machine-readable under REUSE, it satisfies the requirement to keep the
licence notice intact, and it points at LICENSE-ADDENDUM instead of
restating the attribution term in every file.

The template and stylesheet headers do not reach the bundle: Angular
discards template comments and the production build strips CSS ones. The
one in index.html survives, since that file is served as written, which is
no loss.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 12:10:18 +02:00
Lucio Lelii f5386cdb7e Say which subject a bias intervention moved, and let an LLM assess it
The report now reads the per-subject sections the service produces: a row per
iterated subject with what its labelled numbers did (Score 7 -> 4), the two
texts behind it a click away, and the changed ones listed first. The band at
the top leads with what the run actually did - whether the final decision
changed, how many subjects moved, the largest delta - because the counts of
changed nodes that used to open the report were the least actionable thing in
it. A container's iterations are listed with the inner node that changed,
which is what a per-subject iterator run needs and the accumulated list could
never show. Reports produced before any of this exists still render, from
their raw outputs.

"Evaluate impact with LLM" sits next to the report it is about, in all three
places one is mounted, and opens the provider and model picker the interaction
simulator uses - extracted so both call the same dialog rather than two of
their own, with temperature 0 offered by default because a verdict that reads
differently every time it is asked for is worse than none. The assessment runs
as a job, polled like the isolated experiment, and is stored on the report, so
reopening it later shows the same verdicts and the model that produced them.
It is labelled an assessment throughout, and a pair the model could not answer
for is marked without hiding that pair's own figures.

A rerun of a simulated run now opens the Simulate dialog on the simulator it
inherited, with the inherited sampling out where it can be seen - a seed
carried over is the reason the two runs are comparable, and behind a closed
section nobody would find it. Before it is started, a run says which simulator
the run it repeats used; afterwards, both the bias report and the run-to-run
comparison say so when the two sides were not answered the same way, since
that difference is not the intervention's doing and nothing said it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 10:50:54 +02:00
Lucio Lelii bf4ed026b2 Tell the two run buttons apart by who answers the human steps
A green play and a flask said nothing about the one thing that separates them:
whether you answer the flow's chat and decision steps, or an LLM answers them for
you. The flask was actively misleading - "experiment" is the bias experiments'
word, so it pointed at a different feature.

Both buttons now keep the play, and a small badge carries the difference: a
person, or a robot. A gear was the first thought and is the wrong glyph - it is
the settings icon everywhere else in the product, so on a run button it reads as
"configure this run" rather than "a machine runs it".

They also sit next to each other now. Stop and resume used to separate them, so
each had to be understood alone - which is precisely the job a 17px badge cannot
do. Side by side they are one choice, and each is what makes the other legible.

The tooltips stop repeating the button and say who answers: "Run - you answer the
human steps", "Run simulated - an LLM answers the human steps". That is the
sentence the icon can only gesture at.

557 frontend tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-04 09:26:49 +02:00
Lucio Lelii ad8a65fb18 Make a bias variant recognisable, and quieten the run history
Nothing in the run history said whether a run was a bias variant. Worse, the
viewer header claimed "Bias variant" for *every* execution: it tested for the
presence of biasExecutionContext, which the backend sets on all of them,
defaulting to NORMAL. Only the mode distinguishes them. The same faulty test also
offered the bias comparison action on any plain rerun.

Each history row now carries a kind badge - Run, Rerun, Bias, Mitigation, or Bias
+ mitigation - coloured and with a tooltip saying what it means for the result. A
bias variant reads as a variant even when it is also a rerun, because carrying
probes is what changes how its output should be read.

The rows also stop showing raw uuids. A rerun names the run it came from by
number ("from #1") instead of repeating a 36-character id, the execution id is
gone from the row entirely - the viewer header owns it - and the group header no
longer prints the source flow id. "1 runs" reads "1 run".

The viewer header keeps only what changes how a result should be read: Simulated,
the bias variant, Subflow. Execution id, simulator descriptor, experiment id,
baseline and probe internals moved behind a Details toggle, closed by default.
Two Italian strings in that block are now English, like the rest of the app.

Cards are flat here too, matching the flows list: no gradients, no lift on hover,
a left accent for the selected run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-03 11:34:36 +02:00
Lucio Lelii f086064e61 Make the outcomes section collapsible
An outcome payload is often a generated document - the local example is a full
rejection email - which pushed the execution graph down the page. The section now
collapses.

It starts open, because this is the flow's answer and it was invisible until a
moment ago. Collapsed, the header keeps the outcome codes, so the conclusion is
never hidden entirely: only the payload folds away.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-03 11:29:35 +02:00
Lucio Lelii 0dcdc6d135 Show a run's outcomes, where an End node puts the flow's answer
A flow result is built only from *unconnected* outputs, so wiring the last block
into an End node moves its value out of the result and into the outcome payload -
which nothing in the UI read. The run then looked like it produced nothing at all.
There is already an execution in a local database whose outcome payload is a full
generated rejection email that was invisible for exactly this reason.

The run view now has an Outcomes section above the graph, listing each End the run
passed through: its code, its label, the step it came from, and its payload. An
End reached with no value says so, rather than showing an empty box - the two
cases mean different things.

A text payload renders as text rather than through the JSON tree. The tree would
have kept the line breaks but wraps strings in quotes, and the common case here is
a generated document, not a data structure.

This is the smallest of the options for the underlying gap: the data was already
persisted and already in the API payload, so nothing in the engine changed. The
gap itself remains - End cannot be attached as a pure ordering dependency
(canDependOnOtherNodes is false), so a single-exit flow still has to choose
between a labelled end and a value in `result`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-03 11:21:33 +02:00
Lucio Lelii abf21b1c24 Add fullscreen mode, wider zoom-out, and draggable execution graph positions
- Lower rete editor min zoom (0.35 -> 0.1) so wide flows fit on screen.
- Remove the Export action from the read-only subflow preview dialog,
  since it's only ever opened from the execution view.
- Add a fullscreen toggle to the flow editor and task execution viewer,
  placed next to the title; disables the "flows" sidebar tab and hides
  the right assistant/validation panel while the flow editor is
  fullscreen, and collapses the left sidebar if it was open.
- Allow dragging nodes in the read-only execution graph and persist
  custom positions client-side in sessionStorage, keyed by execution id,
  restored on reload/refresh.
2026-09-02 20:00:34 +02:00
Lucio Lelii 4d92cbc51c feat(execution): refine tree and subflow layouts 2026-08-03 17:21:52 +02:00
Lucio Lelii 8faee0fbaf feat: support interactive container subflows 2026-07-24 07:22:38 +02:00
Lucio Lelii 94ce814a6f feat(bias-impact): complete side-effect policy, compare, report list and canvas highlighting
Finishes the bias impact experiments plan (docs/bias-impact-experiments-plan.md
steps 7-12):

- Side-effect policy selector: reuse the .llm-warning visual language for the
  external side-effects banner, add a REQUIRE_CONFIRMATION note.
- Full-flow compare ("Compare with baseline"): new bias-compare-dialog
  (service + host) triggering compareBiasExecutions and opening the shared
  report viewer; inline errors read from errors[].message/detail.
- Persisted reports: new bias-impact-report-list (list + detail in one view)
  wired into a new "Bias impact reports" tab in task-execution-viewer;
  404/403 on report detail show the same inline message on purpose.
- Canvas: annotation badge on generic-node (count, executable-probe
  indicator, severity from the backend catalog); new
  BiasComparisonViewStateService driving bias-active / downstream-changed /
  routing-change highlighting on task-step-node and custom-connection, fed by
  a highlightOnCanvas event from bias-impact-report-viewer wired in all three
  places that render it; legend + "back to normal view" action in the canvas
  toolbar.
- Fixed a bug where the bias variant context badge only rendered for
  simulated executions.
- Fixed "Measure bias impact" to stay visible-but-disabled with an
  explanatory tooltip while the baseline hasn't reached a final state,
  instead of being hidden outright, per the §12 checklist.
- Added a Retry action to the compare dialog's error state for parity with
  the report list.
- Added the end-to-end facade flow test (annotation -> capability -> isolated
  experiment -> report) plus coverage for all new components/services.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-21 11:20:57 +02:00
Lucio Lelii 7bfef72f79 feat: add bias impact experiments and biased reruns 2026-07-21 10:34:22 +02:00
Lucio Lelii 877844fca0 feat(execution): add intermediate inputs tab and preview modal 2026-05-13 12:18:26 +02:00
Lucio Lelii 52bb12edae Reposition task execution controls into graph action rail 2026-04-24 12:18:48 +02:00
Lucio Lelii d508ccf6d1 Refine container rendering and assistant runtime config 2026-03-27 17:39:59 +01:00
Lucio Lelii a2a4f699c7 look and feel modified in node outputs 2026-03-24 14:48:21 +01:00
Lucio Lelii 3bf669180d look and feel in the longtext parmeter changed 2026-03-24 14:05:21 +01:00
Lucio Lelii 55418d2226 Refine execution viewer and schema-driven node handling 2026-03-24 13:45:32 +01:00
Lucio Lelii dc45a5581b Improve execution input and output presentation 2026-03-11 17:11:37 +01:00
Lucio Lelii 36167f13ac added task step part and make nodes generic 2026-03-04 16:55:14 +01:00
Lucio Lelii 4504ec0087 added task layout 2026-03-03 20:01:03 +01:00