Commit Graph

3 Commits

Author SHA1 Message Date
Lucio Lelii 074c8fb763 Give the model parameters the same treatment a node's own panel gives them
The picker for a simulator or a judge and a node's parameter panel are the
same dialog component, and the dialog has always known how to offer "Use
default" on an optional field - showUseDefault, the "Using the default" hint,
the reset. It reads one flag, defaultsWhenEmpty, which the node panels set on
every optional field and this hand-written list never set at all. So a
temperature typed here by mistake had no way back to unset: clearing the box
by hand looks the same as never having decided.

Set it on all five, and brought the rest of each field in line with what the
schema-driven panel produces for the same object: an arrow step a decimal can
actually move by, an integer step on the integers, and the tips ModelParameters
itself declares, so the same explanation appears in both places.

Temperature's max was 2 here and is 1.0 on the server, which has a comment
explaining why - the dialog was offering a value the run would be rejected
for.

Still hand-written rather than derived from the published schema: this dialog
picks a model for a run, not a node's configuration, and reaching for a block
type's schema to render five known fields would buy a network call and a way
to fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 15:15:10 +02:00
Lucio Lelii 03ac4901ed Put the provider/model picker above the dialog that opens it
The picker (node-settings-dialog) shared --z-modal with every other dialog
guest, so opening it from "Evaluate impact with LLM" - itself a modal - tied
on z-index with the report behind it and lost on DOM order: it rendered, and
its backdrop even blocked clicks, but neither was visible. It read as a
button that did nothing.

Named the layer this actually is - --z-dialog-over-modal, the same one the
confirmation dialog already needed and had defined ad hoc as --z-confirm -
and moved the picker onto it.

While chasing this, closed a real silence next to it: with no LLM provider
published at all, the picker answered null and the caller treated that like
a dismissal, so the button did nothing for a second, unrelated reason. It now
throws with a message, and Simulate surfaces it as a notification instead of
swallowing it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 11:41:41 +02:00
Lucio Lelii f5386cdb7e Say which subject a bias intervention moved, and let an LLM assess it
The report now reads the per-subject sections the service produces: a row per
iterated subject with what its labelled numbers did (Score 7 -> 4), the two
texts behind it a click away, and the changed ones listed first. The band at
the top leads with what the run actually did - whether the final decision
changed, how many subjects moved, the largest delta - because the counts of
changed nodes that used to open the report were the least actionable thing in
it. A container's iterations are listed with the inner node that changed,
which is what a per-subject iterator run needs and the accumulated list could
never show. Reports produced before any of this exists still render, from
their raw outputs.

"Evaluate impact with LLM" sits next to the report it is about, in all three
places one is mounted, and opens the provider and model picker the interaction
simulator uses - extracted so both call the same dialog rather than two of
their own, with temperature 0 offered by default because a verdict that reads
differently every time it is asked for is worse than none. The assessment runs
as a job, polled like the isolated experiment, and is stored on the report, so
reopening it later shows the same verdicts and the model that produced them.
It is labelled an assessment throughout, and a pair the model could not answer
for is marked without hiding that pair's own figures.

A rerun of a simulated run now opens the Simulate dialog on the simulator it
inherited, with the inherited sampling out where it can be seen - a seed
carried over is the reason the two runs are comparable, and behind a closed
section nobody would find it. Before it is started, a run says which simulator
the run it repeats used; afterwards, both the bias report and the run-to-run
comparison say so when the two sides were not answered the same way, since
that difference is not the intervention's doing and nothing said it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 10:50:54 +02:00