A full-flow comparison measured one number: an edit distance between the
outputs of every activated node stringified into one map. On a flow that
evaluates one subject per iteration - the multi-CV case - that number was
dominated by the map's punctuation and could not say which subject moved,
by how much, or whether anything of substance had changed.
The comparison is now field by field, and list values are paired element by
element, so an iterated node reports per subject. Where both sides label a
number the same way, its movement is reported in the flow's own units
(Score 8 -> 5), which is the only figure here a reviewer can act on. Values
that only one side produced are marked rather than guessed at, and both the
element count and the edit distance are capped, the latter because it was
quadratic with no bound on exactly the long model outputs it runs on.
Iterator and loop containers are compared iteration by iteration by joining
the child executions on parentIterationIndex, which they already record.
Guard subflows are excluded: a loop creates one per main iteration, and
pairing a guard with a main run reported every loop as rewritten. This is
what turns "the accumulated list changed" into "iteration 3, on this inner
node".
None of that says whether the change is the intervention doing what its
probe described or the same model answering differently, so a report can now
be handed to a model: one call per aligned pair, answering on two separate
axes - how far the meaning moved, and whether the change carries the
intervention's fingerprint. Provider, model and sampling arrive as an
LLMDescriptor and are resolved the way the interaction simulator's are. The
level and the attribution shown are rolled up from the per-pair verdicts in
code; only the narrative comes from the model, so re-reading a report cannot
show a different headline than the pairs it is made of. A model that answers
with something unusable leaves an error on its pair and nothing else.
Finally, a rerun now carries the simulator of the run it repeats - the
descriptor only, never the enabled flag - and the report records what
answered the interactive steps on each side. Two runs answered by different
simulators differ for a reason the intervention had no part in, and until now
nothing said so.
Reports are persisted and served from storage, so they carry a schema version
and an older one is recomputed instead of serving its new sections empty.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>