A report held one BiasJudgeSummary, so asking a second model overwrote the
first: there was no way to compare two opinions on the same comparison, and
a report survived only until someone reevaluated it. BiasImpactReport now
keeps a capped history (10) of assessments, newest first, each carrying its
own verdicts per compared pair rather than a single shared set - the point
of keeping several is being able to trust each one's own reasoning, not just
its headline. Recomputing an outdated report now carries the history forward
instead of discarding it.
schemaVersion had to become Integer rather than int in the same change:
Jackson reads an absent field as null, and a primitive rejected every report
persisted before today - a bug this history change would otherwise have
inherited silently, caught by a new compatibility test that reads an old
report's JSON.
Separately, the judge asked for its verdict in JSON mode only, and Ollama's
format=json makes some models - reasoning models especially - answer with an
empty body or a degenerate {}, since they have nowhere to put their thinking
under that flag. Every compared pair failed with "the model returned an
empty answer". The flow assistant has ripped this exact seam out before and
falls back to a text call; the judge now does the same, extracting the
verdict object from whatever prose it comes wrapped in.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|---|---|---|
| .mvn/wrapper | ||
| src | ||
| .gitattributes | ||
| .gitignore | ||
| Dockerfile | ||
| mvnw | ||
| mvnw.cmd | ||
| ollama-preloaded-docker | ||
| pom.xml | ||
| run_service.sh | ||