feat(evals): add LLM-as-judge scoring - #8515
sudoKrishna wants to merge 7 commits into
Conversation
Add a deterministic eval layer for the agent harness. Scenarios script the OpenAI-compatible streaming tool loop with model turns and stub tool results, then score tool selection, planning, retrieval, and recovery without a provider key. - apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report - `bun run test:evals` from apps/sim runs the suite and writes the report - picked up by the normal vitest run so a regression fails CI - README documents the contract and how to add a case
Replay the same scenarios against a real model. The model is the only thing that changes: runScenario now takes an optional completion transport and a live mode that relaxes exact assertions (ordered subsequence, minimum successes) and skips scripted-only recovery cases. - live.ts: OpenAI-compatible transport + DeepSeek factory - agent-tool-use.live.test.ts: K trials per scenario, gated on EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI - live report with pass rates, avg iterations, latency, failed checks - test:evals:live script and README knobs
…ve mode The first live DeepSeek run exposed brittle assertions, not harness bugs: the model chained the tools correctly but the checks were case-sensitive and required an internal order id. Match the retrieved value case-insensitively and let live runs accept the grounded status rather than the internal id.
Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor, with only executeProviderRequest mocked at the provider boundary. This covers agent-block input wiring, variable resolution from Start outputs, and executor run/error handling, which the direct loop harness cannot see. - executor-harness.ts: workflow builder + runExecutorScenario - shares the scorer (scoreExpectations) and report with the loop suite - two scenarios: Start->Agent output, and <start.message> resolution - README documents adding an executor-level scenario
Add executor-retries-failed-block: the first provider call rejects, the Agent block has retry enabled, and the executor replays it. The run must complete with the second response. Verifies providerCalls === 2, and fails without the retry policy (checked locally: expected 2, got 1).
Add executor-falls-back-to-secondary-model: the primary call rejects, the Agent block has a fallback model, and the handler serves the answer from gpt-4o-mini. Asserts providerCalls === 2 and lastRequestModel, and fails without the fallback row (checked locally: got gpt-4o, run errored).
Substring checks measure phrasing, not correctness. judgeAnswer scores an answer against a weighted rubric with a judge model and returns structured scores; runScenario gains an optional judge that adds a judge check. The judge transport is an injectable OpenAI-compatible completion, so a recorded transcript can replay it deterministically. - judge.ts: rubric, prompt, JSON parsing/clamping, verdict - judge.test.ts: parsing/weighting/clamping (key-free) - judge.live.test.ts: grounded answer outscores an invented one (opt-in) - test:evals:judge script; README documents it
|
@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel. A member of the Team first needs to authorize it. |
|
| rubric: options.judge.rubric, | ||
| userMessage: scenario.userMessage, | ||
| answer: finalContent, | ||
| evidence: JSON.stringify(invocations), |
There was a problem hiding this comment.
Judge lacks tool results When a scenario enables the judge, it receives tool names, arguments, and success flags, but not the returned outputs. If the answer says the rate limit is 100 requests per minute, the judge cannot see the retrieved rate limit and cannot reliably score whether that claim is grounded. Pass the tool results as evidence.
| function createScriptedCompletion(scenario: AgentToolUseScenario): OpenAICompatCreateCompletion { | ||
| let turnIndex = 0 | ||
| return async () => { | ||
| const turn = scenario.script[turnIndex] | ||
| turnIndex += 1 | ||
| if (!turn) { | ||
| throw new Error( | ||
| `Scenario "${scenario.id}" requested model turn ${turnIndex} but only ${scenario.script.length} are scripted` | ||
| ) | ||
| } | ||
| return (async function* () { | ||
| for (const next of turnToChunks(turn)) yield next | ||
| })() | ||
| } | ||
| } |
There was a problem hiding this comment.
Scripted turns ignore feedback The scripted completion advances by turn number without reading the messages sent to it. If the loop stops feeding back a tool result or error, the script still emits its planned retry or answer, so retrieval and recovery scenarios can pass despite that regression. Check the tool messages at the completion boundary.
| toolsMockFns.mockExecuteTool.mockImplementation( | ||
| async (toolId: string, params: Record<string, unknown>): Promise<ToolResponse> => { | ||
| const startedAt = Date.now() | ||
| const response = resultQueues.get(toolId)?.shift() ?? { success: true, output: {} } |
There was a problem hiding this comment.
Live calls receive fabricated results Live tool calls consume results queued for the scripted calls, without matching the model’s actual arguments. If the model makes an extra retry, the exhausted queue returns a successful empty result. That fabricated result reaches the model, so its answer and the reported pass rate no longer reflect the intended tool evidence. Match results to live calls or fail when a fixture runs out.
| expect: { | ||
| toolCallSequence: ['get_weather', 'get_news'], | ||
| finalContent: /12°C[\s\S]*transit strike ends/i, | ||
| maxIterations: 3, | ||
| successfulToolCalls: 2, | ||
| }, | ||
| /** The two tools are independent; a real model may emit them in either order. */ | ||
| liveExpect: { toolCallSequence: undefined, requiredTools: ['get_weather', 'get_news'] }, |
There was a problem hiding this comment.
Live checks reject paraphrases The live expectations keep this exact weather-and-news answer pattern. A correct answer such as “12 degrees Celsius” with a paraphrased headline fails the check, making the pass rate depend on wording rather than correctness. The weather scenario likewise retains an exact
17°C check; relax or judge these live answers.
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
| const weight = criterion.weight ?? 1 | ||
| scores[criterion.id] = score | ||
| weighted += score * weight | ||
| totalWeight += weight | ||
| } | ||
|
|
||
| const weightedScore = totalWeight === 0 ? 0 : weighted / totalWeight | ||
| return { | ||
| scores, | ||
| rationale: typeof parsed.rationale === 'string' ? parsed.rationale : '', | ||
| weightedScore, | ||
| passed: weightedScore >= (rubric.minScore ?? 0.5), |
There was a problem hiding this comment.
| import { adaptOpenAIChatToolSchema } from '@/providers/tool-schema-adapter' | ||
| import type { ProviderToolConfig, TimeSegment } from '@/providers/types' | ||
| import type { ToolResponse } from '@/tools/types' | ||
| import { type JudgeRubric, judgeAnswer } from './judge' |
There was a problem hiding this comment.
Relative imports violate app convention This module imports
./judge, although the repository requires absolute @/... imports in apps/sim. The same pattern appears in report.ts and executor-harness.ts. These imports need to use the app alias before merging to satisfy that requirement.
Context Used: CLAUDE.md (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Summary
Substring checks measure phrasing, not correctness — every live failure so far
was a valid paraphrase or an over-specific assertion. This adds a judge model
that scores an answer against a weighted rubric and returns structured numbers,
so the suite can assert behavior instead of wording.
Stacked on #8409 (the eval harness) — base branch is
feat/agent-tool-use-evals.Closes #8514
What changed
judge.ts— rubric + prompt,judgeAnswer, JSON parsing/clamping, verdictjudge.test.ts— key-free coverage: parsing, fenced JSON, clamping, weights,missing criteria, non-JSON
judge.live.test.ts— opt-in: a grounded answer outscores an invented oneharness.ts—runScenarioaccepts an optionaljudgeand adds ajudgecheck; deterministic checks are unchanged
test:evals:judgescript; README documents itHow it works
Because the judge transport is an ordinary OpenAI-compatible completion, it can
be recorded and replayed deterministically with the record/replay layer.
Test plan
judge.test.ts→ 7/7 (parsing, weights, clamping, failures)bun run test:evals→ 19/19bun run check:test-patternspassesjudge.tsand the harness change type-check against the real loop signaturebun run test:evals:judgelive run (grounded > invented)bun run type-check— run in CIFollow-up
Wire the judge into selected scenarios (replace brittle
finalContentregexeswhere a rubric is more honest) once the live judge run confirms the prompt.