feat(evals): record and replay live model transcripts - #8425
sudoKrishna wants to merge 7 commits into
Conversation
Add a deterministic eval layer for the agent harness. Scenarios script the OpenAI-compatible streaming tool loop with model turns and stub tool results, then score tool selection, planning, retrieval, and recovery without a provider key. - apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report - `bun run test:evals` from apps/sim runs the suite and writes the report - picked up by the normal vitest run so a regression fails CI - README documents the contract and how to add a case
Replay the same scenarios against a real model. The model is the only thing that changes: runScenario now takes an optional completion transport and a live mode that relaxes exact assertions (ordered subsequence, minimum successes) and skips scripted-only recovery cases. - live.ts: OpenAI-compatible transport + DeepSeek factory - agent-tool-use.live.test.ts: K trials per scenario, gated on EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI - live report with pass rates, avg iterations, latency, failed checks - test:evals:live script and README knobs
…ve mode The first live DeepSeek run exposed brittle assertions, not harness bugs: the model chained the tools correctly but the checks were case-sensitive and required an internal order id. Match the retrieved value case-insensitively and let live runs accept the grounded status rather than the internal id.
Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor, with only executeProviderRequest mocked at the provider boundary. This covers agent-block input wiring, variable resolution from Start outputs, and executor run/error handling, which the direct loop harness cannot see. - executor-harness.ts: workflow builder + runExecutorScenario - shares the scorer (scoreExpectations) and report with the loop suite - two scenarios: Start->Agent output, and <start.message> resolution - README documents adding an executor-level scenario
Add executor-retries-failed-block: the first provider call rejects, the Agent block has retry enabled, and the executor replays it. The run must complete with the second response. Verifies providerCalls === 2, and fails without the retry policy (checked locally: expected 2, got 1).
Add executor-falls-back-to-secondary-model: the primary call rejects, the Agent block has a fallback model, and the handler serves the answer from gpt-4o-mini. Asserts providerCalls === 2 and lastRequestModel, and fails without the fallback row (checked locally: got gpt-4o, run errored).
Record a live run once, replay it forever through the real tool loop with no key. EVAL_RECORD=1 wraps the live completion and writes each model call's streamed chunks to fixtures/<scenario>.json; agent-tool-use.replay.test.ts feeds them back through createOpenAICompatStreamingToolLoopStream and scores them with the same checks. - replay.ts: recording/replay completions + fixture I/O - replay.test.ts: chunk round-trip and fixture I/O (key-free) - live test records on EVAL_RECORD=1; test:evals:record script - replay suite skips until a fixture exists; README documents the loop
|
@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel. A member of the Team first needs to authorize it. |
|
| if (RECORD && trial === 0) { | ||
| fixturesToWrite.push({ | ||
| scenarioId: scenario.id, | ||
| model: MODEL, | ||
| recordedAt: new Date().toISOString(), | ||
| turns: recordedTurns, | ||
| }) | ||
| } |
There was a problem hiding this comment.
Failed trials produce fixtures When the first live trial fails during a model stream, the recorder keeps only completed turns, but this code still writes them as a fixture. Because the default pass-rate floor is zero, the record command can succeed while producing an empty or truncated fixture that fails the offline replay suite. Write a fixture only after a complete, passing trial.
| for (const fixture of listReplayFixtures(FIXTURES_DIR)) { | ||
| const scenario = AGENT_TOOL_USE_SCENARIOS.find((candidate) => candidate.id === fixture.scenarioId) | ||
| if (scenario) replayCases.push({ id: fixture.scenarioId, fixture, scenario }) | ||
| } |
There was a problem hiding this comment.
Unmatched fixtures disappear silently If a scenario is renamed or removed, its committed fixture is discarded here without a warning. When no matching fixtures remain, the replay suite skips entirely and CI stays green despite running no recorded cases. Fail on unmatched fixture IDs so lost replay coverage is visible.
| let index = 0 | ||
| return async () => { | ||
| const turn = turns[index] | ||
| index += 1 |
There was a problem hiding this comment.
Stale fixtures can pass Replay returns chunks by turn number without checking the current model request. If a scenario’s prompt or available tools change, an old fixture can still satisfy the output checks, leaving the replay green without testing the changed request. Record and check which request each fixture belongs to.
| import { type EvalRunMode, type ScoredToolCall, scoreExpectations } from './harness' | ||
| import type { | ||
| AgentToolUseExpectations, | ||
| AgentToolUseResult, | ||
| EvalCategory, | ||
| EvalToolInvocation, | ||
| } from './types' |
There was a problem hiding this comment.
Relative imports violate app rules These neighboring-module imports use relative paths, contrary to the Sim app directive: “Always use absolute imports. Never use relative imports.” The same pattern appears in
harness.ts, report.ts, and scenarios.ts. Use the @/evals/agent-tool-use/... alias in these files; this repository requirement must be satisfied before merging.
Context Used: Import patterns for the Sim application (source)
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Summary
Record a live model run once, replay it forever through the real tool loop with
no key and no network. This turns today's nondeterministic live finding into a
durable, deterministic regression test, and unblocks suites (context, subagents)
that need many real model turns.
Stacked on #8409 (the eval harness) — base branch is
feat/agent-tool-use-evals.Closes #8424
What changed
replay.ts—createRecordingCompletion(wraps a live completion, captureseach call's chunks) and
createReplayCompletion(feeds them back), plusfixture read/write/list helpers
agent-tool-use.replay.test.ts— replays every committed fixture throughcreateOpenAICompatStreamingToolLoopStreamwith the same scoring; skips untila fixture exists
replay.test.ts— key-free unit coverage of chunk round-trip and fixture I/Oagent-tool-use.live.test.ts—EVAL_RECORD=1records the first trialtest:evals:recordscript; README documents recording, replaying, refreshingHow to record and replay
Fixtures hold the raw streamed chunks per model call, so a diff shows a behavior
change exactly as the model produced it. They are committed and reviewed like
snapshots.
Test plan
bun run test:evals→ 16/16 (12 harness + 4 replay primitives), replaysuite skips while no fixtures exist
single-tool-lookupreplayed through the real loop and passed), then the synthetic fixture was
removed
replay.tstype-checks against the real loop signaturebun run check:test-patternspassesbun run type-check— run in CIFollow-up
Once real fixtures are committed, the context/memory and subagent suites can
replay real multi-turn transcripts instead of scripting every turn.