RETURN TO BLUEPRINTTDD-019 · THE EVAL HARNESS
EXPERIMENT LOG · AUG 04 202660/60 CELLS · AUDIT VALID

M5 MAX LOCAL INFERENCE · VERCEL AI GATEWAY · CLAUDE CODE 2.1.221

The Coding Benchmark Flight Recorder

Sixty isolated coding-agent runs across twelve real repository tasks. Inspect each patch outcome, the scoring boundary, provider recovery, elapsed time, cache traffic, and exact cloud bill. Codex built the scaffolding and test harness; the matrix then ran unattended through the night, finishing near noon the next day.

RUNS
60
TASKS
12
ROUTES
3
WALL CLOCK
14H 39M
GATEWAY BILL
$34.79
POWER SOURCE
60/60 AC
PLATE 01

Choose the scoring boundary

Strict acceptance requires every held-out command to pass and rejects any tracked-test, lockfile, or Git-metadata change.

22/60 cells pass this lens. Select a route card to focus its marks; select it again to restore all routes.
First-attempt counts include R10, whose passing patch was captured when the process reached the 20-minute deadline.
Local API spend excludes hardware and electricity; neither was measured.
PLATE 02

Every run, in one matrix

first attempt after repair checks pass / integrity reject× fail

Four anchor tasks ran three times per route and show as three cells; the other eight ran once and show as one wide cell spanning the route group. Hover or focus a cell for its run probe; click to pin the full record and a shareable URL. A corner cut marks a process timeout. Task names and every term below are explained in the probe, the pinned record, and the list that follows.

HOW TO READ A CELL · TERMS IN THIS RECORD
Strict acceptance
Every held-out command passed and the patch touched no tracked tests, lockfiles, or Git metadata.
Held-out checks
Every required command passed after target tests were overlaid. Executable correctness only; the integrity gate is not applied.
First attempt
The patch from the first pass, including one captured when the 20-minute deadline forced a stop.
After repair
Passed only after one 10-minute repair pass that received failure labels and output tails.
Integrity reject
Held-out checks passed, but the patch changed tracked tests, lockfiles, or Git metadata, so strict acceptance refused it.
Repair
One optional 10-minute second pass, told to leave existing tests alone, given only when the first attempt finished but failed.
Serving
The cloud host that completed a Gateway request (Fireworks, Alibaba, Baseten, Novita), or the on-device runtime for local runs.
Recovery
A request that succeeded only after a prior provider or credential attempt failed first.
Anchor
A task repeated three times per route to expose run-to-run variance. Four tasks were anchored.
Repetition (R1–R3)
The run index for the same task on the same route. R1 is the run every task gets; R2 and R3 are the extra anchor runs.
All 60 benchmark cells arranged by task, route, and repetition
Task / expected scopeLOCAL DS4GATEWAY DS4GATEWAY SONNET
R1R2R3R1R2R3R1R2R3
Canonical UUID validationctx · 1 expected file
Shared upstream timeoutctx · 3 expected files
Public model catalog parserctx · 1 expected file
Catalog priority and filteringctx · 1 expected file
Callable gateway route slugsctx · 1 expected file
Race-safe credential readsctx · 4 expected files
Fail-closed sensitivity screenctx · 4 expected files
OAuth and connector hardeningctx · 3 expected files
Canonical inquiry labelsportfolio · 2 expected files
Newsletter conversion correctnessportfolio · 2 expected files
Concurrent submit deduplicationportfolio · 1 expected file
Server-decided editorial railportfolio · 6 expected files
R01
Canonical UUID validation

Passed on first attempt

Route
LOCAL DS4
Agent time
2.5 min
Requests
11
Repair
No repair pass
Patch
0.9 KiB · 1 files
Exact bill
$0 API · HW/energy unmeasured
Cache read
77,011 tokens
Serving
On-device runtime
THE TASK

Enforce one canonical UUID form across the ctx public API so mismatched formats stop producing duplicate records.

  • ctx — personal context engine
  • One expected file
  • validation · security
  • Repetition R1
SAME TASK · REPETITION 1Compare the matched route cells
READING THIS RECORD
Strict acceptance
Every held-out command passed and the patch touched no tracked tests, lockfiles, or Git metadata.
Held-out checks
Every required command passed after target tests were overlaid. Executable correctness only; the integrity gate is not applied.
Integrity reject
Held-out checks passed, but the patch changed tracked tests, lockfiles, or Git metadata, so strict acceptance refused it.
Repair
One optional 10-minute second pass, told to leave existing tests alone, given only when the first attempt finished but failed.
Serving
The cloud host that completed a Gateway request (Fireworks, Alibaba, Baseten, Novita), or the on-device runtime for local runs.
Recovery
A request that succeeded only after a prior provider or credential attempt failed first.
PLATE 03

The local lane spent time instead of API dollars

Each mark is one complete agent run. Local DeepSeek reached an agent deadline in 13/20 cells; Gateway DeepSeek did so once. Medians: 20.0, 8.7, and 12.4 minutes.

05101520253035
LOCAL DS4
20m first-attempt cap
GATEWAY DS4
20m first-attempt cap
GATEWAY SONNET
20m first-attempt cap
PLATE 04

Routing recovery was part of the measured product

The Gateway selected among providers and credentials on every request. Failed attempts remain visible beside the provider that completed the response.

REQUESTED MODELDeepSeek V4 Flash
663 completed requests
fireworks411 final
411 ok · 169 fail · p50 0.9s first byte
alibaba136 final
136 ok · 10 fail · p50 2.5s first byte
baseten106 final
106 ok · 300 fail · p50 0.8s first byte
novita10 final
10 ok · 0 fail · p50 7.3s first byte
323 successful requests recovered after a failed provider attempt · 3 requests failed outright
REQUESTED MODELClaude Sonnet 5
1325 completed requests
6 per-run budget stops · 9 first-byte or stream-idle timeouts · exact bill $33.44
PLATE 05

Exact billing beat frozen list-price math

Each step on the line adds one run’s exact Gateway bill, climbing to $34.79 across all 60 runs; Sonnet accounted for $33.44 of it. Dots mark accepted patches in their route color, hollow marks are failed patches, the ring is the run selected elsewhere on this page, and the dashed line is the final total. Click any point to pin that run and sync the rest of the page. Sonnet billed one-third below the frozen list-price estimate; routed DeepSeek billed 4.7% above it.

Cumulative exact Gateway spendAccepted patch · route colorFailed patchSelected run · click any pointFinal bill $34.79
Cumulative exact Gateway spend across all 60 runsThe cumulative bill climbed from zero to 34.79 dollars; Sonnet accounted for 33.44 of it. Click any point to inspect that run.Cumulative spend ($)$0$10$20$30R01 · LOCAL DS4 · Passed on first attempt · $0.00 cumulativeR02 · GATEWAY DS4 · Failed the selected outcome · $0.05 cumulativeR03 · GATEWAY SONNET · Failed the selected outcome · $1.46 cumulativeR04 · LOCAL DS4 · Passed after repair · $1.46 cumulativeR05 · GATEWAY DS4 · Passed on first attempt · $1.49 cumulativeR06 · GATEWAY SONNET · Failed the selected outcome · $3.22 cumulativeR07 · LOCAL DS4 · Failed the selected outcome · $3.22 cumulativeR08 · GATEWAY DS4 · Passed after repair · $3.33 cumulativeR09 · GATEWAY SONNET · Failed the selected outcome · $3.54 cumulativeR10 · LOCAL DS4 · Passed from the patch captured when the first attempt reached its deadline · $3.54 cumulativeR11 · GATEWAY DS4 · Failed the selected outcome · $3.55 cumulativeR12 · GATEWAY SONNET · Failed the selected outcome · $6.24 cumulativeR13 · GATEWAY DS4 · Passed after repair · $6.29 cumulativeR14 · GATEWAY SONNET · Failed the selected outcome · $6.81 cumulativeR15 · LOCAL DS4 · Failed the selected outcome · $6.81 cumulativeR16 · GATEWAY DS4 · Failed the selected outcome · $6.85 cumulativeR17 · GATEWAY SONNET · Failed the selected outcome · $7.51 cumulativeR18 · LOCAL DS4 · Failed the selected outcome · $7.51 cumulativeR19 · GATEWAY DS4 · Failed the selected outcome · $7.60 cumulativeR20 · GATEWAY SONNET · Failed the selected outcome · $10.71 cumulativeR21 · LOCAL DS4 · Passed on first attempt · $10.71 cumulativeR22 · GATEWAY DS4 · Passed on first attempt · $10.77 cumulativeR23 · GATEWAY SONNET · Passed after repair · $13.44 cumulativeR24 · LOCAL DS4 · Failed the selected outcome · $13.44 cumulativeR25 · GATEWAY SONNET · Passed on first attempt · $13.50 cumulativeR26 · LOCAL DS4 · Passed after repair · $13.50 cumulativeR27 · GATEWAY DS4 · Failed the selected outcome · $13.57 cumulativeR28 · GATEWAY SONNET · Failed the selected outcome · $14.15 cumulativeR29 · LOCAL DS4 · Failed the selected outcome · $14.15 cumulativeR30 · GATEWAY DS4 · Failed the selected outcome · $14.24 cumulativeR31 · GATEWAY SONNET · Failed the selected outcome · $17.13 cumulativeR32 · LOCAL DS4 · Failed the selected outcome · $17.13 cumulativeR33 · GATEWAY DS4 · Passed on first attempt · $17.18 cumulativeR34 · GATEWAY SONNET · Passed on first attempt · $19.72 cumulativeR35 · LOCAL DS4 · Failed the selected outcome · $19.72 cumulativeR36 · GATEWAY DS4 · Failed the selected outcome · $19.82 cumulativeR37 · GATEWAY DS4 · Passed on first attempt · $19.83 cumulativeR38 · GATEWAY SONNET · Failed the selected outcome · $21.34 cumulativeR39 · LOCAL DS4 · Failed the selected outcome · $21.34 cumulativeR40 · GATEWAY DS4 · Passed after repair · $21.43 cumulativeR41 · GATEWAY SONNET · Failed the selected outcome · $22.57 cumulativeR42 · LOCAL DS4 · Failed the selected outcome · $22.57 cumulativeR43 · GATEWAY DS4 · Failed the selected outcome · $22.68 cumulativeR44 · GATEWAY SONNET · Passed after repair · $25.36 cumulativeR45 · LOCAL DS4 · Passed on first attempt · $25.36 cumulativeR46 · GATEWAY DS4 · Passed after repair · $25.43 cumulativeR47 · GATEWAY SONNET · Failed the selected outcome · $28.11 cumulativeR48 · LOCAL DS4 · Failed the selected outcome · $28.11 cumulativeR49 · GATEWAY SONNET · Failed the selected outcome · $28.72 cumulativeR50 · LOCAL DS4 · Failed the selected outcome · $28.72 cumulativeR51 · GATEWAY DS4 · Failed the selected outcome · $28.80 cumulativeR52 · GATEWAY SONNET · Passed after repair · $30.61 cumulativeR53 · LOCAL DS4 · Passed on first attempt · $30.61 cumulativeR54 · GATEWAY DS4 · Failed the selected outcome · $30.70 cumulativeR55 · GATEWAY SONNET · Failed the selected outcome · $33.56 cumulativeR56 · LOCAL DS4 · Passed on first attempt · $33.56 cumulativeR57 · GATEWAY DS4 · Failed the selected outcome · $33.62 cumulativeR58 · GATEWAY SONNET · Failed the selected outcome · $34.70 cumulativeR59 · LOCAL DS4 · Failed the selected outcome · $34.70 cumulativeR60 · GATEWAY DS4 · Passed after repair · $34.79 cumulativeR01: selected run · $0.00 cumulative$34.79 final · Sonnet $33.44 (96%)RUN 01RUN 60Run sequence (1 → 60)
Gateway DS4: $1.29 estimated → $1.35 billedGateway Sonnet: $50.16 estimated → $33.44 billed
PLATE 06

Patch footprint exposed the suite’s hard boundary

All three routes fell to 3/7 held-out-check passes on tasks expecting four or more changed files. The repeated sensitivity screen failed on every route; the broad editorial rail failed in all three single trials.

One expected file9 cells per route
L
5/9
D
5/9
S
4/9
Two–three expected files4 cells per route
L
3/4
D
3/4
S
1/4
Four or more expected files7 cells per route
L
0/7
D
1/7
S
0/7
PLATE 07

Price the local machine with your assumptions

The cloud line uses measured API spend per run. The local line combines measured mean agent time with your hardware allocation and power scenario.

Local scenario / run$0.214$0.208 allocated hardware + $0.005 assumed electricity
Measured cloud API / run$0.068GATEWAY DS4 · exact Gateway metadata

671 tasks per month amortize the allocated hardware cost against this cloud route under the entered assumptions.

Defaults are illustrative, not measured or vendor-specified. Power was not measured during the matrix. Hardware, allocation, lifespan, power, and electricity fields are reader-supplied scenarios.
PLATE 08

Instrument record

Machine
Apple M5 Max · 128 GiB unified memory
Local weights
80.76 GiB mixed quant · SHA-256 pinned
Agent shell
Claude Code 2.1.221 pinned for all runs · preflight snapshot recorded 2.1.220
Harness build
Codex assembled the fixture builder, grader, scheduler, and publication pipeline, separate from the Claude Code agent under test
Run cadence
Launched unattended and ran overnight, 01:30Z to 16:09Z, 14h 39m, all 60 cells
Local runtime
Anthropic-compatible relay in front of the resident weights; runtime pinned by commit
Publication
Raw traces stay private; a deterministic script emits sanitized per-run and aggregate artifacts
Context
Local ceiling pinned at 131,072 tokens · hosted ceilings provider-controlled
Tasks
12 historical repository changes · hidden target tests
Budget
20m first attempt · 10m repair · one repair maximum
Telemetry
Requests, timing, cache, provider attempts, grades, patches, power, billing
Local token usage
Fresh prompt tokens reported as cache writes · plain input field remains zero
Statistics
Wilson intervals treat cells as independent · three pairwise p-values are unadjusted
Power control
60/60 endpoint samples on AC · CPU and thermal state were not isolated
Energy limit
No joule measurement; local energy remains a reader-supplied scenario
Audit
60/60 run records reconciled with 60 relay logs