M5 MAX LOCAL INFERENCE · VERCEL AI GATEWAY · CLAUDE CODE 2.1.221
The Coding Benchmark Flight Recorder
Sixty isolated coding-agent runs across twelve real repository tasks. Inspect each patch outcome, the scoring boundary, provider recovery, elapsed time, cache traffic, and exact cloud bill. Codex built the scaffolding and test harness; the matrix then ran unattended through the night, finishing near noon the next day.
- RUNS
- 60
- TASKS
- 12
- ROUTES
- 3
- WALL CLOCK
- 14H 39M
- GATEWAY BILL
- $34.79
- POWER SOURCE
- 60/60 AC
Choose the scoring boundary
Strict acceptance requires every held-out command to pass and rejects any tracked-test, lockfile, or Git-metadata change.
First-attempt counts include R10, whose passing patch was captured when the process reached the 20-minute deadline.
Local API spend excludes hardware and electricity; neither was measured.
Every run, in one matrix
Four anchor tasks ran three times per route and show as three cells; the other eight ran once and show as one wide cell spanning the route group. Hover or focus a cell for its run probe; click to pin the full record and a shareable URL. A corner cut marks a process timeout. Task names and every term below are explained in the probe, the pinned record, and the list that follows.
- Strict acceptance
- Every held-out command passed and the patch touched no tracked tests, lockfiles, or Git metadata.
- Held-out checks
- Every required command passed after target tests were overlaid. Executable correctness only; the integrity gate is not applied.
- First attempt
- The patch from the first pass, including one captured when the 20-minute deadline forced a stop.
- After repair
- Passed only after one 10-minute repair pass that received failure labels and output tails.
- Integrity reject
- Held-out checks passed, but the patch changed tracked tests, lockfiles, or Git metadata, so strict acceptance refused it.
- Repair
- One optional 10-minute second pass, told to leave existing tests alone, given only when the first attempt finished but failed.
- Serving
- The cloud host that completed a Gateway request (Fireworks, Alibaba, Baseten, Novita), or the on-device runtime for local runs.
- Recovery
- A request that succeeded only after a prior provider or credential attempt failed first.
- Anchor
- A task repeated three times per route to expose run-to-run variance. Four tasks were anchored.
- Repetition (R1–R3)
- The run index for the same task on the same route. R1 is the run every task gets; R2 and R3 are the extra anchor runs.
| Task / expected scope | LOCAL DS4 | GATEWAY DS4 | GATEWAY SONNET | ||||||
|---|---|---|---|---|---|---|---|---|---|
| R1 | R2 | R3 | R1 | R2 | R3 | R1 | R2 | R3 | |
| Canonical UUID validationctx · 1 expected file | |||||||||
| Shared upstream timeoutctx · 3 expected files | |||||||||
| Public model catalog parserctx · 1 expected file | |||||||||
| Catalog priority and filteringctx · 1 expected file | |||||||||
| Callable gateway route slugsctx · 1 expected file | |||||||||
| Race-safe credential readsctx · 4 expected files | |||||||||
| Fail-closed sensitivity screenctx · 4 expected files | |||||||||
| OAuth and connector hardeningctx · 3 expected files | |||||||||
| Canonical inquiry labelsportfolio · 2 expected files | |||||||||
| Newsletter conversion correctnessportfolio · 2 expected files | |||||||||
| Concurrent submit deduplicationportfolio · 1 expected file | |||||||||
| Server-decided editorial railportfolio · 6 expected files | |||||||||
Passed on first attempt
- Route
- LOCAL DS4
- Agent time
- 2.5 min
- Requests
- 11
- Repair
- No repair pass
- Patch
- 0.9 KiB · 1 files
- Exact bill
- $0 API · HW/energy unmeasured
- Cache read
- 77,011 tokens
- Serving
- On-device runtime
Enforce one canonical UUID form across the ctx public API so mismatched formats stop producing duplicate records.
- ctx — personal context engine
- One expected file
- validation · security
- Repetition R1
The local lane spent time instead of API dollars
Each mark is one complete agent run. Local DeepSeek reached an agent deadline in 13/20 cells; Gateway DeepSeek did so once. Medians: 20.0, 8.7, and 12.4 minutes.
Routing recovery was part of the measured product
The Gateway selected among providers and credentials on every request. Failed attempts remain visible beside the provider that completed the response.
Exact billing beat frozen list-price math
Each step on the line adds one run’s exact Gateway bill, climbing to $34.79 across all 60 runs; Sonnet accounted for $33.44 of it. Dots mark accepted patches in their route color, hollow marks are failed patches, the ring is the run selected elsewhere on this page, and the dashed line is the final total. Click any point to pin that run and sync the rest of the page. Sonnet billed one-third below the frozen list-price estimate; routed DeepSeek billed 4.7% above it.
Patch footprint exposed the suite’s hard boundary
All three routes fell to 3/7 held-out-check passes on tasks expecting four or more changed files. The repeated sensitivity screen failed on every route; the broad editorial rail failed in all three single trials.
Price the local machine with your assumptions
The cloud line uses measured API spend per run. The local line combines measured mean agent time with your hardware allocation and power scenario.
671 tasks per month amortize the allocated hardware cost against this cloud route under the entered assumptions.
Defaults are illustrative, not measured or vendor-specified. Power was not measured during the matrix. Hardware, allocation, lifespan, power, and electricity fields are reader-supplied scenarios.Instrument record
- Machine
- Apple M5 Max · 128 GiB unified memory
- Local weights
- 80.76 GiB mixed quant · SHA-256 pinned
- Agent shell
- Claude Code 2.1.221 pinned for all runs · preflight snapshot recorded 2.1.220
- Harness build
- Codex assembled the fixture builder, grader, scheduler, and publication pipeline, separate from the Claude Code agent under test
- Run cadence
- Launched unattended and ran overnight, 01:30Z to 16:09Z, 14h 39m, all 60 cells
- Local runtime
- Anthropic-compatible relay in front of the resident weights; runtime pinned by commit
- Publication
- Raw traces stay private; a deterministic script emits sanitized per-run and aggregate artifacts
- Context
- Local ceiling pinned at 131,072 tokens · hosted ceilings provider-controlled
- Tasks
- 12 historical repository changes · hidden target tests
- Budget
- 20m first attempt · 10m repair · one repair maximum
- Telemetry
- Requests, timing, cache, provider attempts, grades, patches, power, billing
- Local token usage
- Fresh prompt tokens reported as cache writes · plain input field remains zero
- Statistics
- Wilson intervals treat cells as independent · three pairwise p-values are unadjusted
- Power control
- 60/60 endpoint samples on AC · CPU and thermal state were not isolated
- Energy limit
- No joule measurement; local energy remains a reader-supplied scenario
- Audit
- 60/60 run records reconciled with 60 relay logs