In July I wrote The Model Sheet — how a filename like Qwen2.5-Coder-32B-Instruct-Q5_K_M.gguf breaks into family, domain, params, training stage, quantization, and format. Those fields predict behavior before you download. What comes after the spec sheet is the trial: load the weights, run a real prompt, see if the model behaves as advertised.
Lately that trial has been Thinking Machines' Inkling: 975B total / 41B active MoE, Apache 2.0, 1M context. Most modern open-weight releases now ship with the OpenAI messages format (system / user / assistant roles with content arrays). That is the only reason the claude-inkling wrapper works. No translation layer needed — just point ANTHROPIC_BASE_URL at the endpoint.
The Baseten route (and the rate limit that killed it)
Baseten's direct endpoint (https://inference.baseten.co) has a native /v1/messages route. For a verified workspace, it skips the Basic-tier rate cap (100k tokens/min) — a single Claude Code turn exhausts that instantly. The key lives at ~/.config/baseten/key. The wrapper (~/.local/bin/claude-inkling) reads it, exports the endpoint and model identifier (thinkingmachines/inkling or inkling-small), then executes claude.
In early trials we hit the cap immediately. Reading files with Inkling Large consumed tokens fast enough that the rate limit interrupted work before anything finished. The fix was temporary: switch back to the Vercel AI Gateway (claude-gateway) to complete the task. The Gateway pools through a single bill with provider failover, so the interruption was minimal. Once the workspace is verified or a dedicated deployment is provisioned, Baseten direct is the stable path. For now the wrapper supports both: INKLING_VIA=baseten for direct, default for Gateway.
The alias: claude-inkling
The wrapper is a thin bash script. It exports Anthropic endpoint variables, disables telemetry and non-essential model calls (so local decode isn't wasted on background tasks), sets a generous output budget (32k by default), and executes claude. Same pattern as claude-laguna (local llama-server) and claude-deepseek (local DeepSeek V4-Flash via ds4). One alias per model, one endpoint per alias, zero proxy overhead.
Because Inkling — like most new releases — speaks the OpenAI messages format, the client sends {"model":"thinkingmachines/inkling","messages":[...]} and gets back {"type":"message","content":"..."}. The alias just bridges Anthropic's CLI naming convention and the provider's model registry.
What the trial looks like
Sampling is not a benchmark run. It is a short, representative task from real work — a coding snippet, a reasoning chain, a format conversion — scored pass or fail against a known answer. The model either produces working output or it doesn't. Binary. I keep the prompt set in a plain file under version control. The next release has to earn the swap.
Inkling's behavior so far aligns with its MoE spec: 41B active out of 975B total. Fast prefill, slower decode (bandwidth-bound, like any large model), and a generous output budget needed for reasoning spans before visible answers appear. That matches other reasoning-model releases — if max_tokens is too small, the budget goes to hidden thinking and the content field returns empty. Fix: large token budget, or a non-reasoning variant for short replies.
This connects back to The Benchmark: a leaderboard rank collapses too many dimensions into one number. The fields from The Model Sheet — dense vs MoE, active params, quantization level, training stage — are the actual variables. The alias wrapper just makes them accessible to the CLI I already use.
Where this fits: the local lab keeps Llama-3.3-70B (storyteller) and Qwen2.5-Coder-32B (architect) hot on Metal via llama-server (llama70 :8080, coder32 :8081). The claude-inkling alias is the hosted complement: frontier reasoning or very long context goes through the Gateway; once verified, Baseten direct takes over.
Next: finish the Baseten workspace verification so direct routing replaces the Gateway fallback, and update the prompt set file with the Inkling trial results. See the running notes in ~/llm-lab/session-log.md for the live measurements.

Discussion
Giscus