Codeficum 42

Measured notes from the workshop and the rack.

Sizing a self-hosted coding LLM: an estimated 200 developers at 100 tokens per second on one 8× B200 server

One server with eight B200 GPUs serves an estimated 200 developers at 100 tokens per second each, by our planning model. It runs Qwen3.6-35B-A3B in FP8 as four copies of two GPUs. That assumes 200,000-token agent prompts, a 30-second pause between prompts and every request reaching the copy that holds its cache. The server then generates about 12,000 tokens per second and receives about 474,000 input tokens per second, of which about 91,000 need computing.

At a 40-tokens-per-second target the same server reaches about 430 developers. With five teams in five time zones, each at high load five hours a day, it serves about 330. Beyond about 200 developers in one time zone, extra developers mostly wait. With 1,000 working continuously on one server, a task takes about 6.5 minutes on average instead of 54 seconds.

These are model estimates, not measurements, and we have not measured them on B200 hardware. We tested the planning model on the one GPU we have measured ourselves, an RTX 5060 Ti 16 GB. It came within about 10 % of aggregate throughput at 8 to 32 streams. It under-predicted single-stream decode by 23 % and predicted first-token time at low load at up to 5.6 times the measured value. That test does not exercise the long-context regime of this post; What the planning model gets wrong has the details. Every input is adjustable in the simulator further down.

What was measured, and what was modelled

Two things in this post are measurements. We counted tokens with the published tokenizers of four Qwen models on 2026-10-05. We compared the planning model with our own RTX 5060 Ti results from 2026-08-12. Model-card scores, configuration files, list prices and rents are published values we read on 2026-10-05. Everything else is the output of a bandwidth-and-compute planning model, run on 2026-10-05 for the workload described next. The delay section adds a finite-source queueing model on top of its output; the simulator does not run that step. From here on, the planning model means that model, and the LLMs it sizes are named.

What workload does the planning model size?

Agentic coding, not chat. Every example in this post assumes the following request shape.

Workload assumed for every table in this post. Planning assumptions, not measured traffic.
ParameterValue
Input per request: repository files, history, tool output200,000 tokens
Share of the input reused from the previous turn85 %
Calls where an early file changed, so no cached prefix is usable5 %
Thinking tokens2,048
Answer tokens, about 65 % of them code3,000
Pause between one developer's prompts30 s
Speculative decoding (multi-token prediction)on
Cache storage per serverGPU memory and 1 TB CPU RAM

Agents resend most of their context on every turn. Input is therefore large but mostly repeated, and the prompt cache matters as much as the GPU.

Why is per-user speed a memory-bandwidth problem?

Per-user speed is set by how fast each decoding step reads weights and KV cache from GPU memory. In the planning model's worked example, those reads take 7.6 ms of the 10 ms step; the arithmetic takes about 0.07 ms.

Each decoding step produces one token per running request, or more when speculative-decoding drafts are accepted. Each step reads the weights of every expert that any request in the batch routes to. It also reads the KV cache and recurrent state of every request. On top comes a fixed overhead per step, plus communication time when a copy spans several GPUs. Two consequences follow.

At 100 tokens per second a step must finish in about 10 ms. At 32 requests per copy the planning model computes a 10.0 ms step; its speculative-decoding gain (+21 %) and its batch-interference penalty (−17 %) roughly cancel.

Which open models does the planning model cover?

Four open-weight LLMs, all of them hybrids in which only some layers hold KV cache. The score column repeats what each publisher reports; we ran none of these benchmarks.

Parameters, layer mix and coding scores as published on each model card and configuration file, read 2026-10-05. Scores come from different harnesses; compare across vendors only as a direction.
LLMParameters and layersPublished coding scoreNotes
Qwen3.6-35B-A3B35B total, 3B active, MoE; 30 Gated DeltaNet and 10 attention layersSWE-bench Verified 70.1 in NVIDIA's comparison; Qwen's own card reports 73.4262,144-token native context; 2.05 GB of KV per 200,000-token context at 8-bit
Qwen3.8-27B27B, dense; 48 Gated DeltaNet and 16 attention layersSWE-bench Pro 61.7, Terminal-Bench 2.1 73.0 (Qwen card)Planning-model estimate: 2 to 2.4 times the GPUs per developer of Qwen3.6-35B-A3B at 100 t/s
Qwen3.8-Flash-Next125B, 6B active, MoE, plus a 51B n-gram embedding and 4B MTP (180B total); 36 Gated DeltaNet and 12 sparse-attention layersSWE-bench Pro 62.5 (Qwen card)About 100 GB of 4-bit weights: one copy fits on one B200, H200 or MI300X, or on two H100s
Nemotron 3.5 Lightning30B total, 3B active, MoE; 23 Mamba-2, 23 MoE and 6 attention layersSWE-bench Verified 51.6 (NVIDIA card)3,072 bytes of KV per token; lowest SWE-bench Verified score in this table

How many developers get 100 tokens per second on one 8× B200 server?

By the planning model, about 200 with Qwen3.6-35B-A3B FP8 as four copies of two GPUs. The other layouts in the table give 45 to 225. Concurrent slots are requests generating at the same moment while each still gets its answer at 100 tokens per second. The target applies to answer tokens; at that load thinking tokens run at about 92 tokens per second, because speculative decoding gains less on prose.

Planning-model estimate, run 2026-10-05: 8× B200, workload above, answers at 100 t/s, cache-aware routing. Developer counts are saturation points, as the following paragraphs explain.
LLM, layoutConcurrent slotsDevelopers with 30 s pauses
Qwen3.6-35B-A3B FP8, 4 copies × 2 GPUs128about 200
Qwen3.6-35B-A3B FP8, 8 copies × 1 GPU144about 225
Qwen3.6-35B-A3B FP8, 2 copies × 4 GPUs68about 105
Qwen3.6-35B-A3B FP8, 1 copy × 8 GPUs34about 55
Qwen3.8-27B 4-bit, 4 copies × 2 GPUs60about 95
Qwen3.8-27B 4-bit, 2 copies × 4 GPUs52about 80
Qwen3.8-Flash-Next 4-bit, 1 copy × 8 GPUs28about 45

Qwen3.6-35B-A3B and Qwen3.8-Flash-Next have two KV heads each. With tensor-parallel attention, a copy spread over four or eight GPUs stores every KV head on several GPUs, which is why those layouts lose slots in the planning model. Engines that run attention data-parallel and the experts expert-parallel avoid that copy, at a communication cost the planning model does not include. Eight single-GPU copies give the most slots, but they split the cache eight ways; the routing section shows what that costs. The worked example uses four copies of two GPUs.

By the planning model, eight AMD Instinct MI355X GPUs serve about 160 developers at 100 tokens per second with Qwen3.6-35B-A3B FP8 as four copies of two GPUs (104 slots), and about 185 as eight single-GPU copies (120 slots). Each MI355X has 288 GB of memory and the same 8 TB/s of memory bandwidth as a B200. The planning model applies its MI300X parameters, none with a published basis: kernels derated by 15 %, 2.5 ms of fixed overhead per step against 2 ms on the B200, and 0.6 ms of link time per GPU doubling against 0.4 ms. With the B200 values the MI355X would match the B200 slot counts. None of the price pages we opened shows an on-demand MI355X price.

The planning model treats the 12 sparse-attention layers of Qwen3.8-Flash-Next as dense attention, so each step reads the KV cache of every stored token. The model card limits each query to 2,048 selected tokens. Read the Qwen3.8-Flash-Next row as a likely underestimate.

Developers served come from Little's law: developers = slots × (task time + pause) ÷ task time. At 100 tokens per second a task takes about 54 seconds, so a 30-second pause adds about 56 % more developers. That count is the saturation point. At about 200 developers the 128 slots are full on average. If developers act independently, more than 128 want to generate 47 % of the time, and requests queue. About 180 developers keep that below 2.2 % (binomial estimate on the same task time and pause).

Does a lower speed target serve more developers?

Yes, but total output changes little: the server generates 11,900 to 13,600 tokens per second whatever the target. A lower target spreads the same output over more people.

Planning-model estimate, run 2026-10-05: Qwen3.6-35B-A3B FP8, 4 copies × 2 GPUs on 8× B200, workload above. Developer counts are saturation points.
Answer speed per userDevelopers (30 s pause)Task timeTokens generated, totalInput computed
100 t/sabout 20054 sabout 12,000 t/sabout 91,000 t/s
80 t/sabout 23066 sabout 12,000 t/sabout 91,000 t/s
60 t/sabout 27587 sabout 11,900 t/sabout 91,000 t/s
40 t/sabout 430131 sabout 13,600 t/sabout 104,000 t/s

At 40 tokens per second the limit is still decoding, at 88 requests per copy. Computing new input then uses 37 % of the planning model's idle prefill capacity of about 277,000 tokens per second.

What delay do developers feel when more of them share one server?

By the planning model, a task on one 8× B200 server takes about 6.5 minutes on average with 1,000 developers, from submit to finished answer. It takes about 2.3 minutes with 400 and 30 minutes with 4,000. These figures hold answers at 100 tokens per second and queue the rest. Each developer works continuously, all in one time zone: submits a task, waits for the answer, reads it for 30 seconds and submits the next.

Planning-model estimate, run 2026-10-05: one 8× B200 server, Qwen3.6-35B-A3B FP8 as 4 copies × 2 GPUs, workload above, every developer in a continuous loop with a 30 s pause, cache-aware routing with one queue per copy. Finite-source queueing model with exponential task and pause times; values are means.
DevelopersHold speed, queue the rest: mean time per taskOf which, queueingServe all, slower: mean time per taskAnswer speed, serve allTasks per developer per hour, hold / serve all
10024 snone24 sabout 245 t/s67 / 67
20055 s3 s55 sabout 100 t/s42.5 / 42.5
4002.3 min1.4 min2.0 minabout 43 t/s21 / 24
1,0006.5 min5.6 min5.7 minabout 29 t/s8.5 / 9.8
4,00030.5 min29.5 min25.7 minabout 29 t/s1.9 / 2.3

Below capacity, answers run faster than the target: with 100 developers a request shares its copy with few others and streams at about 245 tokens per second. Above it, holding speed keeps every answer at 100 tokens per second once it starts, but developers queue first. Serving all admits up to 133 requests per copy, the KV-cache limit, and slows every answer. At 400 and 1,000 developers it finishes slightly more work, 2.6 to 2.7 tasks per second against 2.4, because a larger batch spreads each step's weight reads and fixed overhead over more requests.

Neither policy keeps up with headcount. Tasks per developer per hour fall in inverse proportion to it: 21 at 400, 8.5 at 1,000 and 1.9 at 4,000 when holding speed. Holding speed keeps every context in the cache up to 1,000 developers. At 4,000, three calls in four start cold and throughput drops to 2.15 tasks per second. Serving all fills GPU memory with running requests, so about half of all calls already start cold at 1,000 developers.

Keeping the mean queueing time under 5 seconds takes 2 servers for 400 developers, 5 for 1,000 and 20 for 4,000. At $5.98 per GPU-hour around the clock, the lowest on-demand B200 price we found on 2026-10-05, that is about $70,000, $175,000 and $698,000 a month.

How much faster is a warm cache?

In the planning model, a warm cache cuts first-token time at full load from 9.3 s to 1.4 s for Qwen3.6-35B-A3B FP8. When the 170,000-token prefix is still stored, the server only computes the 30,000 new tokens.

Planning-model estimate, run 2026-10-05: 8× B200 as 4 copies × 2 GPUs, workload above, no queue. Full load is 32 requests per copy for Qwen3.6-35B-A3B FP8 and 15 for Qwen3.8-27B 4-bit.
Where the cached context is foundQwen3.6, 1 per copyQwen3.6, full loadQwen3.8-27B, 1 per copyQwen3.8-27B, full load
GPU memory0.5 s1.4 s1.0 s1.5 s
CPU RAM, reloaded over PCIe0.5 s1.5 s1.1 s1.5 s
Nowhere: all 200,000 tokens recomputed2.9 s9.3 s6.3 s9.4 s

Reloading the prefix from CPU RAM adds 0.04 s to 0.07 s over keeping it in GPU memory. Recomputing it adds 2.4 s to 7.9 s. LMCache, the vLLM KV offloading connector and SGLang HiCache all reload stored KV this way.

The limit is capacity: about 1.8 GB per developer for Qwen3.6-35B-A3B and 5.7 GB for Qwen3.8-27B. Two hundred idle developers fill 0.36 TB or 1.1 TB. Plan 1 TB to 2 TB of CPU RAM per server for the cache tier.

Only the attention layers of these hybrids hold KV cache, which keeps the figures small. To reuse a prefix, the engine also has to store recurrent-state checkpoints for the Gated DeltaNet or Mamba-2 layers. For the CPU-RAM tier, the offload connector must store and reload those checkpoints with the KV cache. At 40 tokens per second the planning model finds about half of the warm prefixes in CPU RAM, so the 430-developer estimate depends on it. Confirm that your engine version supports prefix caching and KV offloading, including recurrent state, for the LLM you pick.

Why does round-robin load balancing break this?

Round robin ignores where a developer's cache lives. With four model copies and a cache per copy, a returning request lands on the copy that served its previous turn one time in four.

Arithmetic on the planning model's per-task input, run 2026-10-05: 8× B200, Qwen3.6-35B-A3B FP8 at each layout's 100 t/s operating point. Warm counts only the copy that served the previous turn; earlier prefixes left on other copies can add partial hits. Caches are not shared between copies. Idle prefill capacity is the planning model's 277,000 t/s for 200,000-token prompts.
RoutingCalls that start warmInput to computeShare of idle prefill capacity
Round robin over 4 copies, cache in GPU memory only24 %about 379,000 t/s137 %: queues grow
Round robin over 8 single-GPU copies, GPU memory only12 %about 481,000 t/s174 %: queues grow
Cache-aware or sticky routing, 4 copies95 %about 91,000 t/s33 %
Cache-aware or sticky routing, 8 copies95 %about 103,000 t/s37 %

Use sticky sessions keyed on developer or repository, or a cache-aware router such as SGLang Model Gateway, llm-d or NVIDIA Dynamo. In this workload the routing choice changes input demand by a factor of about four to five.

How many developers fit when five teams share one server?

About 330 by the planning model, as long as the busy hours fall at different times in UTC. Take five teams of equal size: US West, US East, EU West, EU East and Central Asia. Each works at high load for five hours, from 10:00 to 15:00 local time.

Assumed high-load windows, 10:00 to 15:00 local, converted to UTC with standard (winter) and daylight-saving (summer) offsets.
TeamUTC offset, winter / summerHigh load in UTC, winterHigh load in UTC, summer
US West−8 / −718:00 to 23:0017:00 to 22:00
US East−5 / −415:00 to 20:0014:00 to 19:00
EU West+0 / +110:00 to 15:0009:00 to 14:00
EU East+2 / +308:00 to 13:0007:00 to 12:00
Central Asia+5 / +505:00 to 10:0005:00 to 10:00

In winter at most two teams overlap. In summer, daylight saving time moves the European blocks an hour earlier in UTC. From 09:00 to 10:00 UTC three teams then overlap: Central Asia, EU East and EU West. Size for the summer peak.

At saturation one server carries about 200 developers at high load, at 100 tokens per second each. With three teams at the peak, each team can have 66 developers: 330 on one server, against about 200 if everyone worked in one time zone. With the burst headroom of about 180 developers, each team can have 60, or 300 in all. No team is at high load from 22:00 to 05:00 UTC in summer, or from 23:00 to 05:00 UTC in winter.

Where do the night-time scheduled tasks go?

Into the quiet window, run from one queue. Add scheduled agent tasks worth half of each team's interactive volume, with the same request shape. That is about 35,000 tasks a night across the five teams. Scheduled tasks do not need 100 tokens per second. At 40 tokens per second the server finishes about 9,700 tasks an hour, so the whole batch needs 3.7 hours of the quiet window.

Scheduling by each team's local night breaks this. If every team runs its batch from 00:00 to 05:00 local time, the US West batch lands on the summer three-team peak at 09:00 UTC. Generation demand then reaches 116 % of what the server produces at the 100-tokens-per-second operating point, so interactive users slow down or queue. In winter the worst hour reaches 99 %. Schedule by server load in UTC, not by each team's clock.

What does the same workload cost on a frontier API?

On Claude Sonnet 5.5 list prices read 2026-10-05, the five-team workload costs about $399,000 a month with scheduled tasks and $266,000 without. One 8× B200 server rented around the clock costs about $35,000 a month. The workload is 330 developers at 100 tokens per second, five busy hours a day and 21 working days. Scheduled tasks add half the interactive volume. That is about 1.49 million interactive and 0.74 million scheduled tasks a month.

The list prices are $2 per million input tokens and $10 per million output tokens. Cache writes with a 5-minute lifetime cost $2.50 per million and cache reads $0.20 per million. Anthropic bills thinking as output. For this workload that is about $0.18 per task.

Cost per task on Claude Sonnet 5.5 list prices read 2026-10-05, workload above. Line items are rounded; the unrounded total is $0.179.
Line item per taskCost
Cache read, 170,000 tokens on 95 % of calls$0.032
New tokens written to cache, 30,000 on 95 % of calls$0.071
Cold calls writing the whole 200,000-token prefix, 5 % of calls$0.025
Output, 5,048 tokens including thinking$0.050
Total$0.179
Planning-model estimate, computed 2026-10-05: monthly cost of the five-team workload. Task counts come from the planning model. The B200 rent is the lowest on-demand price per GPU-hour on a provider price page opened that day (RunPod Community Cloud), paid 24/7.
OptionMonthly
Sonnet 5.5 API, interactive tasks onlyabout $266,000
Sonnet 5.5 API, interactive and scheduled tasksabout $399,000
Rent one 8× B200 server at $5.98 per GPU-hourabout $35,000
Rent two such servers, if real capacity is half the planning model'sabout $70,000

If the planning model's capacity holds, the API with scheduled tasks costs about 11 times as much as one rented server. It costs about 6 times as much as two. The rented server runs around the clock, so the scheduled tasks fill hours that would otherwise sit idle. The Message Batches API halves API prices, per its documentation read 2026-10-05. Each batch request is single-shot, without a tool loop. An agent loop, in which each request needs the tool results of the one before, does not run on it as it stands.

The LLMs are not equivalent, though. We have not measured how many turns each needs per finished change, and that ratio decides the comparison; measure it on your own tasks. Self-hosting also adds operations work, hardware risk and capacity you pay for when nobody is working. A hybrid is worth testing: routine edits on a local LLM, hard tasks sent to the API.

Three changes cut the API bill for this workload:

  1. Send fewer new tokens per turn. New-token cache writes are the largest item, $0.071 of $0.179, and large tool outputs resent every turn inflate them.
  2. Cap thinking. Each extra 1,000 thinking tokens per task costs about $22,000 a month.
  3. Keep calls warm. Each 1 % of calls that start cold costs about $8,700 a month.

Which part of a build log costs the most tokens?

The timestamps. The published tokenizers of four Qwen models (two distinct tokenizer files) give every digit its own token. Each of the seven separators takes one more. A 28-character ISO 8601 timestamp therefore costs 28 tokens. On three sample build lines it is 49 % to 80 % of the tokens.

Measured 2026-10-05 with the published Qwen3.6-35B-A3B tokenizer (tokenizers 0.22.1). Each line carries the prefix 2026-01-15T09:00:00.0000000Z. Qwen3.8-27B, Qwen3.8-Flash-Next and Qwen3.5-9B give the same 28 tokens for the timestamp.
Line after the timestampTokens with timestampTokens withoutTimestamp share
make[2]: Entering directory35780 %
[ 42%] Building C object src/CMakeFiles/app.dir/drivers/uart.c.obj482058 %
src/drivers/uart.h:41:20: warning: 'uart_flush' declared 'static' but never defined [-Wunused-function]572949 %

Take a 10,000-line build log whose lines average the 19 tokens of our three samples. By our calculation it holds about 467,000 tokens, of which 280,000, 60 %, are timestamps. That is beyond the 262,144-token native context of Qwen3.6-35B-A3B. Without the timestamps the same log holds about 187,000 tokens and fits. Sent to Claude Sonnet 5.5 as new input, the timestamps alone cost $0.56 per read at $2 per million tokens.

Strip the timestamps before a log reaches an LLM. Keep a short marker where the log went silent for long, so a build that stopped responding stays visible.

What the planning model gets wrong

We have not calibrated the planning model for the regime of this post. We compared it with our own published measurements of an RTX 5060 Ti 16 GB serving Qwen3.5-9B, measured 2026-08-12. For the AWQ 4-bit runs we used the planning model's W4A16 path, which matches that kernel class.

Planning-model prediction against C42 measurements on one RTX 5060 Ti 16 GB, Qwen3.5-9B. Measured 2026-08-12; prediction run 2026-10-05. The simulator has no RTX 5060 Ti preset, so we added one: 16 GB, 448 GB/s, and 47.4, 94.9 and 379.5 TFLOPS dense BF16, FP8 and FP4 with FP32 accumulate. Other parameters repeat the RTX 5090 row; speculative decoding is off and no cache is reused. Concurrency rows: vLLM 0.26.0, AWQ 4-bit, CUDA graphs, 512-token prompts, 128 output tokens.
ConfigurationMeasuredPlanning modelError
Single-stream decode, llama.cpp Q4_K_M69.4 t/s53.1 t/s−23 %
Aggregate output, 1 stream59.4 t/s36.0 t/s−39 %
Aggregate output, 8 streams287.8 t/s258.7 t/s−10 %
Aggregate output, 16 streams384.9 t/s366.7 t/s−5 %
Aggregate output, 32 streams407.1 t/s377.8 t/s−7 %
First token, median, 1 stream203 ms1,145 ms+464 %
First token, median, 8 streams1,180 ms1,145 ms−3 %
First token, median, 32 streams2,077 ms3,621 ms+74 %

Aggregate throughput at 8 to 32 streams lands within about 10 % (−5 % to −10 %). That agreement mixes two errors. With the measured first-token medians taken out, decode alone ranges from 15 % low at 8 streams to 11 % high at 32, where the over-predicted first-token time offsets it.

Prefill is the weak part. At one stream the planning model reads the 512-token prompt at about 450 tokens per second, against at least 2,500 measured, so its first-token time is 5.6 times the measured value. It also has no prefill queue: its first-token time stays flat from one to eight streams, while the measured median rose from 203 ms to 1,180 ms. The two errors meet near eight streams. The 1-stream aggregate is low mostly because it includes the first-token time; with the measured first-token time it would still be 17 % low, from the decode under-prediction.

The test used 512-token prompts on a dense 9B model, so it exercises weight reads and compute. It does not test the terms that set the B200 figures: KV reads at 200,000 tokens, mixture-of-experts weight spread, speculative decoding and long-prompt prefill. We have not measured those.

Checking the numbers before publication caught several errors in our first version:

How the planning model calculates the numbers

Benchmark your own prompts on rented hardware for a day before you buy.

Try the planning model yourself

The Coding Speed Bench runs the same planning model in your browser. Pick GPUs, LLM, precision and layout, and set a speed target. Choose between holding speed and queuing or serving everyone more slowly, and add CPU-RAM or NVMe cache tiers. It then plays back a coding task at the estimated speed, using synthetic text and C code. It calls no LLM and fetches nothing.

The simulator shows slot counts, task time and first-token times directly. The developer, routing and cost figures follow from them by the arithmetic stated in each section. Its cluster output counts the answer phase only, about 12,800 t/s at the opening view. It opens on the worked example's hardware and workload, with 123 requests generating and 190 open sessions: the nearest slider steps to the 128 and 199 in the text.

Synthetic simulator · no LLM is called

Coding Speed Bench

Pick GPUs, LLM and layout and set a per-user speed target. See how many users each server can carry, how long the queue gets and how much a warm KV cache (GPU, CPU RAM or NVMe) shortens the second run. Press Run task to watch a sample embedded-C task stream at that speed.

Setup
Service level
When the server is full
Requests in flight at the same moment, across all servers.
Prompt & cache
Share of the input that repeats between calls (system prompt, repository files, earlier turns).
Calls where an early file changed, so no stored cache can be reused.
Everyone whose cached context should survive until their next call.
Response
Playback
Cache outcome for the next run
Speed
First token–
Thinking–
Answer–
Total per task–

Second run: where the cached context is found

Waiting in queue Reading input Thinking Writing answer
Ready

Press Run task to watch a simulated coding request. The cache outcome is drawn from the probabilities above unless you force one under Playback.

Setup variations for this workload

Each row is one server design evaluated with your current prompt, response and target speed. "Users / server at target" is how many can generate at once while each still gets the target speed. Click a row to load it.

Setup1 userUsers / server at targetServers for your loadGPUsFirst token, warmFirst token, coldRough rent / month

Planning model: each decoding step reads the active weights (more MoE experts get touched as more users share a step) plus every running request's KV cache, plus a fixed per-step overhead that grows when a model is split over GPUs, more over PCIe than over NVLink. Reading input is compute-bound. Warm-cache reloads move stored KV over PCIe (CPU RAM) or from NVMe instead of recomputing it, as LMCache, vLLM KV offloading or SGLang HiCache do. A copy split over more GPUs than the model has KV heads stores those heads on several GPUs. All five models here are Mamba or Gated DeltaNet hybrids, so the engine must store state checkpoints for prefix reuse; check that your engine version supports prefix caching for the chosen model.

All figures are planning-model estimates, not measurements. GPU figures are vendor dense-compute specifications; rents are on-demand prices per GPU-hour from provider pages opened on 2026-10-05: RunPod Community Cloud, and Crusoe for MI300X, where search results showed lower quotes we did not open. MI355X shows no rent: we found no on-demand price on a page we could open, and TensorWave lists it from $2.95 per GPU-hour on commitment. Its derating parameters are copied from the MI300X row. Benchmark your own workload before buying.

Sources

  1. Hugging Face model cards and configuration files: Qwen3.6-35B-A3B, Qwen3.8-27B, Qwen3.8-Flash-Next, Qwen3.5-9B and NVIDIA Nemotron 3.5 Lightning 30B-A3B, read 2026-10-05
  2. NVIDIA, Nemotron 3 Nano technical report, arXiv 2512.20848
  3. Anthropic, Claude API pricing, extended-thinking pricing and Message Batches documentation, read 2026-10-05
  4. NVIDIA H100, H200 and DGX B200 product pages; NVIDIA RTX Blackwell and RTX PRO Blackwell architecture whitepapers; AMD Instinct MI300X product page, read 2026-10-05; TensorWave MI355X product page for MI355X specifications, read 2026-10-05
  5. RunPod and Crusoe public GPU price pages, read 2026-10-05
  6. Project documentation: LMCache, vLLM KV offloading connector, SGLang HiCache, SGLang Model Gateway, llm-d, NVIDIA Dynamo router design, consulted 2026-10-05
  7. C42, RTX 5060 Ti 16 GB as an LLM server, measured 2026-08-12 and 2026-08-13

Measured on 2026-10-05 (tokenizer counts) and 2026-08-12 (RTX 5060 Ti comparison, one card, one LLM). The planning model ran on 2026-10-05 with the simulator on this page; every capacity, latency and cost figure is an estimate, not a measurement. Prices and rents change; re-check them before you rely on the cost tables.