Sizing a self-hosted coding LLM: an estimated 200 developers at 100 tokens per second on one 8× B200 server
One server with eight B200 GPUs serves an estimated 200 developers at 100 tokens per second each, by our planning model. It runs Qwen3.6-35B-A3B in FP8 as four copies of two GPUs. That assumes 200,000-token agent prompts, a 30-second pause between prompts and every request reaching the copy that holds its cache. The server then generates about 12,000 tokens per second and receives about 474,000 input tokens per second, of which about 91,000 need computing.
At a 40-tokens-per-second target the same server reaches about 430 developers. With five teams in five time zones, each at high load five hours a day, it serves about 330. Beyond about 200 developers in one time zone, extra developers mostly wait. With 1,000 working continuously on one server, a task takes about 6.5 minutes on average instead of 54 seconds.
These are model estimates, not measurements, and we have not measured them on B200 hardware. We tested the planning model on the one GPU we have measured ourselves, an RTX 5060 Ti 16 GB. It came within about 10 % of aggregate throughput at 8 to 32 streams. It under-predicted single-stream decode by 23 % and predicted first-token time at low load at up to 5.6 times the measured value. That test does not exercise the long-context regime of this post; What the planning model gets wrong has the details. Every input is adjustable in the simulator further down.
What was measured, and what was modelled
Two things in this post are measurements. We counted tokens with the published tokenizers of four Qwen models on 2026-10-05. We compared the planning model with our own RTX 5060 Ti results from 2026-08-12. Model-card scores, configuration files, list prices and rents are published values we read on 2026-10-05. Everything else is the output of a bandwidth-and-compute planning model, run on 2026-10-05 for the workload described next. The delay section adds a finite-source queueing model on top of its output; the simulator does not run that step. From here on, the planning model means that model, and the LLMs it sizes are named.
What workload does the planning model size?
Agentic coding, not chat. Every example in this post assumes the following request shape.
| Parameter | Value |
|---|---|
| Input per request: repository files, history, tool output | 200,000 tokens |
| Share of the input reused from the previous turn | 85 % |
| Calls where an early file changed, so no cached prefix is usable | 5 % |
| Thinking tokens | 2,048 |
| Answer tokens, about 65 % of them code | 3,000 |
| Pause between one developer's prompts | 30 s |
| Speculative decoding (multi-token prediction) | on |
| Cache storage per server | GPU memory and 1 TB CPU RAM |
Agents resend most of their context on every turn. Input is therefore large but mostly repeated, and the prompt cache matters as much as the GPU.
Why is per-user speed a memory-bandwidth problem?
Per-user speed is set by how fast each decoding step reads weights and KV cache from GPU memory. In the planning model's worked example, those reads take 7.6 ms of the 10 ms step; the arithmetic takes about 0.07 ms.
Each decoding step produces one token per running request, or more when speculative-decoding drafts are accepted. Each step reads the weights of every expert that any request in the batch routes to. It also reads the KV cache and recurrent state of every request. On top comes a fixed overhead per step, plus communication time when a copy spans several GPUs. Two consequences follow.
- More users per step means slower steps. A 200,000-token context on Qwen3.6-35B-A3B holds 2.05 GB of KV cache at 8-bit. At 32 requests per copy the planning model reads about 91 GB per step: 68 GB of KV cache and recurrent state, and 23 GB of expert weights, because the batch touches 62 % of the experts, counting one token per request.
- Splitting an LLM across GPUs adds overhead. Spreading a small LLM over eight GPUs gains bandwidth but adds communication time to every step. In the planning model that time is 1.5 ms per doubling of GPUs over PCIe and 0.4 ms to 0.5 ms over NVLink.
At 100 tokens per second a step must finish in about 10 ms. At 32 requests per copy the planning model computes a 10.0 ms step; its speculative-decoding gain (+21 %) and its batch-interference penalty (−17 %) roughly cancel.
Which open models does the planning model cover?
Four open-weight LLMs, all of them hybrids in which only some layers hold KV cache. The score column repeats what each publisher reports; we ran none of these benchmarks.
| LLM | Parameters and layers | Published coding score | Notes |
|---|---|---|---|
| Qwen3.6-35B-A3B | 35B total, 3B active, MoE; 30 Gated DeltaNet and 10 attention layers | SWE-bench Verified 70.1 in NVIDIA's comparison; Qwen's own card reports 73.4 | 262,144-token native context; 2.05 GB of KV per 200,000-token context at 8-bit |
| Qwen3.8-27B | 27B, dense; 48 Gated DeltaNet and 16 attention layers | SWE-bench Pro 61.7, Terminal-Bench 2.1 73.0 (Qwen card) | Planning-model estimate: 2 to 2.4 times the GPUs per developer of Qwen3.6-35B-A3B at 100 t/s |
| Qwen3.8-Flash-Next | 125B, 6B active, MoE, plus a 51B n-gram embedding and 4B MTP (180B total); 36 Gated DeltaNet and 12 sparse-attention layers | SWE-bench Pro 62.5 (Qwen card) | About 100 GB of 4-bit weights: one copy fits on one B200, H200 or MI300X, or on two H100s |
| Nemotron 3.5 Lightning | 30B total, 3B active, MoE; 23 Mamba-2, 23 MoE and 6 attention layers | SWE-bench Verified 51.6 (NVIDIA card) | 3,072 bytes of KV per token; lowest SWE-bench Verified score in this table |
How many developers get 100 tokens per second on one 8× B200 server?
By the planning model, about 200 with Qwen3.6-35B-A3B FP8 as four copies of two GPUs. The other layouts in the table give 45 to 225. Concurrent slots are requests generating at the same moment while each still gets its answer at 100 tokens per second. The target applies to answer tokens; at that load thinking tokens run at about 92 tokens per second, because speculative decoding gains less on prose.
| LLM, layout | Concurrent slots | Developers with 30 s pauses |
|---|---|---|
| Qwen3.6-35B-A3B FP8, 4 copies × 2 GPUs | 128 | about 200 |
| Qwen3.6-35B-A3B FP8, 8 copies × 1 GPU | 144 | about 225 |
| Qwen3.6-35B-A3B FP8, 2 copies × 4 GPUs | 68 | about 105 |
| Qwen3.6-35B-A3B FP8, 1 copy × 8 GPUs | 34 | about 55 |
| Qwen3.8-27B 4-bit, 4 copies × 2 GPUs | 60 | about 95 |
| Qwen3.8-27B 4-bit, 2 copies × 4 GPUs | 52 | about 80 |
| Qwen3.8-Flash-Next 4-bit, 1 copy × 8 GPUs | 28 | about 45 |
Qwen3.6-35B-A3B and Qwen3.8-Flash-Next have two KV heads each. With tensor-parallel attention, a copy spread over four or eight GPUs stores every KV head on several GPUs, which is why those layouts lose slots in the planning model. Engines that run attention data-parallel and the experts expert-parallel avoid that copy, at a communication cost the planning model does not include. Eight single-GPU copies give the most slots, but they split the cache eight ways; the routing section shows what that costs. The worked example uses four copies of two GPUs.
By the planning model, eight AMD Instinct MI355X GPUs serve about 160 developers at 100 tokens per second with Qwen3.6-35B-A3B FP8 as four copies of two GPUs (104 slots), and about 185 as eight single-GPU copies (120 slots). Each MI355X has 288 GB of memory and the same 8 TB/s of memory bandwidth as a B200. The planning model applies its MI300X parameters, none with a published basis: kernels derated by 15 %, 2.5 ms of fixed overhead per step against 2 ms on the B200, and 0.6 ms of link time per GPU doubling against 0.4 ms. With the B200 values the MI355X would match the B200 slot counts. None of the price pages we opened shows an on-demand MI355X price.
The planning model treats the 12 sparse-attention layers of Qwen3.8-Flash-Next as dense attention, so each step reads the KV cache of every stored token. The model card limits each query to 2,048 selected tokens. Read the Qwen3.8-Flash-Next row as a likely underestimate.
Developers served come from Little's law: developers = slots × (task time + pause) ÷ task time. At 100 tokens per second a task takes about 54 seconds, so a 30-second pause adds about 56 % more developers. That count is the saturation point. At about 200 developers the 128 slots are full on average. If developers act independently, more than 128 want to generate 47 % of the time, and requests queue. About 180 developers keep that below 2.2 % (binomial estimate on the same task time and pause).
Does a lower speed target serve more developers?
Yes, but total output changes little: the server generates 11,900 to 13,600 tokens per second whatever the target. A lower target spreads the same output over more people.
| Answer speed per user | Developers (30 s pause) | Task time | Tokens generated, total | Input computed |
|---|---|---|---|---|
| 100 t/s | about 200 | 54 s | about 12,000 t/s | about 91,000 t/s |
| 80 t/s | about 230 | 66 s | about 12,000 t/s | about 91,000 t/s |
| 60 t/s | about 275 | 87 s | about 11,900 t/s | about 91,000 t/s |
| 40 t/s | about 430 | 131 s | about 13,600 t/s | about 104,000 t/s |
At 40 tokens per second the limit is still decoding, at 88 requests per copy. Computing new input then uses 37 % of the planning model's idle prefill capacity of about 277,000 tokens per second.
What delay do developers feel when more of them share one server?
By the planning model, a task on one 8× B200 server takes about 6.5 minutes on average with 1,000 developers, from submit to finished answer. It takes about 2.3 minutes with 400 and 30 minutes with 4,000. These figures hold answers at 100 tokens per second and queue the rest. Each developer works continuously, all in one time zone: submits a task, waits for the answer, reads it for 30 seconds and submits the next.
| Developers | Hold speed, queue the rest: mean time per task | Of which, queueing | Serve all, slower: mean time per task | Answer speed, serve all | Tasks per developer per hour, hold / serve all |
|---|---|---|---|---|---|
| 100 | 24 s | none | 24 s | about 245 t/s | 67 / 67 |
| 200 | 55 s | 3 s | 55 s | about 100 t/s | 42.5 / 42.5 |
| 400 | 2.3 min | 1.4 min | 2.0 min | about 43 t/s | 21 / 24 |
| 1,000 | 6.5 min | 5.6 min | 5.7 min | about 29 t/s | 8.5 / 9.8 |
| 4,000 | 30.5 min | 29.5 min | 25.7 min | about 29 t/s | 1.9 / 2.3 |
Below capacity, answers run faster than the target: with 100 developers a request shares its copy with few others and streams at about 245 tokens per second. Above it, holding speed keeps every answer at 100 tokens per second once it starts, but developers queue first. Serving all admits up to 133 requests per copy, the KV-cache limit, and slows every answer. At 400 and 1,000 developers it finishes slightly more work, 2.6 to 2.7 tasks per second against 2.4, because a larger batch spreads each step's weight reads and fixed overhead over more requests.
Neither policy keeps up with headcount. Tasks per developer per hour fall in inverse proportion to it: 21 at 400, 8.5 at 1,000 and 1.9 at 4,000 when holding speed. Holding speed keeps every context in the cache up to 1,000 developers. At 4,000, three calls in four start cold and throughput drops to 2.15 tasks per second. Serving all fills GPU memory with running requests, so about half of all calls already start cold at 1,000 developers.
Keeping the mean queueing time under 5 seconds takes 2 servers for 400 developers, 5 for 1,000 and 20 for 4,000. At $5.98 per GPU-hour around the clock, the lowest on-demand B200 price we found on 2026-10-05, that is about $70,000, $175,000 and $698,000 a month.
How much faster is a warm cache?
In the planning model, a warm cache cuts first-token time at full load from 9.3 s to 1.4 s for Qwen3.6-35B-A3B FP8. When the 170,000-token prefix is still stored, the server only computes the 30,000 new tokens.
| Where the cached context is found | Qwen3.6, 1 per copy | Qwen3.6, full load | Qwen3.8-27B, 1 per copy | Qwen3.8-27B, full load |
|---|---|---|---|---|
| GPU memory | 0.5 s | 1.4 s | 1.0 s | 1.5 s |
| CPU RAM, reloaded over PCIe | 0.5 s | 1.5 s | 1.1 s | 1.5 s |
| Nowhere: all 200,000 tokens recomputed | 2.9 s | 9.3 s | 6.3 s | 9.4 s |
Reloading the prefix from CPU RAM adds 0.04 s to 0.07 s over keeping it in GPU memory. Recomputing it adds 2.4 s to 7.9 s. LMCache, the vLLM KV offloading connector and SGLang HiCache all reload stored KV this way.
The limit is capacity: about 1.8 GB per developer for Qwen3.6-35B-A3B and 5.7 GB for Qwen3.8-27B. Two hundred idle developers fill 0.36 TB or 1.1 TB. Plan 1 TB to 2 TB of CPU RAM per server for the cache tier.
Only the attention layers of these hybrids hold KV cache, which keeps the figures small. To reuse a prefix, the engine also has to store recurrent-state checkpoints for the Gated DeltaNet or Mamba-2 layers. For the CPU-RAM tier, the offload connector must store and reload those checkpoints with the KV cache. At 40 tokens per second the planning model finds about half of the warm prefixes in CPU RAM, so the 430-developer estimate depends on it. Confirm that your engine version supports prefix caching and KV offloading, including recurrent state, for the LLM you pick.
Why does round-robin load balancing break this?
Round robin ignores where a developer's cache lives. With four model copies and a cache per copy, a returning request lands on the copy that served its previous turn one time in four.
| Routing | Calls that start warm | Input to compute | Share of idle prefill capacity |
|---|---|---|---|
| Round robin over 4 copies, cache in GPU memory only | 24 % | about 379,000 t/s | 137 %: queues grow |
| Round robin over 8 single-GPU copies, GPU memory only | 12 % | about 481,000 t/s | 174 %: queues grow |
| Cache-aware or sticky routing, 4 copies | 95 % | about 91,000 t/s | 33 % |
| Cache-aware or sticky routing, 8 copies | 95 % | about 103,000 t/s | 37 % |
Use sticky sessions keyed on developer or repository, or a cache-aware router such as SGLang Model Gateway, llm-d or NVIDIA Dynamo. In this workload the routing choice changes input demand by a factor of about four to five.
How many developers fit when five teams share one server?
About 330 by the planning model, as long as the busy hours fall at different times in UTC. Take five teams of equal size: US West, US East, EU West, EU East and Central Asia. Each works at high load for five hours, from 10:00 to 15:00 local time.
| Team | UTC offset, winter / summer | High load in UTC, winter | High load in UTC, summer |
|---|---|---|---|
| US West | −8 / −7 | 18:00 to 23:00 | 17:00 to 22:00 |
| US East | −5 / −4 | 15:00 to 20:00 | 14:00 to 19:00 |
| EU West | +0 / +1 | 10:00 to 15:00 | 09:00 to 14:00 |
| EU East | +2 / +3 | 08:00 to 13:00 | 07:00 to 12:00 |
| Central Asia | +5 / +5 | 05:00 to 10:00 | 05:00 to 10:00 |
In winter at most two teams overlap. In summer, daylight saving time moves the European blocks an hour earlier in UTC. From 09:00 to 10:00 UTC three teams then overlap: Central Asia, EU East and EU West. Size for the summer peak.
At saturation one server carries about 200 developers at high load, at 100 tokens per second each. With three teams at the peak, each team can have 66 developers: 330 on one server, against about 200 if everyone worked in one time zone. With the burst headroom of about 180 developers, each team can have 60, or 300 in all. No team is at high load from 22:00 to 05:00 UTC in summer, or from 23:00 to 05:00 UTC in winter.
Where do the night-time scheduled tasks go?
Into the quiet window, run from one queue. Add scheduled agent tasks worth half of each team's interactive volume, with the same request shape. That is about 35,000 tasks a night across the five teams. Scheduled tasks do not need 100 tokens per second. At 40 tokens per second the server finishes about 9,700 tasks an hour, so the whole batch needs 3.7 hours of the quiet window.
Scheduling by each team's local night breaks this. If every team runs its batch from 00:00 to 05:00 local time, the US West batch lands on the summer three-team peak at 09:00 UTC. Generation demand then reaches 116 % of what the server produces at the 100-tokens-per-second operating point, so interactive users slow down or queue. In winter the worst hour reaches 99 %. Schedule by server load in UTC, not by each team's clock.
What does the same workload cost on a frontier API?
On Claude Sonnet 5.5 list prices read 2026-10-05, the five-team workload costs about $399,000 a month with scheduled tasks and $266,000 without. One 8× B200 server rented around the clock costs about $35,000 a month. The workload is 330 developers at 100 tokens per second, five busy hours a day and 21 working days. Scheduled tasks add half the interactive volume. That is about 1.49 million interactive and 0.74 million scheduled tasks a month.
The list prices are $2 per million input tokens and $10 per million output tokens. Cache writes with a 5-minute lifetime cost $2.50 per million and cache reads $0.20 per million. Anthropic bills thinking as output. For this workload that is about $0.18 per task.
| Line item per task | Cost |
|---|---|
| Cache read, 170,000 tokens on 95 % of calls | $0.032 |
| New tokens written to cache, 30,000 on 95 % of calls | $0.071 |
| Cold calls writing the whole 200,000-token prefix, 5 % of calls | $0.025 |
| Output, 5,048 tokens including thinking | $0.050 |
| Total | $0.179 |
| Option | Monthly |
|---|---|
| Sonnet 5.5 API, interactive tasks only | about $266,000 |
| Sonnet 5.5 API, interactive and scheduled tasks | about $399,000 |
| Rent one 8× B200 server at $5.98 per GPU-hour | about $35,000 |
| Rent two such servers, if real capacity is half the planning model's | about $70,000 |
If the planning model's capacity holds, the API with scheduled tasks costs about 11 times as much as one rented server. It costs about 6 times as much as two. The rented server runs around the clock, so the scheduled tasks fill hours that would otherwise sit idle. The Message Batches API halves API prices, per its documentation read 2026-10-05. Each batch request is single-shot, without a tool loop. An agent loop, in which each request needs the tool results of the one before, does not run on it as it stands.
The LLMs are not equivalent, though. We have not measured how many turns each needs per finished change, and that ratio decides the comparison; measure it on your own tasks. Self-hosting also adds operations work, hardware risk and capacity you pay for when nobody is working. A hybrid is worth testing: routine edits on a local LLM, hard tasks sent to the API.
Three changes cut the API bill for this workload:
- Send fewer new tokens per turn. New-token cache writes are the largest item, $0.071 of $0.179, and large tool outputs resent every turn inflate them.
- Cap thinking. Each extra 1,000 thinking tokens per task costs about $22,000 a month.
- Keep calls warm. Each 1 % of calls that start cold costs about $8,700 a month.
Which part of a build log costs the most tokens?
The timestamps. The published tokenizers of four Qwen models (two distinct tokenizer files) give every digit its own token. Each of the seven separators takes one more. A 28-character ISO 8601 timestamp therefore costs 28 tokens. On three sample build lines it is 49 % to 80 % of the tokens.
| Line after the timestamp | Tokens with timestamp | Tokens without | Timestamp share |
|---|---|---|---|
make[2]: Entering directory | 35 | 7 | 80 % |
[ 42%] Building C object src/CMakeFiles/app.dir/drivers/uart.c.obj | 48 | 20 | 58 % |
src/drivers/uart.h:41:20: warning: 'uart_flush' declared 'static' but never defined [-Wunused-function] | 57 | 29 | 49 % |
Take a 10,000-line build log whose lines average the 19 tokens of our three samples. By our calculation it holds about 467,000 tokens, of which 280,000, 60 %, are timestamps. That is beyond the 262,144-token native context of Qwen3.6-35B-A3B. Without the timestamps the same log holds about 187,000 tokens and fits. Sent to Claude Sonnet 5.5 as new input, the timestamps alone cost $0.56 per read at $2 per million tokens.
Strip the timestamps before a log reaches an LLM. Keep a short marker where the log went silent for long, so a build that stopped responding stays visible.
What the planning model gets wrong
We have not calibrated the planning model for the regime of this post. We compared it with our own published measurements of an RTX 5060 Ti 16 GB serving Qwen3.5-9B, measured 2026-08-12. For the AWQ 4-bit runs we used the planning model's W4A16 path, which matches that kernel class.
| Configuration | Measured | Planning model | Error |
|---|---|---|---|
| Single-stream decode, llama.cpp Q4_K_M | 69.4 t/s | 53.1 t/s | −23 % |
| Aggregate output, 1 stream | 59.4 t/s | 36.0 t/s | −39 % |
| Aggregate output, 8 streams | 287.8 t/s | 258.7 t/s | −10 % |
| Aggregate output, 16 streams | 384.9 t/s | 366.7 t/s | −5 % |
| Aggregate output, 32 streams | 407.1 t/s | 377.8 t/s | −7 % |
| First token, median, 1 stream | 203 ms | 1,145 ms | +464 % |
| First token, median, 8 streams | 1,180 ms | 1,145 ms | −3 % |
| First token, median, 32 streams | 2,077 ms | 3,621 ms | +74 % |
Aggregate throughput at 8 to 32 streams lands within about 10 % (−5 % to −10 %). That agreement mixes two errors. With the measured first-token medians taken out, decode alone ranges from 15 % low at 8 streams to 11 % high at 32, where the over-predicted first-token time offsets it.
Prefill is the weak part. At one stream the planning model reads the 512-token prompt at about 450 tokens per second, against at least 2,500 measured, so its first-token time is 5.6 times the measured value. It also has no prefill queue: its first-token time stays flat from one to eight streams, while the measured median rose from 203 ms to 1,180 ms. The two errors meet near eight streams. The 1-stream aggregate is low mostly because it includes the first-token time; with the measured first-token time it would still be 17 % low, from the decode under-prediction.
The test used 512-token prompts on a dense 9B model, so it exercises weight reads and compute. It does not test the terms that set the B200 figures: KV reads at 200,000 tokens, mixture-of-experts weight spread, speculative decoding and long-prompt prefill. We have not measured those.
Checking the numbers before publication caught several errors in our first version:
- Our first version of this comparison used FP16-accumulate compute for the RTX 5060 Ti and reported throughput at 32 streams 80 % above the measurement. With FP32 accumulate, the convention the simulator uses for every GeForce card, the error is −7 %.
- Our first version of the planning model split KV cache evenly over any number of GPUs. Tensor-parallel attention replicates KV heads when a copy spans more GPUs than the LLM has KV heads. With that corrected, Qwen3.6-35B-A3B at 2 copies × 4 GPUs drops from 98 to 68 slots, and at 1 copy × 8 GPUs from 99 to 34.
- Our first draft quoted warm-cache first-token times at an unstated light load instead of full load, and priced GPU time below every on-demand quote we found.
- It carried KV sizes off by 30 % to a factor of two against the published configuration files, and claimed an accuracy that nothing we measured supports.
- The planning model still counts one token per request when it works out how many experts a step touches. With speculative decoding on, each step also verifies draft tokens. Counting one draft token per request would cut the worked example from 128 to 120 slots.
How the planning model calculates the numbers
- Decoding: step time = fixed overhead + max(bytes read ÷ effective bandwidth, compute ÷ effective FLOPS). Bytes read are the active weights, more mixture-of-experts weights as more requests share a step, and each running request's KV cache and recurrent state. Per-request speed is 1 ÷ step time, cut by a batch-interference penalty of up to 17 %, full at 16 or more requests per copy. A speculative-decoding gain raises it and falls linearly to zero at 48 requests per copy.
- Input processing: compute-bound, slower for very long prompts and shared with other requests.
- Cache reuse: stored prefixes reload over PCIe from CPU RAM or from NVMe instead of recomputing them, limited by the space left for the cache.
- KV size: 2 × attention layers × KV heads × head dimension × 1 byte per token, from each LLM's published configuration file. It is multiplied when a copy spans more GPUs than the LLM has KV heads.
- Hardware: vendor dense-compute and bandwidth specifications, derated for real kernels.
- Delay under load: each developer is a closed loop of submit, wait and a 30-second pause, routed to one model copy with its own queue. A finite-source queueing model with exponential task and pause times gives the mean time per task. Deep in saturation that mean equals developers ÷ throughput − pause, whatever the distributions.
Benchmark your own prompts on rented hardware for a day before you buy.
Try the planning model yourself
The Coding Speed Bench runs the same planning model in your browser. Pick GPUs, LLM, precision and layout, and set a speed target. Choose between holding speed and queuing or serving everyone more slowly, and add CPU-RAM or NVMe cache tiers. It then plays back a coding task at the estimated speed, using synthetic text and C code. It calls no LLM and fetches nothing.
The simulator shows slot counts, task time and first-token times directly. The developer, routing and cost figures follow from them by the arithmetic stated in each section. Its cluster output counts the answer phase only, about 12,800 t/s at the opening view. It opens on the worked example's hardware and workload, with 123 requests generating and 190 open sessions: the nearest slider steps to the 128 and 199 in the text.
Coding Speed Bench
Pick GPUs, LLM and layout and set a per-user speed target. See how many users each server can carry, how long the queue gets and how much a warm KV cache (GPU, CPU RAM or NVMe) shortens the second run. Press Run task to watch a sample embedded-C task stream at that speed.
Second run: where the cached context is found
Press Run task to watch a simulated coding request. The cache outcome is drawn from the probabilities above unless you force one under Playback.
Setup variations for this workload
Each row is one server design evaluated with your current prompt, response and target speed. "Users / server at target" is how many can generate at once while each still gets the target speed. Click a row to load it.
| Setup | 1 user | Users / server at target | Servers for your load | GPUs | First token, warm | First token, cold | Rough rent / month |
|---|
Planning model: each decoding step reads the active weights (more MoE experts get touched as more users share a step) plus every running request's KV cache, plus a fixed per-step overhead that grows when a model is split over GPUs, more over PCIe than over NVLink. Reading input is compute-bound. Warm-cache reloads move stored KV over PCIe (CPU RAM) or from NVMe instead of recomputing it, as LMCache, vLLM KV offloading or SGLang HiCache do. A copy split over more GPUs than the model has KV heads stores those heads on several GPUs. All five models here are Mamba or Gated DeltaNet hybrids, so the engine must store state checkpoints for prefix reuse; check that your engine version supports prefix caching for the chosen model.
All figures are planning-model estimates, not measurements. GPU figures are vendor dense-compute specifications; rents are on-demand prices per GPU-hour from provider pages opened on 2026-10-05: RunPod Community Cloud, and Crusoe for MI300X, where search results showed lower quotes we did not open. MI355X shows no rent: we found no on-demand price on a page we could open, and TensorWave lists it from $2.95 per GPU-hour on commitment. Its derating parameters are copied from the MI300X row. Benchmark your own workload before buying.
Sources
- Hugging Face model cards and configuration files: Qwen3.6-35B-A3B, Qwen3.8-27B, Qwen3.8-Flash-Next, Qwen3.5-9B and NVIDIA Nemotron 3.5 Lightning 30B-A3B, read 2026-10-05
- NVIDIA, Nemotron 3 Nano technical report, arXiv 2512.20848
- Anthropic, Claude API pricing, extended-thinking pricing and Message Batches documentation, read 2026-10-05
- NVIDIA H100, H200 and DGX B200 product pages; NVIDIA RTX Blackwell and RTX PRO Blackwell architecture whitepapers; AMD Instinct MI300X product page, read 2026-10-05; TensorWave MI355X product page for MI355X specifications, read 2026-10-05
- RunPod and Crusoe public GPU price pages, read 2026-10-05
- Project documentation: LMCache, vLLM KV offloading connector, SGLang HiCache, SGLang Model Gateway, llm-d, NVIDIA Dynamo router design, consulted 2026-10-05
- C42, RTX 5060 Ti 16 GB as an LLM server, measured 2026-08-12 and 2026-08-13