RTX 5060 Ti 16 GB as an LLM server: 69 to 407 tokens per second, and the thermal wall
Measured LLM serving on a consumer RTX 5060 Ti 16 GB in a Dell PowerEdge R740: engine comparison, prefill against decode, the concurrency knee, what a full attention cache costs, and a hardware thermal slowdown that default chassis cooling never reported.