Skip to Content

INFERENCE IS THE NEW TRAINING: WHY HPC’S JOB ISN’T DONE WHEN THE MODEL SHIPS

October 6, 2026

For years, the High-Performance Computing (HPC) conversation around AI has centered on training: bigger clusters, more GPUs, faster interconnects, all in service of getting a model from raw data to a finished set of weights. Once that model shipped, the assumption was that the hard computational problem was solved — inference was “just” running the model forward, a comparatively cheap operation you could push to a modest server or even a phone.

That assumption is breaking down. The newest generation of AI models doesn’t just answer — it reasons, generating intermediate chains of thought before producing a final response.1 And reasoning takes compute. Understanding why is the next chapter in the HPC-AI story.

Why inference stopped being cheap

Reasoning-oriented models generate intermediate steps — chains of thought, self-checks, multiple candidate answers evaluated before a final one is selected — before producing a response. Every one of those intermediate steps is itself a forward pass through the model. A question that used to cost one inference call can now cost dozens, sometimes hundreds, depending on how much “thinking” the model is allowed to do.

This is often described as trading training-time compute for inference-time compute, and recent research suggests the tradeoff can be favorable: allocating more compute at test time can be more effective than simply scaling model parameters further, at least on certain reasoning-heavy tasks. 2 But the cost doesn’t disappear — it moves. It shifts from a one-time, amortizable training run to a recurring, per-query expense that scales with usage.

Why this is, again, an HPC problem

Serving reasoning-heavy models at scale reproduces many of the same engineering problems that HPC has spent decades solving for training — just under different constraints.

Latency now competes with throughput. A training cluster can batch work for hours; a user waiting on a reasoning response cannot. Serving infrastructure has to parallelize aggressively across GPUs to keep response times reasonable while still batching enough requests to use hardware efficiently — a scheduling problem with a much tighter time budget than training ever had.

Memory bandwidth becomes the bottleneck sooner. Generating a long chain of intermediate reasoning tokens means the model’s key-value cache keeps growing over the course of a single response. Serving systems such as vLLM were built specifically to address this: by managing the KV cache the way an operating system manages virtual memory pages, they cut memory waste and roughly double to quadruple throughput at the same latency compared to earlier serving stacks.3 High-bandwidth memory, once mainly a training-cluster concern, is now directly on the critical path for how many concurrent reasoning conversations a system can serve at once.

Cost predictability gets harder. When a query might trigger five reasoning steps or five hundred depending on its difficulty, capacity planning stops being a matter of counting expected requests and starts requiring the same kind of statistical modeling HPC teams use for variable, bursty workloads.

What this changes, practically

For organizations deploying these models, the shift has concrete consequences. Inference infrastructure needs the same discipline that training infrastructure has had for years: profiling, capacity planning, and hardware choices driven by measured bottlenecks rather than by whatever configuration happened to be available.

It also reopens a question that looked settled: on-premises versus cloud for inference is no longer just a cost-and-compliance conversation — it’s an architecture conversation, because the answer depends on how much inference-time compute a given use case actually needs, and that number is far less stable than it used to be.

The pattern repeats

The broader lesson echoes the argument for HPC as AI’s backbone: every time AI capability takes a step forward, the compute demand it creates doesn’t stay where it started. It used to be exclusively a training-time problem. Now it’s a training-time and inference-time problem, and the second half of that sentence is growing faster than most serving infrastructure was built to handle.

HPC’s role in AI was never only about making training runs finish faster. It’s about making sure the infrastructure keeps up with wherever the compute demand moves next — and right now, a good part of it is moving to the moment a model actually answers.

About the author

Project Manger | France
Dr. Ait Kaci Célia is a project manager specializing in digital systems and high-performance computing (HPC). She holds a PhD in HPC and simulation from the University of Bordeaux, completed in collaboration with INRIA and Atos, as well as dual master’s degrees in applied mathematics, statistics, and high-performance computing.

References

  1. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., & Zhou, D. (2023). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv:2201.11903. ↩︎
  2. Snell, C., Lee, J., Xu, K., & Kumar, A. (2024). Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. arXiv:2408.03314. ↩︎
  3. Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., & Stoica, I. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention. Proceedings of the 29th ACM SIGOPS Symposium on Operating Systems Principles (SOSP ’23). arXiv:2309.06180. ↩︎

Leave a Reply

Your email address will not be published. Required fields are marked *

Slide to submit