The AI inference revolution is reaching a crucial turning point, with major technology companies reassessing their hardware strategies to keep up with rising demand for inference tasks. While AI training—where models learn from data—has historically been the focus, inference—the process of using trained models to generate outputs like text, code, or images—has now become central. This shift is largely due to the rapid growth of large language models (LLMs), which have expanded from millions to trillions of parameters since 2020. For example, OpenAI’s GPT-3 scored 43.9 percent on a knowledge-and-reasoning benchmark in 2020, but by 2026, GPT-4o achieved 88.7 percent, nearly matching human performance.
Inference has become more computationally demanding, especially with the rise of reasoning models that produce far more output than traditional models. Moreover, agentic AI systems—those that perform tasks autonomously—operate continuously, further increasing the need for efficient inference capabilities. This growing demand has led to unexpected partnerships among tech giants. For example, OpenAI and Amazon have deployed Cerebras’s wafer-scale engine chips, even though Amazon has its own Trainium chips. Nvidia has acquired key talent and intellectual property from Groq in a $20 billion deal, and Anthropic is paying SpaceXAI over a billion dollars per month to access spare computing resources.
AI training and inference are fundamentally different in how they use computation. Training involves organizing unstructured data using a technique called backpropagation to refine model parameters, while inference uses a trained model to generate outputs based on input. Inference presents unique challenges, especially with autoregressive models, where each generated token depends on the previous one. This requires the model to read all previous weights and context, a process that has two main stages: prefill, where the model reads a prompt and calculates relationships between tokens, and decode, where the model generates tokens one at a time based on context stored in a key-value (KV) cache.
The KV cache, which stores the context of previous tokens, can grow to dozens of gigabytes, making the decode phase particularly memory-intensive. To address this, companies like d-Matrix and Majestic Labs are exploring new approaches. d-Matrix’s Raptor architecture minimizes the distance between compute and memory by stacking an AI accelerator directly on a DRAM die, reducing data travel to micrometers. Majestic Labs, in contrast, is improving the memory interface to handle longer wire traces, allowing the use of standard DRAM chips in larger quantities.
Nvidia and Amazon are adopting hybrid strategies, using different chips to handle various inference tasks. Nvidia’s Groq 3 language-processing unit (LPU) has 500 megabytes of on-die SRAM, offering seven times the memory bandwidth of a standard GPU. The company plans to use its newest Rubin GPUs for the compute-heavy prefill phase and the Groq 3 LPU for the memory-heavy decode phase. Amazon Web Services (AWS) has partnered with Cerebras to pair its Trainium chips with Cerebras’s Wafer-Scale Engine 3 (WSE-3), which integrates 44 gigabytes of SRAM directly into the chip.
Researchers are also working on optimizing AI inference through both software and hardware. Quantization, which converts models into lower-precision formats, is being used to reduce memory and computational needs without significantly affecting performance. Nvidia’s NVFP4 and AMD’s MXFP4 formats are examples of such efforts. Startups like Tensordyne and Etched are developing new architectures, including logarithmic number systems and direct silicon implementations of the transformer architecture, to improve inference efficiency.
The growing need for AI inference is expected to drive significant innovation in both hardware and software, much like the evolution of the CPU. As AI becomes more integrated into industries, the demand for efficient inference solutions will likely continue to rise, leading to a diverse range of approaches and technologies.
AI Inference Becomes Central to Tech Innovation as Hardware and Software Evolve
AI-rewritten from original reportingHow it works
ai-inferencehardwarellmnvidiacomputememory



