Chip targets low-latency token generation for AI agents

Nvidia said Monday that its Groq 3 LPX inference chip is now in full production, with AI cloud provider Nebius as the first AI cloud to deploy the new chip. The product stems from the chipmaker’s roughly $20 billion agreement with inference-chip specialist Groq, announced about eight months ago, and is designed to extend Nvidia’s Vera Rubin systems by generating tokens with low latency for AI agents handling long, complex tasks.

The launch marks the first commercial product from an arrangement that began in December, when Nvidia licensed Groq’s technology and hired founder and CEO Jonathan Ross and other team members. The deal gave Nvidia access to Groq’s expertise in inference computing — the use of already-trained AI models to solve problems and produce outputs — rather than the training of large language models that has driven most chip demand in recent years.

Nvidia’s graphics processing units have been the go-to chips for training large language models, but the focus is increasingly on inference, The Wall Street Journal reported. “Agentic AI creates two distinct computing challenges: efficiently processing enormous amounts of context and generating tokens with extremely low latency,” Nvidia said in announcing the Groq 3 LPX’s production status.

The chip is designed to extend the company’s Vera Rubin systems. According to Nvidia, the Groq 3 LPX is “purpose-built to extend Vera Rubin’s interactivity — the rate at which tokens are generated for an individual user, determining how quickly an agent can complete each step of its work.” Faster token generation allows agents to do more while “maintaining a responsive user experience,” the company said.

Advanced Micro Devices, a smaller GPU maker, said earlier this year it would integrate its rack-scale systems with chips from Cerebras, which recently went public, for low-latency inference, according to a CNBC report cited by The Wall Street Journal.

Separately, executives are encouraging staff to send messages in public Slack channels rather than direct messages, The Wall Street Journal reported, so that AI agents crawling through company communications can access updates on products, earnings, and strategic initiatives without being blocked by private channels.

The deal’s success may tell us a lot about the extent to which AI agents in the enterprise live up to expectations, according to The Wall Street Journal.