Nebius has announced that its Token Factory platform is the first AI cloud to integrate NVIDIA Groq 3 LPX, a purpose-built solution for ultra-fast token generation, designed to accelerate inference workloads on the NVIDIA Vera Rubin NVL72 architecture. The move positions Nebius at the forefront of high-performance inference, particularly for agentic applications where speed and responsiveness are critical to real-time decision-making and user interaction.
Token generation speed has become a defining factor in the performance of production-grade AI systems. Unlike training, which processes data in bulk, inference is split into two distinct phases: context ingestion and generation. Context involves handling long prompts, large codebases, and persistent multi-turn agent states, while generation produces the actual tokens that users or agents see and act upon. An agent may execute dozens of sequential model calls to complete a single task, with each step dependent on the previous one, meaning latency compounds across the entire workflow. As a result, generation speed increasingly determines how quickly an agent can deliver useful outcomes.
NVIDIA Groq 3 LPX is engineered specifically for long-context, low-latency performance. Each LPX rack pairs 256 LPU accelerators with 128 GB of on-chip SRAM and 640 TB/s of scale-up bandwidth, fully liquid-cooled using the NVIDIA MGX rack architecture. In independent benchmarking conducted by Artificial Analysis, the platform delivered 3,400 output tokens per second for a single user running Gemma 4 31B, the fastest performance ever recorded for that model. For larger-scale workloads, NVIDIA projects up to 35x higher inference throughput per megawatt for 2T-parameter models at long context and low latency compared to GB200 NVL72, a metric that underscores the gains achievable through extreme hardware-software codesign at agentic-scale token volumes.
A key advantage of the integration is that developers can access the new performance without re-architecting their existing stacks. Nebius designed Token Factory to run NVIDIA Vera Rubin NVL72 and Groq 3 LPX on the same platform that developers already use in production, initially supporting a subset of models alongside familiar features such as autoscaling and observability. This approach minimizes friction, allowing teams to adopt faster generation as a model-selection change rather than a migration, with no new SDK, vendor relationship, or billing setup required.
Token Factory already provides the building blocks for multi-agent systems, including native function calling, structured JSON outputs, and built-in safety guardrails for tool-calling agents, all on top of dedicated endpoints with sub-second latency targets. With the addition of NVIDIA Groq 3 LPX, developers building multi-agent pipelines, real-time coding assistants, and latency-sensitive reasoning applications can now achieve faster generation through the same operating model they use today. The combination of speed, scale, and lowest token cost positions Nebius as a compelling option for organizations pushing the boundaries of production AI inference.