AI Inference: Meta Teams with Cerebras on Llama API

Sunnyvale, CA — Meta has teamed with Cerebras on AI inference in Meta’s new Llama API, combining  Meta’s open-source Llama models with inference technology from Cerebras.

Developers building on the Llama 4 Cerebras model in the API can expect speeds up to 18 times faster than traditional GPU-based solutions, according to Cerebras. “This acceleration unlocks an entirely new generation of applications that are impossible to build on other technology. Conversational low latency voice, interactive code generation, instant multi-step reasoning, and real-time agents — all of which require chaining multiple LLM calls — can now be completed in seconds rather than minutes,” Cerebras said.

By partnering with Meta to serve Llama models from Meta’s new API service, Cerebras gains exposure to an expanded developer audience and deepens its business and partnership with Meta and their incredible teams.

Since launching its inference solutions in 2024, Cerebras has delivered the world’s fastest Llama inference, serving billions of tokens through its own AI infrastructure. The broad developer community now has direct access to a robust, OpenAI-class alternative for building intelligent, real-time systems — backed by Cerebras speed and scale.

“Cerebras is proud to make Llama API the fastest inference API in the world,” said Andrew Feldman, CEO and co-founder of Cerebras. “Developers building agentic and real-time apps need speed. With Cerebras on Llama API, they can build AI systems that are fundamentally out of reach for leading GPU-based inference clouds.”

Cerebras is the fastest AI inference solution as measured by third party benchmarking site Artificial Analysis, reaching over 2,600 token/s for Llama 4 Scout compared to ChatGPT at ~130 tokens/sec and DeepSeek at ~25 tokens/sec.

Developers will be able to access to the fastest Llama 4 inference by selecting Cerebras from the model options within the Llama API. This streamlined experience will make it easy to prototype, build, and scale real-time AI applications. To sign up for early access to the Llama API and to experience Cerebras speed today, visit www.cerebras.ai/inference.