Gimlet Labs Adds Cerebras to Deliver Ultrafast AI Inference through Gimlet Cloud
SAN FRANCISCO and SUNNYVALE, Calif., Sept. 28, 2026 (GLOBE NEWSWIRE) -- Gimlet Labs and Cerebras Systems (NASDAQ: CBRS) today announced a collaboration to deliver a new class of ultrafast AI inference at massive scale. The collaboration brings together Cerebras’ wafer-scale compute with the Gimlet Cloud to deliver a purpose-built disaggregated inference cloud spanning datacenter infrastructure to developer APIs. Together, the companies plan to deliver speeds of up to 3,000 tokens per second for demanding agentic and real-time applications, with the first Cerebras-powered Gimlet Cloud datacenter expected to come online later this year.
In AI, speed drives user experience and engagement and shapes what users can build. For real-time applications, from voice and video AI to agents and assistants, latency can be the difference between an interaction that feels seamless and one that feels slow. Real-time AI feels like an active collaborator that is immediate, fluid and responsive. When AI responds in real time, users do more with it, stay longer and run higher value workloads. As a result, fast tokens are more valuable tokens.
Get started
Create a free account to read the full story and Whiz Insights.
