No reserved GPU commitments. Pay only for the tokens and agent compute you consume with sub-15ms edge routing included.
Perfect for indie developers, prototypes, and testing our global inference latency.
Get Started FreeIdeal for fast-growing SaaS startups shipping production LLM and agent features.
Start 14-Day TrialFor scale-ups needing dedicated throughput, priority GPU routing, and security.
Upgrade to ScaleDedicated private GPU clusters, sovereign data fencing, and 99.999% SLA.
Everything included in our enterprise cloud platform at a glance.
| Feature & Metric | Developer | Growth | Scale | Enterprise |
|---|---|---|---|---|
| Monthly Included Tokens | 5 Million | 25 Million | 120 Million | Bespoke Volume |
| Global Edge Regions | 3 Regions | All 38 Regions | All 38 Regions | Custom Peering |
| Average P90 Latency | < 35ms | < 18ms | < 14ms | < 10ms |
| VectorFlow™ L1 Cache | — | 1M Vectors | 10M Vectors | Unlimited Shards |
| CloudGuard™ Security Suite | — | Standard | Advanced PII + Prompt | Custom Rules & SIEM |
| SLA Guarantee | — | 99.9% | 99.99% | 99.999% |
| Dedicated Technical Architect | — | — | — | Included |
Test our 14ms edge latency with \$50 in complimentary compute tokens.