Inference
From prototype to production, with the models you choose.
Serverless
Pay per token · no idle GPUs
Dedicated endpoint
Your weights or LoRA · autoscaling
Endpoints
| Endpoint | Model | Path | Replicas | Tokens / s | p50 latency | Status |
|---|---|---|---|---|---|---|
Serverless model catalog
Prices per 1M tokens · $0 egress| Model | Provider | Context | Input / 1M | Output / 1M | |
|---|---|---|---|---|---|