Docs

Inference

From prototype to production, with the models you choose.

Serverless

Pay per token · no idle GPUs

Dedicated endpoint

Your weights or LoRA · autoscaling

On your cluster

KServe · full infrastructure control

Pick a cluster

Endpoints

EndpointModelPathReplicasTokens / sp50 latencyStatus

Serverless model catalog

Prices per 1M tokens · $0 egress
ModelProviderContextInput / 1MOutput / 1M