Instant
AI Inference
via API
Integrate 30+ pre-optimized AI models into your applications with our developer-friendly API and robust Inference infrastructure.
- Llama3.3 70B
- Deepseek-V4 Pro
- Deepseek-V4 Flash
- Qwen 3.6
- MiniMax-M3
- GLM5.2
- Kimi-K3
- Kimi-K2.6
Serverless Inference
Get access to 30+ models through API endpoints like Llama, DeepSeek, Qwen, MiniMax, Kimi and many others.
Integrate powerful open-source and multimodal AI capabilities.
From conversational AI to image generation and code assistance – into your applications.
Deploy any of our 30+ open-source AI models instantly without infrastructure setup. Get your AI endpoints live in seconds with automatic scaling and enterprise-grade reliability.
Test and prototype with models directly in our web-based playground before integrating. Experiment with parameters, compare model outputs, and generate API code snippets to accelerate your development workflow.
Sub-second response times with 99.9% uptime SLA, automatic load balancing, and global edge deployment. Built on our optimized inference stack to handle production workloads at any scale.
Secure API access with token-based authentication, custom rate limits, and usage analytics. Monitor consumption in real-time.
Affordable text generation and image creation with pay-as-you-use pricing model available.
| Model | Input | Output |
|---|---|---|
| Kimi-K3 | $5/M tokens | $15/M tokens |
| Kimi-K2.6 | $1.2/M tokens | $4.5/M tokens |
| Kimi-K2.5 | $0.6/M tokens | $3.0/M tokens |
| MiniMax-M3 | $0.3/M tokens | $1.2/M tokens |
| MiniMax-M2.7 | $0.4/M tokens | $1.5/M tokens |
| MiniMax-M2.5 | $0.3/M tokens | $1.2/M tokens |
| Nemotron-3-Super-120B-A12B | $0.4/M tokens | $0.8/M tokens |
| Nemotron-3-Ultra-550B-A55B | $0.6/M tokens | $3.5/M tokens |
| Qwen3-235B-A22B-Instruct | $0.25/M tokens | $1.0/M tokens |
| Qwen3.6-35B-A3B | $0.25/M tokens | $1.5/M tokens |
| GLM-5.2 | $1.3/M tokens | $4.2/M tokens |
| GLM-5.1 | $1.4/M tokens | $4.4/M tokens |
| GLM-5 | $1.0/M tokens | $3.0/M tokens |
| GLM-4.7 | $0.7/M tokens | $2.8/M tokens |
| Deepseek-V4-Pro | $1.7/M tokens | $3.4/M tokens |
| Deepseek-V4-Flash | $0.2/M tokens | $0.4/M tokens |
| Deepseek-R1 | $1.5/M tokens | $4.5/M tokens |
| Deepseek-V3.2 | $0.5/M tokens | $1.5/M tokens |
| Gemma-4-31B-IT | $0.35/M tokens | $0.8/M tokens |
| Gemma-4-26B-A4B-IT | $0.2/M tokens | $0.5/M tokens |
| GPT-OSS 120B | $0.15/M tokens | $0.6/M tokens |
| GPT-OSS 20B | $0.07/M tokens | $0.25/M tokens |
| Llama 3.3 70B Instruct | $0.7/M tokens | $0.7/M tokens |
| Llama 3.1 70B Instruct | $0.8/M tokens | $0.8/M tokens |
| Llama 3.1 8B Instruct | $0.15/M tokens | $0.15/M tokens |
| Mistral Large 3 | $0.5/M tokens | $1.5/M tokens |
| Mistral Medium 3.5 | $1.5/M tokens | $7.5/M tokens |
| Mistral Small 4 | $0.15/M tokens | $0.6/M tokens |
| Ministral 14B | $0.2/M tokens | $0.2/M tokens |
| Model | Price per image | Images per $1 |
|---|---|---|
| Flux Schnell | $0.003 | 333 |
| Flux Dev | $0.025 | 40 |
| Stable Diffusion XL | $0.005 | 200 |

