Deploy AI models on any GPU.
Pay by the second, not the token.

Choose your model, pick your GPU, configure scaling → deployed in seconds. Time-based pricing from $1.29/hr. No token math, no surprises.

Free playground available · No commitment · Pay only for running replicas
Why choose us

Three reasons to switch

What we do differently from every other GPU inference provider.

1. Billed by time, not tokens

Fixed GPU price per hour. The longer your prompts and outputs, the more you save vs. token-based providers. No input/output cost split to decode.

2. You choose the GPU

RTX PRO 6000, B200, B300, H200 → pick the GPU that fits your model and budget, even in serverless mode. Most providers abstract hardware away.

3. Serverless to dedicated in one click

Start serverless with scale-to-zero. When traffic grows, switch to dedicated GPU → same console, same API key, same workflow. No migration.

You are fully-guided on our console:
Model selector → GPU selector → Scaling config → Deploy

Arkane Cloud Console serverless process
Screenshot of Arkane Cloud's console: serverless onboarding

Trusted by 1,000+ AI startups, labs and enterprises.

How it works

From zero to production in 2 minutes

Choose a model

Browse the catalog or paste any Hugging Face model ID

Pick your GPU

Select GPU type and quantity per replica

Configure scaling

Light, Heavy or Custom : set min/max replicas and queue

Deploy & monitor

Hit deploy, get your API endpoint, track performance live

Pricing

Simple GPU pricing.
No token math.

Pay per GPU-hour, billed to the second. Auto-scaling adjusts the number of active GPUs to match your traffic.
RTX PRO 6000
$2.39/hr

Best for: cost-optimized inference, batch processing, dev/test

Deploy now
H200
$4.59/hr

Best for: large models (70B+), production workloads, low latency

Deploy now
B200
$6.99/hr

Best for: 200B+ models, maximum throughput, enterprise scale

Deploy now
B300
$9.99/hr

Best for: 400B+ models, long context, multi-model serving

Deploy now
Why time-based pricing wins

Same workload.
Different bill.

Same model. Same documents. Only the billing changes.
Arkane Cloud (per hour)

1 hour of RTX PRO 6000 running gpt-oss-120B. Process entire documents : all input tokens included at zero extra cost.

Token-based provider

Same hour, same model on Together AI. At $0.15/M input + $0.60/M output, document-heavy workloads cost 3.5× more.

Scenario: Document analysis · gpt-oss-120B

Processing long documents (16:1 input-to-output ratio) on 1× RTX PRO 6000 at 787 tok/s. Input tokens are free on GPU, while token providers charge for every one. Other example, for RAG workloads (8:1 ratio): $2.39 vs $5.10. Save 53%.

FAQ

Common questions.

1. How does time-based pricing work exactly?

You pay per GPU-hour (billed to the second). While your endpoint is active and processing requests, you're charged the hourly rate for the GPU(s) you selected. With scale-to-zero enabled, you pay $0 when there's no traffic. No token counting, no input/output rate split.

2. What about cold starts?

Cold start times depend on the model size and GPU. For most models, expect 30-60 seconds on first request after scale-to-zero. Set min replicas to 1 to eliminate cold starts entirely : you'll pay for one GPU continuously but get instant responses.

3. Can I deploy my own fine-tuned model?

Yes. Paste any Hugging Face model ID (public or private with an HF token) and we'll deploy it on your chosen GPU. Not limited to our catalog, any vLLM-compatible model works.

4. What are quiet hours?

Quiet hours reflect periods when your endpoint runs at minimum replicas instead of scaling up. For example, an RTX PRO 6000 endpoint set to min 1 / max 3 replicas costs $2.39/hr at quiet (1 replica) and $7.17/hr at peak (3 replicas). You only pay for the GPUs actually running. No fixed discount schedule, just automatic scaling based on your traffic.

5. How do I switch from token-based inference to GPU-hour pricing?

Choose your model, pick a GPU, and deploy. Your endpoint is live in under 2 minutes. You get an OpenAI-compatible API endpoint, so most codebases only need a URL and key swap. No architecture changes, no vendor lock-in.

6. Where are your data centers?

Our GPU infrastructure is located in Paris and Lyon, France : GDPR-compliant by default. We're expanding to additional EU and US regions in 2026.

Ready to Deploy ?

Start with serverless, scale to dedicated.
No credit card required to explore.

Free playground available · No commitment · Scale-to-zero = $0 when idle

Create your account