Deploy open
LLMs on a
managed instance.
One OpenAI-compatible endpoint, 42+ open models, and flat monthly pricing — no per-token bills. Deploy a model, get an API key, and call it from the tools you already use.
Free tier · 250K tokens/mo · nothing to manage
The open-model API, without the per-token roulette.
One endpoint, a curated catalog of open models, and a bill you can predict. Built for developers who just want to ship.
One OpenAI-compatible endpoint
A single /v1/chat/completions endpoint that works with the OpenAI SDK and any tool you already use. Change the base URL and key — that's the whole migration.
42+ open models
DeepSeek, Qwen, Llama, GLM, Kimi, MiniMax, gpt-oss, Gemma and more — deploy any of them onto your instance. Each is routed to the cheapest provider serving it.
Flat monthly pricing
One monthly fee per instance tier — no per-token billing, no metered surprises. You always know the bill before the month starts.
Per-model keys & honest limits
Every deployed model gets its own API key. A dashboard shows your live throughput and monthly capacity, and your tier's limits are stated plainly — no hidden throttling.
42 open models, curated and ready.
The models actually worth running — each routed to the cheapest provider serving it. Deploy any on your instance.
DeepSeek
· 7Qwen
· 9Meta
· 4Mistral
· 2OpenAI
· 2Z.ai
· 6Moonshot
· 3MiniMax
· 4Nvidia
· 1Microsoft
· 1From model to endpoint in three steps.
Pick a model
Choose from 42 open models in one catalog — no juggling accounts and keys across five different providers.
Deploy it
Deploy onto your instance and get a scoped API key in seconds. No Dockerfiles, no GPUs, no DevOps to manage.
Call your endpoint
Point your OpenAI SDK at the endpoint with your key. Flat monthly cost — no surprise token bill at the end of the month.
$ curl https://gateway.ghosterr.io/v1/chat/completions \
-H "Authorization: Bearer sk-ghosterr-…" \
-d '{"model":"deepseek-v3.2","messages":[…]}'
◆ same schema as the OpenAI API
✓ { "choices": [{ "message": { "content": "Hello!" } }] }
billing flat monthly · no per-token charges
→ swap the base URL & key, keep your code
Flat monthly pricing. Start free.
Pick an instance tier — each is a flat monthly fee with a monthly token allocation and a throughput ceiling. No per-token billing, ever. Change tiers any time.
Small Instance
Runs the essential open-source models — great for light and testing workloads.
Monthly capacity: Standard capacity
10 modelsyou can deploy · 7 families
Medium Instance
More monthly capacity, with larger models unlocked including DeepSeek V3.
Monthly capacity: High capacity
23 modelsyou can deploy · 9 families
Large Instance
Top capacity with every supported model unlocked, including the largest.
Monthly capacity: Max capacity
42 modelsyou can deploy · 11 families
Enterprise Instance
Dedicated capacity, higher throughput, priority support, and custom model access — tailored to your workload.
Monthly capacity: Negotiated
Everything in Large,plus:
- Custom throughput & monthly capacity
- Dedicated support & SLAs
- Volume pricing & custom model access
Deploy your first model free.
Sign up, deploy an open model, and call it from your code in minutes. No card, no infra to manage — the Free tier is permanent.
42+ open models · OpenAI-compatible · flat monthly pricing