Three tiers of open-source AI models. Pay only for tokens you use. All deployed in US-West for low-latency access.
Ultra-fast, sub-200ms TTFT
per 1M tokens
per 1M tokens
Balanced reasoning & code
per 1M tokens
per 1M tokens
Premium quality, deep reasoning
per 1M tokens
per 1M tokens
All prices in USD per million tokens. Volume discounts available for enterprise customers.
Built for developers who need reliability, transparency, and control.
Drop-in replacement. Change base_url and api_key — that's it. Works with all OpenAI SDKs and tools.
All servers in Singapore. Low latency for Asian and international users.
Create, name, and revoke keys instantly. Track per-key usage. No limit on number of keys.
Dedicated capacity means no cold starts and no noisy neighbors. Consistent TTFT, every request.
Full SSE streaming support. Usage reporting in final chunk for accurate billing.
Powered by Tencent Cloud TokenHub — DeepSeek, GLM, and Hunyuan. Enterprise-grade reliability.
You only pay for tokens you actually use. Input and output tokens are billed separately at different rates. We deduct from your prepaid balance after each request. No monthly fees, no minimums.
Flash is our fastest model with sub-200ms TTFT, great for high-volume chat. Reason offers the best balance of speed, quality, and price for most production workloads. Pro is our highest quality model for complex reasoning tasks.
Yes! New users get free credit to test all models. No credit card required — just sign up and start building.
All models are powered by Tencent Cloud TokenHub, providing access to DeepSeek V4 Flash, GLM 5.3, and Hunyuan 3. You can use the API output for commercial purposes with all plans.
Yes, full SSE streaming support with usage reporting in the final chunk. Works seamlessly with the OpenAI SDK's stream parameter.
All inference runs on servers in the AP-Singapore region. This provides low latency for Asian and international users.
Sign up for free credit. No credit card required. Ready in 2 minutes.
Create Free Account