Pay as you go · No minimums · No credit card required

Simple, Transparent Pricing

Three tiers of open-source AI models. Pay only for tokens you use. All deployed in US-West for low-latency access.

Flash

Ultra-fast, sub-200ms TTFT

Input$0.08

per 1M tokens

Output$0.18

per 1M tokens

  • DeepSeek V4 Flash
  • 128K context window
  • ~180ms TTFT
  • 120 tokens/s throughput
  • Powered by TokenHub
Start Free
MOST POPULAR

Reason

Balanced reasoning & code

Input$0.30

per 1M tokens

Output$0.65

per 1M tokens

  • GLM 5.3
  • 128K context window
  • ~350ms TTFT
  • 70 tokens/s throughput
  • Strong code & reasoning
  • Powered by TokenHub
Start Building

Pro

Premium quality, deep reasoning

Input$0.75

per 1M tokens

Output$1.60

per 1M tokens

  • Hunyuan 3
  • 128K long context
  • ~600ms TTFT
  • 40 tokens/s throughput
  • Deep reasoning capability
  • Powered by TokenHub
Go Pro

All prices in USD per million tokens. Volume discounts available for enterprise customers.

Everything You Need for Production

Built for developers who need reliability, transparency, and control.

OpenAI-Compatible API

Drop-in replacement. Change base_url and api_key — that's it. Works with all OpenAI SDKs and tools.

AP-Singapore Deployment

All servers in Singapore. Low latency for Asian and international users.

API Key Management

Create, name, and revoke keys instantly. Track per-key usage. No limit on number of keys.

Predictable Performance

Dedicated capacity means no cold starts and no noisy neighbors. Consistent TTFT, every request.

Streaming with Usage

Full SSE streaming support. Usage reporting in final chunk for accurate billing.

Premium AI Models

Powered by Tencent Cloud TokenHub — DeepSeek, GLM, and Hunyuan. Enterprise-grade reliability.

Frequently Asked Questions

How does billing work?

You only pay for tokens you actually use. Input and output tokens are billed separately at different rates. We deduct from your prepaid balance after each request. No monthly fees, no minimums.

What's the difference between the three models?

Flash is our fastest model with sub-200ms TTFT, great for high-volume chat. Reason offers the best balance of speed, quality, and price for most production workloads. Pro is our highest quality model for complex reasoning tasks.

Is there a free trial?

Yes! New users get free credit to test all models. No credit card required — just sign up and start building.

What license are the models under?

All models are powered by Tencent Cloud TokenHub, providing access to DeepSeek V4 Flash, GLM 5.3, and Hunyuan 3. You can use the API output for commercial purposes with all plans.

Do you support streaming?

Yes, full SSE streaming support with usage reporting in the final chunk. Works seamlessly with the OpenAI SDK's stream parameter.

Where are the servers located?

All inference runs on servers in the AP-Singapore region. This provides low latency for Asian and international users.

Ready to Get Started?

Sign up for free credit. No credit card required. Ready in 2 minutes.

Create Free Account