Three tiers of AI models with enterprise-grade reliability. OpenAI-compatible API — drop-in replacement, zero code changes.
from openai import OpenAI
client = OpenAI(
api_key="sk-zacht-...",
base_url="https://zachie.site/api/v1"
)
response = client.chat.completions.create(
model="zachai-reason",
messages=[...]
)100% OpenAI-compatible. Change 2 lines, keep everything else.
Pick the right model for the job. All powered by Tencent Cloud TokenHub, all with 128K context windows.
Ultra-fast, sub-200ms TTFT
Balanced reasoning & code
Premium quality, deep reasoning
Pay only for what you use. No minimums, no monthly fees, no surprises.
| Model | Input / 1M tokens | Output / 1M tokens | Context |
|---|---|---|---|
Flashzachai-flash DeepSeek V4 Flash | $0.08 | $0.18 | 128K |
Reasonzachai-reason GLM 5.3 | $0.30 | $0.65 | 128K |
Prozachai-pro Hunyuan 3 | $0.75 | $1.60 | 128K |
Prices in USD. Volume discounts available for enterprise customers.
Everything you need to ship AI features to production, from day one.
Drop-in replacement for the OpenAI SDK. Just change base_url and api_key — no other code changes needed.
All inference runs on Tencent Cloud TokenHub. Enterprise-grade infrastructure with global CDN access.
Create, name, and revoke API keys instantly. Track per-key usage and set rate limits per plan.
Dedicated capacity means consistent TTFT and throughput. No cold starts, no noisy-neighbor slowdowns.
Full SSE streaming with usage reporting. Real-time token-by-token responses for chat experiences.
No monthly minimums, no long-term contracts. Prepaid balance, pay only for tokens you actually use.
Sign up for free and get instant access to all three models. No credit card required.