Changelog
What changed, and when.
2026-08-25
Output price lowered to $2.09/M (from $2.20). Per-key concurrency
limits with instant 429 + Retry-After semantics. Free shared trial key in docs.
Docs, pricing, security and platform pages published.
2026-08-24
Public model page with pricing. Unified pricing across all
channels. Terms updated with infrastructure-provider AUP clause. Listed in the
LiteLLM and models.dev catalogs (PRs).
2026-08-23
Public status page: hourly uptime strip, measured TTFT and
throughput from live traffic, refreshed every 5 minutes. Latency instrumentation
added to every request. Prompt caching verified and enabled on marketplace
traffic ($0.04/M cached input).
2026-08-22
Production launch: Qwen3.8-27B (NVFP4, 262,144-token context)
serving live marketplace traffic on NanoGPT. Zero-data-retention serving path,
per-request billing journal, automated health checks and recovery.