Changelog

What changed, and when.

2026-08-25 Output price lowered to $2.09/M (from $2.20). Per-key concurrency limits with instant 429 + Retry-After semantics. Free shared trial key in docs. Docs, pricing, security and platform pages published.
2026-08-24 Public model page with pricing. Unified pricing across all channels. Terms updated with infrastructure-provider AUP clause. Listed in the LiteLLM and models.dev catalogs (PRs).
2026-08-23 Public status page: hourly uptime strip, measured TTFT and throughput from live traffic, refreshed every 5 minutes. Latency instrumentation added to every request. Prompt caching verified and enabled on marketplace traffic ($0.04/M cached input).
2026-08-22 Production launch: Qwen3.8-27B (NVFP4, 262,144-token context) serving live marketplace traffic on NanoGPT. Zero-data-retention serving path, per-request billing journal, automated health checks and recovery.