Per token. No subscription, no minimums.
| What | Price per 1M tokens | Notes |
|---|---|---|
| Input | $0.25 | everything you send |
| Cached input | $0.04 | repeated prompt prefixes, automatic — 6x cheaper |
| Output | $2.09 | completions, including reasoning tokens |
Billed monthly by invoice. Every request is journaled with its price snapshot — invoices are verifiable line by line. Committed volume, dedicated capacity and single-tenant EU deployments are quoted separately: artem@llmtech.eu.
The rates above are for direct accounts, and what they buy is not only tokens: a signed Art. 28 DPA, named sub-processors, a support path and a contractual counterparty. Marketplace and gateway traffic is cheaper because none of that is in place — we sign no DPA with the platform's users, there is no named customer behind the request, and nothing applies beyond our standard Terms. Data handling does not change: zero retention is how the service is built, on every channel. Anyone is welcome to buy at wholesale rates on that basis; with a DPA and a contract attached, the price is the one above. Our wholesale rate improves with delivered volume, and the same tier is open to any gateway that reaches it.
Nothing to subscribe to and no tiers to pick from: the rates above are what you pay as you go. These are three workload profiles with the arithmetic done, so you know what a month adds up to. Committed volume, dedicated capacity and EU-only deployments are quoted on their own terms.
If you run a chatbot
1M in + 100K out daily
30% cache hits
input $0.19 + output $0.21 per day
≈ $12/month
If you run an agent
3M in + 500K out daily
80% cache hits (repo context)
input $0.25 + output $1.05 per day
≈ $39/month
If you run bulk extraction
10M in + 300K out daily
50% cache hits
input $1.45 + output $0.63 per day
≈ $62/month
The cache rate is what makes agents and bulk workloads cheap here: repeated system prompts and document prefixes bill at $0.04 instead of $0.25.
| Model | Input per 1M tokens | Output |
|---|---|---|
| decider-4b-nvfp4 | $0.04 | free |
| decider-2b-fp8 | $0.03 | free |
| decider-0.8b-fp8 | $0.02 | free |
A decision model generates no text, so output tokens are 0 and a request bills its input tokens only, with the state counted once. 1,000 requests of 1,000 input tokens each on decider-2b-fp8 are 1M input tokens, $0.03. Up to 12 requests in flight per key, per model. What the models are: model page.