Per token. No subscription, no minimums, no "contact sales".
| What | Price per 1M tokens | Notes |
|---|---|---|
| Input | $0.25 | everything you send |
| Cached input | $0.04 | repeated prompt prefixes, automatic — 6x cheaper |
| Output | $2.09 | completions, including reasoning tokens |
Billed monthly by invoice. Every request is journaled with its price snapshot — invoices are verifiable line by line. Volume pricing for sustained load: artem@llmtech.eu.
There are no tiers and no subscriptions: everyone pays the same per-token rates above. These are just three workload profiles with the arithmetic done, so you know what a month adds up to.
If you run a chatbot
1M in + 100K out daily
30% cache hits
input $0.19 + output $0.21 per day
≈ $12/month
If you run an agent
3M in + 500K out daily
80% cache hits (repo context)
input $0.25 + output $1.05 per day
≈ $39/month
If you run bulk extraction
10M in + 300K out daily
50% cache hits
input $1.45 + output $0.63 per day
≈ $62/month
The cache rate is what makes agents and bulk workloads cheap here: repeated system prompts and document prefixes bill at $0.04 instead of $0.25.
| Provider | Input, $/M | Output, $/M |
|---|---|---|
| LLM Tech | $0.25 | $2.09 |
| Chutes | $0.35 | $2.75 |
| DeepInfra | $0.40 | $3.00 |
| Alibaba Cloud | $0.50 | $3.00 |
prices checked 2026-08-25 on public pricing pages; we intend to stay the lowest for this model.