Service endingThe LLM Tech API stops on 31 October 2026. Keys stop working on 1 November 2026, 00:00 UTC. Our quantized models stay on Hugging Face: huggingface.co/llmtech
Pricing

Per token. No subscription, no minimums.

$0.25/M input · $2.09/M output · $0.04/M cached input
Qwen3.8-27B · 262,144-token context
WhatPrice per 1M tokensNotes
Input$0.25everything you send
Cached input$0.04repeated prompt prefixes, automatic — 6x cheaper
Output$2.09completions, including reasoning tokens

Billed monthly by invoice. Every request is journaled with its price snapshot — invoices are verifiable line by line. Committed volume, dedicated capacity and single-tenant EU deployments are quoted separately: artem@llmtech.eu.

The rates above are for direct accounts, and what they buy is not only tokens: a signed Art. 28 DPA, named sub-processors, a support path and a contractual counterparty. Marketplace and gateway traffic is cheaper because none of that is in place — we sign no DPA with the platform's users, there is no named customer behind the request, and nothing applies beyond our standard Terms. Data handling does not change: zero retention is how the service is built, on every channel. Anyone is welcome to buy at wholesale rates on that basis; with a DPA and a contract attached, the price is the one above. Our wholesale rate improves with delivered volume, and the same tier is open to any gateway that reaches it.

worked examples — not plans

Nothing to subscribe to and no tiers to pick from: the rates above are what you pay as you go. These are three workload profiles with the arithmetic done, so you know what a month adds up to. Committed volume, dedicated capacity and EU-only deployments are quoted on their own terms.

If you run a chatbot

1M in + 100K out daily
30% cache hits

input $0.19 + output $0.21 per day

≈ $12/month

If you run an agent

3M in + 500K out daily
80% cache hits (repo context)

input $0.25 + output $1.05 per day

≈ $39/month

If you run bulk extraction

10M in + 300K out daily
50% cache hits

input $1.45 + output $0.63 per day

≈ $62/month

The cache rate is what makes agents and bulk workloads cheap here: repeated system prompts and document prefixes bill at $0.04 instead of $0.25.

decisions
Decider decision models · POST /v1/systemone
ModelInput per 1M tokensOutput
decider-4b-nvfp4$0.04free
decider-2b-fp8$0.03free
decider-0.8b-fp8$0.02free

A decision model generates no text, so output tokens are 0 and a request bills its input tokens only, with the state counted once. 1,000 requests of 1,000 input tokens each on decider-2b-fp8 are 1M input tokens, $0.03. Up to 12 requests in flight per key, per model. What the models are: model page.

Try the free trial key Get a production key