Service endingThe LLM Tech API stops on 31 October 2026. Keys stop working on 1 November 2026, 00:00 UTC. Our quantized models stay on Hugging Face: huggingface.co/llmtech
Inference provider · EU

Qwen3.8-27B API from EU hardware. Measured, not promised.

262K context · zero data retention · EU end to end
TTFT and speed from client traffic, uptime from a probe every minute; refreshes every 5 min · full status
Serving live traffic since Aug 22, 2026 — currently on NanoGPT
try it now
curl https://api.llmtech.eu/v1/chat/completions \
  -H "Authorization: Bearer $LLMTECH_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nvidia/Qwen3.8-27B-NVFP4",
    "messages": [{"role": "user", "content": "Hello"}],
    "stream": true
  }'

OpenAI-compatible: point your existing SDK at https://api.llmtech.eu/v1 and it works. Free trial key in the docs — no signup, no card.

the model
Qwen3.8-27B · NVFP4 on Blackwell · full card →
context window262,144 tokens
reasoning controlenable_thinking · reasoning_effort
tool calling / JSON modeyes
prompt cachingautomatic, $0.04/M
concurrencyup to 64 in flight, shared on the public endpoint
quantisationNVFP4 — hardware format, near-fp8 quality
decision models
Decider · three models · full card →
what it doesanswers typed questions about a state, no text generation
endpointPOST /v1/systemone, the same API key as chat
pricing$0.02 to $0.04/M input, output free
eu, by construction

Zero data retention

Prompts and completions live in volatile memory only. Never written to disk, logs, analytics or backups. Never used for training.

EU jurisdiction and hardware

Registered in Poland. German edge (Hetzner), Italian GPU datacentre (Seeweb, Frosinone). No US cloud anywhere in the stack. Who we are →

GDPR-ready

Signable Art. 28 DPA, published sub-processor list, documented technical measures. Security page →

faq
How do I get an API key?

Try the free trial key from the docs right now. For a production key, email artem@llmtech.eu — issued the same day, usually within the hour. Self-service is planned.

How does billing work?

Per token, monthly invoice: $0.25/M input, $2.09/M output, $0.04/M cached input. No subscription, no minimums. Every request is journaled with its price snapshot, so an invoice is verifiable line by line.

Are your performance numbers real?

Time to first token and generation speed are measured on client traffic and published at llmtech.eu/status, refreshed every 5 minutes; uptime and the hourly strip come from a probe we send every minute. The speed figures on the model page come from a dated bench run, and the page says so. Nothing on this site is quoted from a datasheet.

Can I control the model's reasoning?

Yes: chat_template_kwargs.enable_thinking (on/off) and reasoning_effort (low / medium / xhigh). Thinking is off unless you ask for it with either parameter.

What about rate limits and SLA?

Up to 64 requests in flight on the public endpoint, shared by the keys on it. Above the limit you get an instant 429 with Retry-After, requests never queue silently. A channel of your own, with concurrency nobody else draws down, is available on request. Uptime history is public on the status page.

Start with the free trial key