Qwen3.8-27B API from EU hardware. Measured, not promised.
curl https://api.llmtech.eu/v1/chat/completions \
-H "Authorization: Bearer $LLMTECH_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/Qwen3.8-27B-NVFP4",
"messages": [{"role": "user", "content": "Hello"}],
"stream": true
}'
OpenAI-compatible: point your existing SDK at https://api.llmtech.eu/v1 and it works. Free trial key in the docs — no signup, no card.
Zero data retention
Prompts and completions live in volatile memory only. Never written to disk, logs, analytics or backups. Never used for training.
EU jurisdiction and hardware
Registered in Poland. German edge (Hetzner), Italian GPU datacentre (Seeweb, Frosinone). No US cloud anywhere in the stack. Who we are →
GDPR-ready
Signable Art. 28 DPA, published sub-processor list, documented technical measures. Security page →
How do I get an API key?
Try the free trial key from the docs right now. For a production key, email artem@llmtech.eu — issued the same day, usually within the hour. Self-service is planned.
How does billing work?
Per token, monthly invoice: $0.25/M input, $2.09/M output, $0.04/M cached input. No subscription, no minimums. Every request is journaled with its price snapshot, so an invoice is verifiable line by line.
Are your performance numbers real?
Time to first token and generation speed are measured on client traffic and published at llmtech.eu/status, refreshed every 5 minutes; uptime and the hourly strip come from a probe we send every minute. The speed figures on the model page come from a dated bench run, and the page says so. Nothing on this site is quoted from a datasheet.
Can I control the model's reasoning?
Yes: chat_template_kwargs.enable_thinking (on/off) and reasoning_effort (low / medium / xhigh). Thinking is off unless you ask for it with either parameter.
What about rate limits and SLA?
Up to 64 requests in flight on the public endpoint, shared by the keys on it. Above the limit you get an instant 429 with Retry-After, requests never queue silently. A channel of your own, with concurrency nobody else draws down, is available on request. Uptime history is public on the status page.