Service endingThe LLM Tech API stops on 31 October 2026. Keys stop working on 1 November 2026, 00:00 UTC. Our quantized models stay on Hugging Face: huggingface.co/llmtech
For routers & marketplaces

Built to be someone's upstream.

This page answers the questions platform integration teams actually ask. We serve live marketplace traffic today and our uptime is measured publicly — by us and by the platforms that route to us.

integration checklist
/chat/completions with SSE streamingyes, OpenAI format
usage in stream and non-streamyes, forced include_usage; cached_tokens reported
per-request id for billing reconciliationInference-Id header on every response from the model
tool calling / JSON mode / logprobsyes
reasoning control passthroughchat_template_kwargs respected per request
rate behaviourinstant 429 + Retry-After above limit — no silent queueing
per-platform channelsseparate keys, pricing and limits per platform
zero data retentionqualifies for ZDR-only routing
operations

Uptime, measured

Public status with an hourly uptime strip and measured TTFT/throughput, refreshed every 5 minutes: llmtech.eu/status. Health probes run every minute with automated recovery.

Restarts wait for a lull

Deployments wait until no request is in flight, and routine updates restart only the request layer, which is back in under a second. The model server itself is not touched.

Billing you can audit

Every request journaled with token counts and price snapshot. Usage reports in your format; per-request reconciliation by Inference-Id. Invoices verifiable line by line.

current serving
modelQwen3.8-27B (NVFP4 on Blackwell) — card
context262,144 tokens
published list price$0.25 / $2.09 / $0.04 cached, per 1M
regionEU only: Hetzner edge in Germany, Seeweb GPU node in Italy
live sinceAug 22, 2026 — serving marketplace traffic
Talk integration See live numbers first

Integration from our side typically takes a day. Test keys for your onboarding monitor issued immediately.