For routers & marketplaces

Built to be someone's upstream.

This page answers the questions platform integration teams actually ask. We serve live marketplace traffic today and our uptime is measured publicly — by us and by the platforms that route to us.

integration checklist
/chat/completions with SSE streamingyes, OpenAI format
usage in stream and non-streamyes, forced include_usage; cached_tokens reported
per-request id for billing reconciliationInference-Id header on every response
tool calling / JSON mode / logprobsyes
reasoning control passthroughchat_template_kwargs respected per request
rate behaviorinstant 429 + Retry-After above limit — no silent queueing
per-platform channelsseparate keys, pricing and limits per platform
zero data retentionqualifies for ZDR-only routing
operations

Uptime, measured

Public status with an hourly uptime strip and measured TTFT/throughput, refreshed every 5 minutes: llmtech.eu/status. Health probes run every minute with automated recovery.

Restarts drain first

Deployments wait for traffic lulls, connections drain before restart, and the model server itself is not touched on routine updates. Restart windows are seconds, not minutes.

Billing you can audit

Every request journaled with token counts and price snapshot. Usage reports in your format; per-request reconciliation by Inference-Id. Invoices verifiable line by line.

current serving
modelQwen3.8-27B (NVFP4 on Blackwell) — card
context262,144 tokens
wholesale pricing$0.25 / $2.09 / $0.04 cached, per 1M
regionEU only (DE edge + FI GPUs)
live sinceAug 22, 2026 — serving marketplace traffic
Talk integration See live numbers first

Integration from our side typically takes a day. Test keys for your onboarding monitor issued immediately.