Service endingThe LLM Tech API stops on 31 October 2026. Keys stop working on 1 November 2026, 00:00 UTC. Our quantized models stay on Hugging Face: huggingface.co/llmtech
Changelog

What changed, and when.

2026-09-30 Scheduled maintenance, 18:00 to 20:00 UTC: the chat model now runs with multi-token prediction and an FP8 key-value cache. Errors in this period came from requests sent during the scheduled maintenance window.
2026-09-30 New endpoint POST /v1/systemone for decision models, with the same API key as chat. Three models: decider-4b-nvfp4 at $0.04, decider-2b-fp8 at $0.03 and decider-0.8b-fp8 at $0.02 per million input tokens. Output is free: a decision model generates no text, and a request bills its input tokens only. A state of up to 32,768 tokens fits with any question of up to 4,096 tokens, and a key can have up to 12 requests in flight per model. The models are Mapika/decider (Apache 2.0), quantised and measured by LLM Tech. Accuracy against bf16 and latency measured on an idle GPU are on the new /models/decider/ page. The cabinet and the monthly invoice show decision models separately, by model.
2026-09-26 Seeweb in Frosinone, Italy, is now the permanent site for GPU inference. The move announced on 19 September will not happen, and the sub-processor list stays as it is.
2026-09-24 Corrections to what the site said. Reasoning is off unless a request asks for it with reasoning_effort or enable_thinking; the docs had said the model decides by itself. The public trial key is shared, and so is its prompt cache. TLS ends at our edge and a WireGuard tunnel carries traffic to the GPU node; the model page had said TLS end to end. On the status page, latency and speed come from client traffic and uptime from a probe sent every minute.
2026-09-22 The trial key now allows 4 concurrent requests instead of 2, still shared by everyone using it, still 2 million tokens a day per address. The docs now describe the prompt cache as it has worked on this node since 18 September: 784-token blocks, read from cache from the second request on.
2026-09-21 A monthly spending cap per key, on request. Past it the API answers 402 with the code monthly_cap_reached until the 1st of the next month, 00:00 UTC.
2026-09-19 The site now says which site is current. GPU inference runs in Italy at Seeweb, whose Milan and Frosinone locations hold ISO 27001:2022, 27017 and 27018; the edge stays in Nuremberg. This location is temporary: a move to another EU location is planned, and the date and the new sub-processor will be announced before it happens rather than after. (Superseded on 26 September: Seeweb in Frosinone is the permanent site for GPU inference, and the move will not happen.)
2026-09-18 Service is back after the maintenance window, on new hardware: one RTX PRO 6000 Blackwell with 96 GB, in Italy. The engine moved from a nightly vLLM build to the 0.29.0 stable release, which closes the risk named in the entry of 28 August, and a second cache tier of 64 GiB now lives in host RAM. Measured on the card in service that day, with nothing else running on it: 61.6 tok/s on a single stream, 50.7 at twelve concurrent, 32.5 at thirty-two; 0.17 s to first token; 13,190 tok/s of fresh prefill; a cached prompt comes back 19.3 times faster than the same prompt cold. The public model id is now nvidia/Qwen3.8-27B-NVFP4, and the earlier id unsloth/Qwen3.8-27B-NVFP4 keeps working as an alias, so nothing built against it breaks. The prompt cache is isolated per key and was verified on this node: a prefix cached under one key returns zero cached tokens for another. min_p and logit_bias are now accepted.
2026-09-15 New /personal/ page for people buying for themselves, and a cabinet that shows a prepaid balance. A balance is topped up by bank transfer, credited the same working day, refundable and does not expire. When it runs out the API answers 402 with the type insufficient_balance instead of running up a debt.
2026-09-13 Trouble on the edge between 11 and 13 September: a configuration change of ours made requests with large bodies fail. Found and fixed. Request bodies now stream straight through to the backend and never touch the disk of the edge. Planned maintenance followed the same morning: the API answered 503 with Retry-After while the service moved to new hardware, and it returned on 18 September.
2026-09-10 The output ceiling of 32,768 tokens that we publish in /v1/models is now enforced, not just advertised. Until this change a request could ask for more and get it: two streams in the small hours of 10 September ran into the upstream time limit and ended as 504, and the client paid for nothing because an aborted stream is not billed. Over 5 to 10 September, 10 requests out of 39,799 had gone past the published ceiling. A larger max_tokens or max_completion_tokens is now lowered to the ceiling, and a request that names neither gets the ceiling as its limit, so the answer stops with finish_reason length instead of failing.
2026-09-08 LLM Tech is now an integration page in Haystack, deepset's open-source framework: the page shows how to call the endpoint through the standard OpenAI generators and a pipeline, with notes on prompt caching. The pull request was reviewed and merged by the deepset team the same day.
2026-09-06 New /demo/ page: drop a scanned page, a photo of a form or a PDF up to five pages and get one JSON object back against a fixed schema, from the same endpoint the paying traffic runs on. Nothing is stored; a PDF is split into page images in memory and discarded when the request ends. Five documents per address per day.
2026-08-31 Pricing is set per channel, not as one rate everywhere. The published list price is what a direct account pays. Gateways and marketplaces are quoted separately — they carry their own margin and their own support path — and committed volume is quoted on its own terms. Every request is journaled with the price that applied at that moment, so an invoice can be checked line by line at the rate that was actually in force. Two billing corrections: invoices no longer include our own smoke-test requests, and client-aborted requests (HTTP 499) are no longer counted as service errors — that is the client closing the connection, not a failure on our side.
2026-08-30 The site is now machine-readable: structured data on all public pages under a single organisation identifier, /llms.txt and /llms-full.txt, a sitemap of sixteen URLs, and a robots.txt that names nineteen crawlers and disallows none. New /agents/ page with working configurations for nine tools — Cline, Kilo Code, Continue, aider, Zed, Open WebUI, LibreChat, LiteLLM and Claude Code. Image input is now declared in /v1/models; it already worked, it is now advertised. Model page corrected: it promised a median TTFT under one second and 80 tok/s, which live measurement did not support. It now publishes measured figures with the date of the measurement. The billing journal now fsyncs on write and is mirrored to a second machine every minute.
2026-08-28 Incident: the inference engine crashed at 13:26 UTC after five days of continuous load (vLLM nightly bug; no OOM or GPU fault). Automated health checks detected it and restarted the node — roughly 7 minutes of degraded availability, no billing data lost. The prefix cache rebuilt from live traffic within minutes. Migration to the stable vLLM release is planned to remove the nightly-build risk.
2026-08-25 Output price lowered to $2.09/M (from $2.20). Per-key concurrency limits with instant 429 + Retry-After semantics. Free shared trial key in docs. Docs, pricing, security and platform pages published.
2026-08-24 Public model page with pricing. Terms updated with infrastructure-provider AUP clause. Pull requests opened to add us to the LiteLLM and models.dev catalogues.
2026-08-23 Public status page: hourly uptime strip, measured TTFT and throughput from live traffic, refreshed every 5 minutes. Latency instrumentation added to every request. Prompt caching verified and enabled on marketplace traffic ($0.04/M cached input).
2026-08-22 Production launch: Qwen3.8-27B (NVFP4, 262,144-token context) serving live marketplace traffic on NanoGPT. Zero-data-retention serving path, per-request billing journal, automated health checks and recovery.