Changelog
What changed, and when.
2026-09-30
Scheduled maintenance, 18:00 to 20:00 UTC: the chat model now runs with
multi-token prediction and an FP8 key-value cache. Errors in this period came from
requests sent during the scheduled maintenance window.
2026-09-30
New endpoint POST /v1/systemone for decision models, with the
same API key as chat. Three models: decider-4b-nvfp4 at $0.04, decider-2b-fp8 at
$0.03 and decider-0.8b-fp8 at $0.02 per million input tokens. Output is free: a
decision model generates no text, and a request bills its input tokens only. A
state of up to 32,768 tokens fits with any question of up to 4,096 tokens, and a
key can have up to 12 requests in flight per model.
The models are Mapika/decider (Apache 2.0), quantised and measured by LLM Tech.
Accuracy against bf16 and latency measured on an idle GPU are on the new
/models/decider/ page. The cabinet and the monthly invoice show decision models
separately, by model.
2026-09-26
Seeweb in Frosinone, Italy, is now the permanent site for GPU inference.
The move announced on 19 September will not happen, and the sub-processor list
stays as it is.
2026-09-24
Corrections to what the site said. Reasoning is off unless a
request asks for it with reasoning_effort or enable_thinking; the docs had said the
model decides by itself. The public trial key is shared, and so is its prompt
cache. TLS ends at our edge and a WireGuard tunnel carries traffic to the GPU node;
the model page had said TLS end to end. On the status page, latency and speed come
from client traffic and uptime from a probe sent every minute.
2026-09-22
The trial key now allows 4 concurrent requests instead of 2,
still shared by everyone using it, still 2 million tokens a day per address. The
docs now describe the prompt cache as it has worked on this node since 18 September:
784-token blocks, read from cache from the second request on.
2026-09-21
A monthly spending cap per key, on request. Past it the API
answers 402 with the code monthly_cap_reached until the 1st of the next month,
00:00 UTC.
2026-09-19
The site now says which site is current. GPU inference runs in
Italy at Seeweb, whose Milan and Frosinone locations hold ISO 27001:2022, 27017
and 27018; the edge stays in Nuremberg. This location is temporary: a move to
another EU location is planned, and the date and the new sub-processor will be
announced before it happens rather than after. (Superseded on 26 September:
Seeweb in Frosinone is the permanent site for GPU inference, and the move will
not happen.)
2026-09-18
Service is back after the maintenance window, on new hardware:
one RTX PRO 6000 Blackwell with 96 GB, in Italy. The engine moved from a nightly
vLLM build to the 0.29.0 stable release, which closes the risk named in the entry
of 28 August, and a second cache tier of 64 GiB now lives in host RAM. Measured
on the card in service that day, with nothing else running on it: 61.6 tok/s on
a single stream, 50.7 at twelve concurrent, 32.5 at thirty-two; 0.17 s to first
token; 13,190 tok/s of fresh prefill; a cached prompt comes back 19.3 times
faster than the same prompt cold. The public model id is now
nvidia/Qwen3.8-27B-NVFP4, and the earlier id unsloth/Qwen3.8-27B-NVFP4 keeps
working as an alias, so nothing built against it breaks. The prompt cache is
isolated per key and was verified on this node: a prefix cached under one key
returns zero cached tokens for another. min_p and logit_bias are now
accepted.
2026-09-15
New /personal/ page for people buying for themselves, and a
cabinet that shows a prepaid balance. A balance is topped up by bank transfer,
credited the same working day, refundable and does not expire. When it runs out
the API answers 402 with the type insufficient_balance instead of running up a
debt.
2026-09-13
Trouble on the edge between 11 and 13 September: a configuration
change of ours made requests with large bodies fail. Found and fixed. Request
bodies now stream straight through to the backend and never touch the disk of
the edge. Planned maintenance followed the same morning: the API answered 503
with Retry-After while the service moved to new hardware, and it returned on 18
September.
2026-09-10
The output ceiling of 32,768 tokens that we publish in /v1/models
is now enforced, not just advertised. Until this change a request could ask for
more and get it: two streams in the small hours of 10 September ran into the
upstream time limit and ended as 504, and the client paid for nothing because an
aborted stream is not billed. Over 5 to 10 September, 10 requests out of 39,799
had gone past the published ceiling. A larger max_tokens or
max_completion_tokens is now lowered to the ceiling, and a request that names
neither gets the ceiling as its limit, so the answer stops with finish_reason
length instead of failing.
2026-09-08
LLM Tech is now an integration page in Haystack, deepset's
open-source framework: the page shows how to call the endpoint through the
standard OpenAI generators and a pipeline, with notes on prompt caching. The
pull request was reviewed and merged by the deepset team the same day.
2026-09-06
New /demo/ page: drop a scanned page, a photo of a form or a PDF
up to five pages and get one JSON object back against a fixed schema, from the
same endpoint the paying traffic runs on. Nothing is stored; a PDF is split
into page images in memory and discarded when the request ends. Five
documents per address per day.
2026-08-31
Pricing is set per channel, not as one rate everywhere. The
published list price is what a direct account pays. Gateways and marketplaces
are quoted separately — they carry their own margin and their own support
path — and committed volume is quoted on its own terms. Every request is
journaled with the price that applied at that moment, so an invoice can be
checked line by line at the rate that was actually in force. Two billing
corrections: invoices no longer include our own smoke-test requests, and
client-aborted requests (HTTP 499) are no longer counted as service errors —
that is the client closing the connection, not a failure on our side.
2026-08-30
The site is now machine-readable: structured data on all
public pages under a single organisation identifier, /llms.txt and
/llms-full.txt, a sitemap of sixteen URLs, and a robots.txt that names nineteen
crawlers and disallows none. New /agents/ page with working configurations for
nine tools — Cline, Kilo Code, Continue, aider, Zed, Open WebUI, LibreChat,
LiteLLM and Claude Code. Image input is now declared in /v1/models; it already
worked, it is now advertised. Model page corrected: it promised a median TTFT
under one second and 80 tok/s, which live measurement did not support. It now publishes measured figures with the date of the measurement. The billing journal now fsyncs on write and is mirrored to a
second machine every minute.
2026-08-28
Incident: the inference engine crashed at 13:26 UTC after
five days of continuous load (vLLM nightly bug; no OOM or GPU fault).
Automated health checks detected it and restarted the node — roughly
7 minutes of degraded availability, no billing data lost. The prefix
cache rebuilt from live traffic within minutes. Migration to the stable
vLLM release is planned to remove the nightly-build risk.
2026-08-25
Output price lowered to $2.09/M (from $2.20). Per-key concurrency
limits with instant 429 + Retry-After semantics. Free shared trial key in docs.
Docs, pricing, security and platform pages published.
2026-08-24
Public model page with pricing. Terms updated with
infrastructure-provider AUP clause. Pull requests opened to add us to the
LiteLLM and models.dev catalogues.
2026-08-23
Public status page: hourly uptime strip, measured TTFT and
throughput from live traffic, refreshed every 5 minutes. Latency instrumentation
added to every request. Prompt caching verified and enabled on marketplace
traffic ($0.04/M cached input).
2026-08-22
Production launch: Qwen3.8-27B (NVFP4, 262,144-token context)
serving live marketplace traffic on NanoGPT. Zero-data-retention serving path,
per-request billing journal, automated health checks and recovery.