First response in under a minute.
The API is OpenAI-compatible. Base URL https://api.llmtech.eu/v1, model nvidia/Qwen3.8-27B-NVFP4. Below is a shared free trial key — try before you talk to anyone.
lt-trial-ba1ef28c6d32ed6980678d8d
Shared and rate-limited: 4 concurrent requests and 2M tokens per day per address, counted across prompt and completion and reset at 00:00 UTC. Enough to evaluate, not enough to run on. For production, email artem@llmtech.eu — a personal key with no daily limit is issued the same day, usually within the hour; it shares the endpoint's 64 requests in flight with other keys. Working with scans or PDFs? The document demo runs the same endpoint with no key at all.
curl https://api.llmtech.eu/v1/chat/completions \
-H "Authorization: Bearer lt-trial-ba1ef28c6d32ed6980678d8d" \
-H "Content-Type: application/json" \
-d '{
"model": "nvidia/Qwen3.8-27B-NVFP4",
"messages": [{"role": "user", "content": "Say hello"}],
"stream": true
}'
from openai import OpenAI
client = OpenAI(
base_url="https://api.llmtech.eu/v1",
api_key="lt-trial-ba1ef28c6d32ed6980678d8d",
)
r = client.chat.completions.create(
model="nvidia/Qwen3.8-27B-NVFP4",
messages=[{"role": "user", "content": "Say hello"}],
stream=True,
)
for chunk in r:
print(chunk.choices[0].delta.content or "", end="")
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.llmtech.eu/v1",
apiKey: "lt-trial-ba1ef28c6d32ed6980678d8d",
});
const stream = await client.chat.completions.create({
model: "nvidia/Qwen3.8-27B-NVFP4",
messages: [{ role: "user", content: "Say hello" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
Reasoning is off by default. Turn it on per request with reasoning_effort or enable_thinking:
{
"model": "nvidia/Qwen3.8-27B-NVFP4",
"messages": [{"role": "user", "content": "Is 1001 a prime number?"}],
"chat_template_kwargs": {
"enable_thinking": true
}
}
Reasoning text comes back in reasoning (non-streaming) or reasoning_content deltas (streaming). Reasoning tokens are billed as output.
Automatic, no code changes. Repeated prompt prefixes are billed at $0.04/M instead of $0.25/M. The first request writes the cache and the second already reads it. The cache works in blocks of 1,600 tokens and reuses every whole block except the last one, so a prompt under 3,200 tokens gets none. Check usage.prompt_tokens_details.cached_tokens in the response to see it working.
The cache belongs to your key. Two of your own keys do not share a cached prefix, and you never share one with anybody else. Inside one key you can split it yourself by sending your own cache_salt. The one exception is the public trial key. Everyone who uses it shares the key, so its cache is split by source IP address: guests behind one address, such as an office network or a VPN, share a cache. Send nothing confidential with it.
| Code | Meaning | What to do |
|---|---|---|
| 401 | invalid or missing API key | check the Authorization header |
| 402 | daily limit reached, or prepaid balance used up | the error type says which: insufficient_quota for a limit, insufficient_balance for a balance; limits reset at 00:00 UTC |
| 402 | monthly spending cap on your key reached | code monthly_cap_reached; resets on the 1st of the month at 00:00 UTC, or ask us to change the cap |
| 403 | this key may not generate | a read-only cabinet key cannot spend tokens; use your working key |
| 404 | unknown model id | use nvidia/Qwen3.8-27B-NVFP4; the earlier id unsloth/Qwen3.8-27B-NVFP4 still works |
| 413 | request body larger than 20 MB | send images at a lower resolution or split the request |
| 429 | concurrency limit reached | retry after the Retry-After header, 1 or 2 seconds; requests never queue silently |
| 400 | malformed request / context overflow | the error message names the exact problem |
| 5xx | server-side failure | retry with backoff; check status |
Default limits: up to 64 requests in flight on the public endpoint, shared by the keys on it; the trial key allows 4 for everyone using it together and a context of 131,072 tokens. Other keys get the full 262,144. Need concurrency that is yours alone, ask: that is what a channel of your own is for.
Responses of the API, errors included, carry X-Inference-Location: IT and X-Inference-Site: Frosinone: the country (ISO 3166-1 alpha-2) and the site of the GPU node that runs the request; a streamed answer carries them too. They are left out where they would not describe that node: on responses the front proxy produces itself (its 429 for too many connections or requests, its 413 for a body over 20 MB, its errors while the API behind it is unreachable, the static list at /openrouter/models and the 404 for /v1/admin/), on a 400 for a request that is not valid HTTP, and on the response to a request served by a GPU outside this node.
A few things behave differently from a plain OpenAI endpoint, and one used to. We tested each one against production rather than guessing, and we would rather you read it here than discover it at three in the morning.
The last one is worth a sentence. Requests should use the public id nvidia/Qwen3.8-27B-NVFP4, and the earlier id unsloth/Qwen3.8-27B-NVFP4 keeps working as an alias so that nothing built against it breaks; anything else is a 404 — but the response and the stream chunks currently carry the internal served name. Harmless for every client we have seen, since nothing routes on it, but if your code asserts that the response model equals the requested one, that assertion will fail. Say so and we will change it.
Everything else you would expect works, including max_completion_tokens, seed, logprobs, response_format with a JSON schema, tool calling with the standard string-encoded arguments, the developer role, and the sampler set temperature / top_p / top_k / repetition_penalty.
POST https://api.llmtech.eu/v1/systemone answers typed questions about a state with the Decider decision models. It takes the same Bearer key as chat completions and is enabled per key. There is no text generation: every question is read in its own prompt row (a score question by default in one row per level), each row in one forward pass, and only input tokens are billed.
Output is free. GET /v1/models with a key that has decision models enabled lists the three models with "type": "decisions" and "endpoint": "/v1/systemone"; without a key, or with a key that does not have them, they are not listed. To have decision models enabled for a key, email artem@llmtech.eu.
The model server keeps a prefix cache per key: an identical state sent again on the same key can be served faster, and the cache is never shared between keys, so no key can tell from timing that another key has sent the same state.
The body is one JSON object of at most 2,097,152 bytes. Content-Type is not checked.
A body may be compressed with Content-Encoding: gzip or deflate (zlib-wrapped or raw deflate). The API decompresses it itself; the 2,097,152-byte limit applies both to the body as sent, with or without Content-Length (chunked), and to the decompressed body (413 payload_too_large), and a body that cannot be decompressed, including a truncated stream, is refused with 400 invalid_request. A gzip body may be several gzip members one after another, at most 64; a body with more is refused with 400 invalid_request. Content-Encoding: br and zstd are not supported and are refused with 400 invalid_request. Any other Content-Encoding value is ignored: the body is read as it was sent.
| Field | Required | Type | Meaning |
|---|---|---|---|
| model | yes | string | one of the three model ids above |
| state | yes | any JSON value except null | what the questions are about |
| questions | yes | object, at least one entry | question id → question |
| independent | no | true (or null, the same as absent) | every question is answered on its own; false is refused |
| layout | no | any | ignored |
Any other top-level field is ignored and not passed to the model server. Arrays and objects may nest at most 128 levels deep in the whole body, the body itself being level 1 (so a state may nest 127 levels). The body may hold at most 500,000 { and [ characters and at most 1,000,000 JSON values, counted as ,, { and [ characters together; both are counted in the (decompressed) body before it is parsed, so the characters inside strings count too. These are hard limits on the body, not on what the model can answer: a body over either is refused with 400 invalid_request even when the model could answer the request. For scale, 1,024 questions of 255 options each hold about 264,000 such values, and 2 MiB of single-digit numbers about 1,048,000. Every string and object key in state and questions must be valid Unicode: an unpaired UTF-16 surrogate (a \ud800 to \udfff escape without its partner) is refused with 400 invalid_request, and the message says where (state or question #N). A proper pair, such as an emoji written as two escapes, is fine.
A string is used as it is. Any other value is rendered as JSON text: separators ", " and ": ", non-ASCII characters kept, and every array of 8 or more elements rewritten so that each element carries its position: an object element e at index i becomes {"_index": i, ...e} (a key _index already in e wins), any other element becomes {"_index": i, "value": e}. Arrays of fewer than 8 elements keep their elements as they are (arrays nested inside them follow the same rule). Numbers, true and false become their JSON text (42, true).
The model reads "Context:\n" followed by that text. Nothing is ever cut from the state: a state that does not fit is refused with 413 (see limits below).
An object from your question id to a question. Ids are any strings; they are never shown to the model. Answers come back under the same ids, in the same order. Every question is an object with these fields:
| Field | choice | score | noul |
|---|---|---|---|
| type | "choice" (the default when type is absent) | "score" | "noul"; "bool" is accepted as the same |
| instructions | required | required | optional if criteria describes "true" or "false" |
| criteria | required: 2 to 255 options | required: 2 to 10 levels | optional |
| isolated | ignored | optional, true (default) or false | ignored |
instructions is the question in words: a non-empty string, or any other JSON value except null, which is sent to the model as its JSON text. question is accepted as a synonym of instructions and options as a synonym of criteria; when both names are present, instructions and criteria win. Other fields in a question are ignored.
choice: criteria is either an object from option name to description, or an array of distinct option names (strings). A description may be null or "" (the option is shown by its name only), a string, or any other JSON value (shown as JSON text). The model sees each option as name or name: description, in the order you give. The answer names the chosen option by its name.
score: criteria is either an array of level descriptions, lowest first, or an object whose keys are distinct decimal numbers ("0", "1", "2.5", "-1"; no exponent, no leading zeros, no spaces). The object form is sorted by the numeric value of its keys and the levels are then renumbered 0, 1, 2, …: the answer uses these numbers, not your keys. A level description is a string or any JSON value (shown as JSON text).
By default every level is judged alone, in its own row: the model sees the question, one level ("Proposed answer: …", with a leading N: removed from the level text) and answers whether it fits. With "isolated": false all levels are shown together as options 0: …, 1: … in one row.
noul: a yes/no question answered with the probability of yes. criteria, if given, is an object with optional "true" and "false" descriptions (other keys are ignored); the model sees the options as no / no: <false description> and yes / yes: <true description>. Without instructions (absent, null or "") the question text is the fixed sentence "Which answer fits the context?", and at least one of the two descriptions must be present and neither null nor "".
The model server interprets a few inputs in Python's way. We refuse them with 400 instead, so that no request is answered in a way its author did not intend:
| Refused with 400 | What the model server would do |
|---|---|
| state: null | read the text null |
| instructions: null in a choice or score question | ask the question "null" |
| a choice criteria array with a non-string element or a repeated name | convert elements with Python's str(), so true becomes True, and merge repeats |
| score criteria object keys that are not plain decimal numbers or are equal as numbers ("01", "1e1", " 1", "nan", "1_0", "0" with "0.0") | accept anything Python's float() reads |
| isolated that is not true or false | take its truth value, so "false" would mean true |
| independent that is not true, false or null | coerce "yes", 1 and similar |
200 OK, Content-Type: application/json, header Inference-Id (see headers below). The body:
{"model": "<the model id you sent>",
"answers": {"<question id>": {...}, ...},
"usage": {"input_tokens": <integer>, "output_tokens": 0}}
answers holds exactly one answer per question id, in the order of questions. Fields per answer, in this order (probabilities and derived values are rounded to 4 decimals, score to 2, each value on its own, so probabilities may not add up to exactly 1):
| Field | Value |
|---|---|
| type | "choice" |
| choice | the option name with the highest probability (the first one on a tie) |
| confidence | (n·p_max − 1)/(n − 1) clipped to [0, 1]: 0 for a uniform distribution, 1 when one option has all of it |
| x_p_max | the highest probability, p_max |
| certainty | 1 − H(p)/ln n floored at 0, where H is the entropy in nats |
| probabilities | option name → probability, in the order of criteria |
| Field | Value |
|---|---|
| type | "score" |
| score | the expected level Σ i·pᵢ over levels 0 … n−1 |
| confidence | max(0, 1 − Σ pᵢ·|i − k| / D), k the most likely level, D = (1/n)·Σ |i − (n−1)/2| |
| x_p_max | the probability of the most likely level |
| certainty | as for choice |
| legend | "0", "1", … → the level description as text (non-string descriptions as JSON text) |
| probabilities | "0", "1", … → probability of the level |
| level_fit | only with isolated levels (the default): "0", "1", … → the probability that this level fits, judged alone |
| fit_mass | only with isolated levels: the sum of level_fit; near 1 when one level fits, low when none does, above 1 when several do. probabilities is level_fit divided by it |
| Field | Value |
|---|---|
| type | "noul" |
| noul | the probability of yes |
output_tokens is always 0. input_tokens is what the request is billed for: every distinct token position the model reads, with the state counted once.
A request becomes scoring rows: one per question, and one per level of a score question with isolated levels. Each row is the same context ("Context:\n" plus the rendered state, C tokens) followed by that row's question block (Qᵢ tokens: the question, its options and the answer slot). The question blocks all begin with the same few tokens (L, the length of their longest common beginning; 0 for a single row). Then
input_tokens = C + L + Σᵢ (Qᵢ − L)
For a request with one choice or noul question, input_tokens is the length of its row. The same number is what your usage and your invoice count for the request; the price is per 1M of them.
| Limit | Value | Over it |
|---|---|---|
| request body, as sent and after decompression | 2,097,152 bytes | 413 payload_too_large |
| JSON nesting in the body | 128 levels | 400 invalid_request |
| { and [ in the body, also inside strings | 500,000 | 400 invalid_request |
| JSON values, counted as , { [ in the body, also inside strings | 1,000,000 | 400 invalid_request |
| requests of one key waiting to be read | 36 (12 per model × 3 models), or 8 MiB of bodies | 429 concurrency_limit |
| requests waiting to be read, all keys together | 256, or 32 MiB of bodies | 503 overloaded |
| options of a choice question | 2 to 255 | 400 invalid_request |
| levels of a score question | 2 to 10 | 400 invalid_request |
| one row: the state plus one question | 36,864 tokens | 413 state_too_long |
| rows per request | 1,024 | 413 too_many_questions |
| all rows of a request together, the state counted in every row | 1,048,576 tokens | 413 request_too_long |
| requests in flight per key and model | 12 | 429 concurrency_limit |
| time for one request | 120 s | 504 upstream_timeout |
A request waits to be read only while its body is parsed. A key may have as many requests waiting as it may have in flight on all models together, 36 (12 per model, three models), and at most 8 MiB of bodies waiting. As long as a key keeps within 12 requests per model, the count cannot refuse it; only the 8 MiB can, and then it gets 429 concurrency_limit, as for its requests in flight. The 503 overloaded of a full queue is the load of all keys together, and no single key can fill that queue.
A state of up to 32,768 tokens always fits with any question of up to 4,096 tokens; a longer state fits as long as the row stays within 36,864. The state is never truncated: a request that does not fit is refused as a whole and not billed. Token counts are those of the model's tokeniser, the same for all three models; a one-question choice or noul request reports its row length in usage.input_tokens.
Every request to api.llmtech.eu first passes our front proxy. Its limits apply to all endpoints together, chat completions included, and it refuses with its own responses, not in the error format below: the body is an HTML page, and there is no Inference-Id and no Retry-After header.
| Limit | Value | Over it |
|---|---|---|
| open connections with the same Authorization header | 64, shared by all endpoints | 429, HTML body |
| request rate from one IP address | 100 per second, burst 200 | 429, HTML body |
| request rate of the whole API | a fixed ceiling | 429, HTML body |
| request body | 20,971,520 bytes | 413, HTML body |
A 429 without Retry-After and without a JSON body comes from the proxy: wait a second and retry. The 64 connections of a key are counted over its chat completions and /v1/systemone requests together, so a key with many open chat streams can reach this limit before the 12 requests per model above. A body over 2,097,152 bytes but not over 20,971,520 reaches the API and gets the JSON 413 payload_too_large.
Inference-Id: a UUID on every response whose request was admitted and recorded in your usage journal: every 200 and every error marked "yes" in the Inference-Id column of the table below. Responses refused before admission (401, 402, 403, 404, 400, 413 payload_too_large, 429, and the 503 overloaded and 500 internal_error returned while the request is being read) carry none, as with chat completions.
Retry-After on 429, 1 second, and on the 503 overloaded we return before the model runs: 1 second for a full reading queue, up to 5 seconds while our side cannot process requests (not expected). Content-Type: application/json on every response. X-Inference-Location: IT and X-Inference-Site: Frosinone: the country (ISO 3166-1 alpha-2) and the site of the GPU node that runs the models, on responses of the API, errors included, with the same exceptions as on chat completions (response headers, above): not on responses the front proxy produces itself, not on a 400 for a request that is not valid HTTP, and not on the response to a request served by a model server outside this node. The front proxy's own 429 and 413 (see limits in front of the api, above) carry none of these headers.
Every error body is exactly {"error": {"message": "...", "type": "...", "code": "..."}}; the 402 answers marked "no code" have no code field (the same bodies as for chat completions). Outside this format are only the front proxy's HTML 429 and 413 (see limits in front of the api, above) and the plain-text 405 described below the table. No error message quotes your state, question ids, option names or any other part of the request. Checks run in the order of the table; the first failing one answers.
| Status | error.type | error.code | When | Billed | Inference-Id |
|---|---|---|---|---|---|
| 401 | invalid_request_error | invalid_api_key | no key, or an unknown key | no | no |
| 403 | invalid_request_error | read_only_key | a read-only (cabinet) key | no | no |
| 403 | invalid_request_error | management_key | a key that manages stream-hour units | no | no |
| 403 | invalid_request_error | model_not_available_on_channel | a stream-hour unit key, or decision models are not enabled for this key | no | no |
| 413 | invalid_request_error | payload_too_large | body over 2,097,152 bytes, as sent or after decompression | no | no |
| 400 | invalid_request_error | invalid_request | the body cannot be decompressed, holds more than 64 gzip members, or is sent with Content-Encoding: br or zstd; over 500,000 { and [ or over 1,000,000 values in the body | no | no |
| 429 | rate_limit_error | concurrency_limit | 36 requests of this key (12 per model × 3 models) are already waiting to be read, or its waiting bodies would pass 8 MiB with this one; retry in a second | no | no |
| 503 | api_error | overloaded | too many requests to the decision models are waiting to be read, all keys together (over 256, or over 32 MiB of bodies), retry in a second; or we cannot read requests right now (not expected), retry after Retry-After | no | no |
| 500 | api_error | internal_error | we could not read the request on our side (not expected); retry | no | no |
| 400 | invalid_request_error | invalid_request | body is not a JSON object; nesting over 128 levels | no | no |
| 404 | invalid_request_error | unknown_model | model missing or not one of the three ids | no | no |
| 400 | invalid_request_error | invalid_request | questions not a non-empty object; state missing or null; independent is false or not a boolean; a question breaks a rule above (the message names it by its position, question #2: …); a string or key in state or questions holds an unpaired UTF-16 surrogate | no | no |
| 402 | insufficient_quota | no code | daily token limit of the address, daily spend cap or daily token cap of the key | no | no |
| 402 | insufficient_quota | monthly_cap_reached | the monthly spend cap you set | no | no |
| 402 | insufficient_balance | no code | prepaid balance used up | no | no |
| 429 | rate_limit_error | concurrency_limit | 12 requests of this key to this model already in flight | no | no |
| 429 | rate_limit_error | channel_concurrency_limit | the channel's concurrency ceiling reached | no | no |
| 503 | api_error | overloaded | we cannot process the model's answers right now (not expected); the request did not reach the model; retry after Retry-After | no | yes |
| 413 | invalid_request_error | state_too_long | a row (the state plus one question) is over 36,864 tokens | no | yes |
| 413 | invalid_request_error | too_many_questions | over 1,024 rows | no | yes |
| 413 | invalid_request_error | request_too_long | all rows together over 1,048,576 tokens | no | yes |
| 400 | invalid_request_error | invalid_request | the model server refused a question the checks above let through (not expected; the message says "the model server rejected …") | no | yes |
| 503 | api_error | overloaded | the model's queue is full (over 4,096 rows waiting); retry in a few seconds | no | yes |
| 503 | api_error | model_unavailable | the model is not running | no | yes |
| 502 | api_error | upstream_bad_response | the model answered without usage or without an answer for every question | no | yes |
| 500 | api_error | internal_error | the model answered, but we could not process its answer on our side (not expected) | no | yes |
| 502 | api_error | upstream_error | the connection to the model broke (after one automatic retry), the model's answer holds values that cannot be returned as JSON (NaN, Infinity, an unpaired surrogate), or any other model-server failure | no | yes |
| 504 | api_error | upstream_timeout | no answer within 120 s | no | yes |
A request is billed only when it returns 200. If you disconnect before the answer arrives, the request still runs to the end and is billed if it succeeds, as a non-streaming chat completion is. The endpoint takes only POST; any other method gets 405 with a plain-text body from the HTTP server.
curl https://api.llmtech.eu/v1/systemone \
-H "Authorization: Bearer YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "decider-2b-fp8",
"state": {"ticket": {"id": "A-104", "messages": [
{"from": "customer",
"text": "I was charged twice for order A-104. Please refund the duplicate."}]},
"refund_policy": "Duplicate charges are refunded within 5 days."},
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "charges, invoices, refunds",
"shipping": "delivery and returns", "other": null}},
"refund_requested": {"type": "noul",
"instructions": "Does ticket.messages[0].text ask for a refund?"},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["calm", "frustrated", "very frustrated"]}
}
}'
{
"model": "decider-2b-fp8",
"answers": {
"department": {
"type": "choice",
"choice": "billing",
"confidence": 0.895,
"x_p_max": 0.93,
"certainty": 0.731,
"probabilities": {"billing": 0.93, "shipping": 0.02, "other": 0.05}
},
"refund_requested": {"type": "noul", "noul": 0.97},
"frustration": {
"type": "score",
"score": 1.1,
"confidence": 0.5571,
"x_p_max": 0.7048,
"certainty": 0.2787,
"legend": {"0": "calm", "1": "frustrated", "2": "very frustrated"},
"probabilities": {"0": 0.0952, "1": 0.7048, "2": 0.2},
"level_fit": {"0": 0.1, "1": 0.74, "2": 0.21},
"fit_mass": 1.05
}
},
"usage": {"input_tokens": 231, "output_tokens": 0}
}
The request and the response shape are real; so is input_tokens (231: five rows over a 63-token context; one row each for department and refund_requested, three for the levels of frustration). The probabilities are illustrative, not a model's judgement; every value derived from them (choice, confidence, x_p_max, certainty, score, level_fit, fit_mass) was computed by the serving code from those probabilities.
Context:
{"ticket": {"id": "A-104", "messages": [{"from": "customer", "text": "I was charged twice for order A-104. Please refund the duplicate."}]}, "refund_policy": "Duplicate charges are refunded within 5 days."}
Question: Does ticket.messages[0].text ask for a refund?
Options:
(A) no
(B) yes
Answer: (