Service endingThe LLM Tech API stops on 31 October 2026. Keys stop working on 1 November 2026, 00:00 UTC. Our quantized models stay on Hugging Face: huggingface.co/llmtech
Docs · quickstart

First response in under a minute.

The API is OpenAI-compatible. Base URL https://api.llmtech.eu/v1, model nvidia/Qwen3.8-27B-NVFP4. Below is a shared free trial key — try before you talk to anyone.

trial key
lt-trial-ba1ef28c6d32ed6980678d8d

Shared and rate-limited: 4 concurrent requests and 2M tokens per day per address, counted across prompt and completion and reset at 00:00 UTC. Enough to evaluate, not enough to run on. For production, email artem@llmtech.eu — a personal key with no daily limit is issued the same day, usually within the hour; it shares the endpoint's 64 requests in flight with other keys. Working with scans or PDFs? The document demo runs the same endpoint with no key at all.

first request
curl
curl https://api.llmtech.eu/v1/chat/completions \
  -H "Authorization: Bearer lt-trial-ba1ef28c6d32ed6980678d8d" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nvidia/Qwen3.8-27B-NVFP4",
    "messages": [{"role": "user", "content": "Say hello"}],
    "stream": true
  }'
python (openai sdk)
from openai import OpenAI

client = OpenAI(
    base_url="https://api.llmtech.eu/v1",
    api_key="lt-trial-ba1ef28c6d32ed6980678d8d",
)

r = client.chat.completions.create(
    model="nvidia/Qwen3.8-27B-NVFP4",
    messages=[{"role": "user", "content": "Say hello"}],
    stream=True,
)
for chunk in r:
    print(chunk.choices[0].delta.content or "", end="")
javascript (openai sdk)
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.llmtech.eu/v1",
  apiKey: "lt-trial-ba1ef28c6d32ed6980678d8d",
});

const stream = await client.chat.completions.create({
  model: "nvidia/Qwen3.8-27B-NVFP4",
  messages: [{ role: "user", content: "Say hello" }],
  stream: true,
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
reasoning control

Reasoning is off by default. Turn it on per request with reasoning_effort or enable_thinking:

{
  "model": "nvidia/Qwen3.8-27B-NVFP4",
  "messages": [{"role": "user", "content": "Is 1001 a prime number?"}],
  "chat_template_kwargs": {
    "enable_thinking": true
  }
}
enable_thinkingtrue / false
reasoning_effort"low" / "medium" / "xhigh"

Reasoning text comes back in reasoning (non-streaming) or reasoning_content deltas (streaming). Reasoning tokens are billed as output.

prompt caching

Automatic, no code changes. Repeated prompt prefixes are billed at $0.04/M instead of $0.25/M. The first request writes the cache and the second already reads it. The cache works in blocks of 1,600 tokens and reuses every whole block except the last one, so a prompt under 3,200 tokens gets none. Check usage.prompt_tokens_details.cached_tokens in the response to see it working.

The cache belongs to your key. Two of your own keys do not share a cached prefix, and you never share one with anybody else. Inside one key you can split it yourself by sending your own cache_salt. The one exception is the public trial key. Everyone who uses it shares the key, so its cache is split by source IP address: guests behind one address, such as an office network or a VPN, share a cache. Send nothing confidential with it.

errors and limits
CodeMeaningWhat to do
401invalid or missing API keycheck the Authorization header
402daily limit reached, or prepaid balance used upthe error type says which: insufficient_quota for a limit, insufficient_balance for a balance; limits reset at 00:00 UTC
402monthly spending cap on your key reachedcode monthly_cap_reached; resets on the 1st of the month at 00:00 UTC, or ask us to change the cap
403this key may not generatea read-only cabinet key cannot spend tokens; use your working key
404unknown model iduse nvidia/Qwen3.8-27B-NVFP4; the earlier id unsloth/Qwen3.8-27B-NVFP4 still works
413request body larger than 20 MBsend images at a lower resolution or split the request
429concurrency limit reachedretry after the Retry-After header, 1 or 2 seconds; requests never queue silently
400malformed request / context overflowthe error message names the exact problem
5xxserver-side failureretry with backoff; check status

Default limits: up to 64 requests in flight on the public endpoint, shared by the keys on it; the trial key allows 4 for everyone using it together and a context of 131,072 tokens. Other keys get the full 262,144. Need concurrency that is yours alone, ask: that is what a channel of your own is for.

response headers

Responses of the API, errors included, carry X-Inference-Location: IT and X-Inference-Site: Frosinone: the country (ISO 3166-1 alpha-2) and the site of the GPU node that runs the request; a streamed answer carries them too. They are left out where they would not describe that node: on responses the front proxy produces itself (its 429 for too many connections or requests, its 413 for a body over 20 MB, its errors while the API behind it is unreachable, the static list at /openrouter/models and the 404 for /v1/admin/), on a 400 for a request that is not valid HTTP, and on the response to a request served by a GPU outside this node.

known quirks

A few things behave differently from a plain OpenAI endpoint, and one used to. We tested each one against production rather than guessing, and we would rather you read it here than discover it at three in the morning.

max output32,768 tokens; a larger max_tokens or max_completion_tokens is lowered to it, and a request that names neither gets it as its limit
calls from a browser pageno CORS headers on the generation endpoints; call them from your server
min_p and logit_biasaccepted since 18 Sep 2026; the engine rejected them before that
system message positionmust come first; several in a row are fine
images in a system messagenot allowed — put them in a user message
reasoning_effortlow / medium / xhigh; minimal maps to low, high and max map to xhigh
model in the responseechoes our internal name qwen38, not the id you sent

The last one is worth a sentence. Requests should use the public id nvidia/Qwen3.8-27B-NVFP4, and the earlier id unsloth/Qwen3.8-27B-NVFP4 keeps working as an alias so that nothing built against it breaks; anything else is a 404 — but the response and the stream chunks currently carry the internal served name. Harmless for every client we have seen, since nothing routes on it, but if your code asserts that the response model equals the requested one, that assertion will fail. Say so and we will change it.

Everything else you would expect works, including max_completion_tokens, seed, logprobs, response_format with a JSON schema, tool calling with the standard string-encoded arguments, the developer role, and the sampler set temperature / top_p / top_k / repetition_penalty.

decisions: /v1/systemone

POST https://api.llmtech.eu/v1/systemone answers typed questions about a state with the Decider decision models. It takes the same Bearer key as chat completions and is enabled per key. There is no text generation: every question is read in its own prompt row (a score question by default in one row per level), each row in one forward pass, and only input tokens are billed.

decider-4b-nvfp4$0.04 per 1M input tokens
decider-2b-fp8$0.03 per 1M input tokens
decider-0.8b-fp8$0.02 per 1M input tokens

Output is free. GET /v1/models with a key that has decision models enabled lists the three models with "type": "decisions" and "endpoint": "/v1/systemone"; without a key, or with a key that does not have them, they are not listed. To have decision models enabled for a key, email artem@llmtech.eu.

The model server keeps a prefix cache per key: an identical state sent again on the same key can be served faster, and the cache is never shared between keys, so no key can tell from timing that another key has sent the same state.

systemone: request

The body is one JSON object of at most 2,097,152 bytes. Content-Type is not checked.

A body may be compressed with Content-Encoding: gzip or deflate (zlib-wrapped or raw deflate). The API decompresses it itself; the 2,097,152-byte limit applies both to the body as sent, with or without Content-Length (chunked), and to the decompressed body (413 payload_too_large), and a body that cannot be decompressed, including a truncated stream, is refused with 400 invalid_request. A gzip body may be several gzip members one after another, at most 64; a body with more is refused with 400 invalid_request. Content-Encoding: br and zstd are not supported and are refused with 400 invalid_request. Any other Content-Encoding value is ignored: the body is read as it was sent.

FieldRequiredTypeMeaning
modelyesstringone of the three model ids above
stateyesany JSON value except nullwhat the questions are about
questionsyesobject, at least one entryquestion id → question
independentnotrue (or null, the same as absent)every question is answered on its own; false is refused
layoutnoanyignored

Any other top-level field is ignored and not passed to the model server. Arrays and objects may nest at most 128 levels deep in the whole body, the body itself being level 1 (so a state may nest 127 levels). The body may hold at most 500,000 { and [ characters and at most 1,000,000 JSON values, counted as ,, { and [ characters together; both are counted in the (decompressed) body before it is parsed, so the characters inside strings count too. These are hard limits on the body, not on what the model can answer: a body over either is refused with 400 invalid_request even when the model could answer the request. For scale, 1,024 questions of 255 options each hold about 264,000 such values, and 2 MiB of single-digit numbers about 1,048,000. Every string and object key in state and questions must be valid Unicode: an unpaired UTF-16 surrogate (a \ud800 to \udfff escape without its partner) is refused with 400 invalid_request, and the message says where (state or question #N). A proper pair, such as an emoji written as two escapes, is fine.

state

A string is used as it is. Any other value is rendered as JSON text: separators ", " and ": ", non-ASCII characters kept, and every array of 8 or more elements rewritten so that each element carries its position: an object element e at index i becomes {"_index": i, ...e} (a key _index already in e wins), any other element becomes {"_index": i, "value": e}. Arrays of fewer than 8 elements keep their elements as they are (arrays nested inside them follow the same rule). Numbers, true and false become their JSON text (42, true).

The model reads "Context:\n" followed by that text. Nothing is ever cut from the state: a state that does not fit is refused with 413 (see limits below).

questions

An object from your question id to a question. Ids are any strings; they are never shown to the model. Answers come back under the same ids, in the same order. Every question is an object with these fields:

Fieldchoicescorenoul
type"choice" (the default when type is absent)"score""noul"; "bool" is accepted as the same
instructionsrequiredrequiredoptional if criteria describes "true" or "false"
criteriarequired: 2 to 255 optionsrequired: 2 to 10 levelsoptional
isolatedignoredoptional, true (default) or falseignored

instructions is the question in words: a non-empty string, or any other JSON value except null, which is sent to the model as its JSON text. question is accepted as a synonym of instructions and options as a synonym of criteria; when both names are present, instructions and criteria win. Other fields in a question are ignored.

choice: criteria is either an object from option name to description, or an array of distinct option names (strings). A description may be null or "" (the option is shown by its name only), a string, or any other JSON value (shown as JSON text). The model sees each option as name or name: description, in the order you give. The answer names the chosen option by its name.

score: criteria is either an array of level descriptions, lowest first, or an object whose keys are distinct decimal numbers ("0", "1", "2.5", "-1"; no exponent, no leading zeros, no spaces). The object form is sorted by the numeric value of its keys and the levels are then renumbered 0, 1, 2, …: the answer uses these numbers, not your keys. A level description is a string or any JSON value (shown as JSON text).

By default every level is judged alone, in its own row: the model sees the question, one level ("Proposed answer: …", with a leading N: removed from the level text) and answers whether it fits. With "isolated": false all levels are shown together as options 0: …, 1: … in one row.

noul: a yes/no question answered with the probability of yes. criteria, if given, is an object with optional "true" and "false" descriptions (other keys are ignored); the model sees the options as no / no: <false description> and yes / yes: <true description>. Without instructions (absent, null or "") the question text is the fixed sentence "Which answer fits the context?", and at least one of the two descriptions must be present and neither null nor "".

where we are stricter than the model server

The model server interprets a few inputs in Python's way. We refuse them with 400 instead, so that no request is answered in a way its author did not intend:

Refused with 400What the model server would do
state: nullread the text null
instructions: null in a choice or score questionask the question "null"
a choice criteria array with a non-string element or a repeated nameconvert elements with Python's str(), so true becomes True, and merge repeats
score criteria object keys that are not plain decimal numbers or are equal as numbers ("01", "1e1", " 1", "nan", "1_0", "0" with "0.0")accept anything Python's float() reads
isolated that is not true or falsetake its truth value, so "false" would mean true
independent that is not true, false or nullcoerce "yes", 1 and similar
systemone: response and usage

200 OK, Content-Type: application/json, header Inference-Id (see headers below). The body:

{"model": "<the model id you sent>",
 "answers": {"<question id>": {...}, ...},
 "usage": {"input_tokens": <integer>, "output_tokens": 0}}

answers holds exactly one answer per question id, in the order of questions. Fields per answer, in this order (probabilities and derived values are rounded to 4 decimals, score to 2, each value on its own, so probabilities may not add up to exactly 1):

choice
FieldValue
type"choice"
choicethe option name with the highest probability (the first one on a tie)
confidence(n·p_max − 1)/(n − 1) clipped to [0, 1]: 0 for a uniform distribution, 1 when one option has all of it
x_p_maxthe highest probability, p_max
certainty1 − H(p)/ln n floored at 0, where H is the entropy in nats
probabilitiesoption name → probability, in the order of criteria
score
FieldValue
type"score"
scorethe expected level Σ i·pᵢ over levels 0 … n−1
confidencemax(0, 1 − Σ pᵢ·|i − k| / D), k the most likely level, D = (1/n)·Σ |i − (n−1)/2|
x_p_maxthe probability of the most likely level
certaintyas for choice
legend"0", "1", … → the level description as text (non-string descriptions as JSON text)
probabilities"0", "1", … → probability of the level
level_fitonly with isolated levels (the default): "0", "1", … → the probability that this level fits, judged alone
fit_massonly with isolated levels: the sum of level_fit; near 1 when one level fits, low when none does, above 1 when several do. probabilities is level_fit divided by it
noul (also for "type": "bool")
FieldValue
type"noul"
noulthe probability of yes
usage

output_tokens is always 0. input_tokens is what the request is billed for: every distinct token position the model reads, with the state counted once.

A request becomes scoring rows: one per question, and one per level of a score question with isolated levels. Each row is the same context ("Context:\n" plus the rendered state, C tokens) followed by that row's question block (Qᵢ tokens: the question, its options and the answer slot). The question blocks all begin with the same few tokens (L, the length of their longest common beginning; 0 for a single row). Then

input_tokens = C + L + Σᵢ (Qᵢ − L)

For a request with one choice or noul question, input_tokens is the length of its row. The same number is what your usage and your invoice count for the request; the price is per 1M of them.

systemone: limits, headers, errors
LimitValueOver it
request body, as sent and after decompression2,097,152 bytes413 payload_too_large
JSON nesting in the body128 levels400 invalid_request
{ and [ in the body, also inside strings500,000400 invalid_request
JSON values, counted as , { [ in the body, also inside strings1,000,000400 invalid_request
requests of one key waiting to be read36 (12 per model × 3 models), or 8 MiB of bodies429 concurrency_limit
requests waiting to be read, all keys together256, or 32 MiB of bodies503 overloaded
options of a choice question2 to 255400 invalid_request
levels of a score question2 to 10400 invalid_request
one row: the state plus one question36,864 tokens413 state_too_long
rows per request1,024413 too_many_questions
all rows of a request together, the state counted in every row1,048,576 tokens413 request_too_long
requests in flight per key and model12429 concurrency_limit
time for one request120 s504 upstream_timeout

A request waits to be read only while its body is parsed. A key may have as many requests waiting as it may have in flight on all models together, 36 (12 per model, three models), and at most 8 MiB of bodies waiting. As long as a key keeps within 12 requests per model, the count cannot refuse it; only the 8 MiB can, and then it gets 429 concurrency_limit, as for its requests in flight. The 503 overloaded of a full queue is the load of all keys together, and no single key can fill that queue.

A state of up to 32,768 tokens always fits with any question of up to 4,096 tokens; a longer state fits as long as the row stays within 36,864. The state is never truncated: a request that does not fit is refused as a whole and not billed. Token counts are those of the model's tokeniser, the same for all three models; a one-question choice or noul request reports its row length in usage.input_tokens.

limits in front of the api

Every request to api.llmtech.eu first passes our front proxy. Its limits apply to all endpoints together, chat completions included, and it refuses with its own responses, not in the error format below: the body is an HTML page, and there is no Inference-Id and no Retry-After header.

LimitValueOver it
open connections with the same Authorization header64, shared by all endpoints429, HTML body
request rate from one IP address100 per second, burst 200429, HTML body
request rate of the whole APIa fixed ceiling429, HTML body
request body20,971,520 bytes413, HTML body

A 429 without Retry-After and without a JSON body comes from the proxy: wait a second and retry. The 64 connections of a key are counted over its chat completions and /v1/systemone requests together, so a key with many open chat streams can reach this limit before the 12 requests per model above. A body over 2,097,152 bytes but not over 20,971,520 reaches the API and gets the JSON 413 payload_too_large.

headers

Inference-Id: a UUID on every response whose request was admitted and recorded in your usage journal: every 200 and every error marked "yes" in the Inference-Id column of the table below. Responses refused before admission (401, 402, 403, 404, 400, 413 payload_too_large, 429, and the 503 overloaded and 500 internal_error returned while the request is being read) carry none, as with chat completions.

Retry-After on 429, 1 second, and on the 503 overloaded we return before the model runs: 1 second for a full reading queue, up to 5 seconds while our side cannot process requests (not expected). Content-Type: application/json on every response. X-Inference-Location: IT and X-Inference-Site: Frosinone: the country (ISO 3166-1 alpha-2) and the site of the GPU node that runs the models, on responses of the API, errors included, with the same exceptions as on chat completions (response headers, above): not on responses the front proxy produces itself, not on a 400 for a request that is not valid HTTP, and not on the response to a request served by a model server outside this node. The front proxy's own 429 and 413 (see limits in front of the api, above) carry none of these headers.

errors

Every error body is exactly {"error": {"message": "...", "type": "...", "code": "..."}}; the 402 answers marked "no code" have no code field (the same bodies as for chat completions). Outside this format are only the front proxy's HTML 429 and 413 (see limits in front of the api, above) and the plain-text 405 described below the table. No error message quotes your state, question ids, option names or any other part of the request. Checks run in the order of the table; the first failing one answers.

Statuserror.typeerror.codeWhenBilledInference-Id
401invalid_request_errorinvalid_api_keyno key, or an unknown keynono
403invalid_request_errorread_only_keya read-only (cabinet) keynono
403invalid_request_errormanagement_keya key that manages stream-hour unitsnono
403invalid_request_errormodel_not_available_on_channela stream-hour unit key, or decision models are not enabled for this keynono
413invalid_request_errorpayload_too_largebody over 2,097,152 bytes, as sent or after decompressionnono
400invalid_request_errorinvalid_requestthe body cannot be decompressed, holds more than 64 gzip members, or is sent with Content-Encoding: br or zstd; over 500,000 { and [ or over 1,000,000 values in the bodynono
429rate_limit_errorconcurrency_limit36 requests of this key (12 per model × 3 models) are already waiting to be read, or its waiting bodies would pass 8 MiB with this one; retry in a secondnono
503api_erroroverloadedtoo many requests to the decision models are waiting to be read, all keys together (over 256, or over 32 MiB of bodies), retry in a second; or we cannot read requests right now (not expected), retry after Retry-Afternono
500api_errorinternal_errorwe could not read the request on our side (not expected); retrynono
400invalid_request_errorinvalid_requestbody is not a JSON object; nesting over 128 levelsnono
404invalid_request_errorunknown_modelmodel missing or not one of the three idsnono
400invalid_request_errorinvalid_requestquestions not a non-empty object; state missing or null; independent is false or not a boolean; a question breaks a rule above (the message names it by its position, question #2: …); a string or key in state or questions holds an unpaired UTF-16 surrogatenono
402insufficient_quotano codedaily token limit of the address, daily spend cap or daily token cap of the keynono
402insufficient_quotamonthly_cap_reachedthe monthly spend cap you setnono
402insufficient_balanceno codeprepaid balance used upnono
429rate_limit_errorconcurrency_limit12 requests of this key to this model already in flightnono
429rate_limit_errorchannel_concurrency_limitthe channel's concurrency ceiling reachednono
503api_erroroverloadedwe cannot process the model's answers right now (not expected); the request did not reach the model; retry after Retry-Afternoyes
413invalid_request_errorstate_too_longa row (the state plus one question) is over 36,864 tokensnoyes
413invalid_request_errortoo_many_questionsover 1,024 rowsnoyes
413invalid_request_errorrequest_too_longall rows together over 1,048,576 tokensnoyes
400invalid_request_errorinvalid_requestthe model server refused a question the checks above let through (not expected; the message says "the model server rejected …")noyes
503api_erroroverloadedthe model's queue is full (over 4,096 rows waiting); retry in a few secondsnoyes
503api_errormodel_unavailablethe model is not runningnoyes
502api_errorupstream_bad_responsethe model answered without usage or without an answer for every questionnoyes
500api_errorinternal_errorthe model answered, but we could not process its answer on our side (not expected)noyes
502api_errorupstream_errorthe connection to the model broke (after one automatic retry), the model's answer holds values that cannot be returned as JSON (NaN, Infinity, an unpaired surrogate), or any other model-server failurenoyes
504api_errorupstream_timeoutno answer within 120 snoyes

A request is billed only when it returns 200. If you disconnect before the answer arrives, the request still runs to the end and is billed if it succeeds, as a non-streaming chat completion is. The endpoint takes only POST; any other method gets 405 with a plain-text body from the HTTP server.

systemone: example
request
curl https://api.llmtech.eu/v1/systemone \
  -H "Authorization: Bearer YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "decider-2b-fp8",
    "state": {"ticket": {"id": "A-104", "messages": [
                {"from": "customer",
                 "text": "I was charged twice for order A-104. Please refund the duplicate."}]},
              "refund_policy": "Duplicate charges are refunded within 5 days."},
    "questions": {
      "department": {"type": "choice", "instructions": "Which team should handle this?",
                     "criteria": {"billing": "charges, invoices, refunds",
                                  "shipping": "delivery and returns", "other": null}},
      "refund_requested": {"type": "noul",
                           "instructions": "Does ticket.messages[0].text ask for a refund?"},
      "frustration": {"type": "score", "instructions": "How frustrated is the customer?",
                      "criteria": ["calm", "frustrated", "very frustrated"]}
    }
  }'
response: 200, headers Content-Type: application/json and Inference-Id: <uuid>
{
  "model": "decider-2b-fp8",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "billing",
      "confidence": 0.895,
      "x_p_max": 0.93,
      "certainty": 0.731,
      "probabilities": {"billing": 0.93, "shipping": 0.02, "other": 0.05}
    },
    "refund_requested": {"type": "noul", "noul": 0.97},
    "frustration": {
      "type": "score",
      "score": 1.1,
      "confidence": 0.5571,
      "x_p_max": 0.7048,
      "certainty": 0.2787,
      "legend": {"0": "calm", "1": "frustrated", "2": "very frustrated"},
      "probabilities": {"0": 0.0952, "1": 0.7048, "2": 0.2},
      "level_fit": {"0": 0.1, "1": 0.74, "2": 0.21},
      "fit_mass": 1.05
    }
  },
  "usage": {"input_tokens": 231, "output_tokens": 0}
}

The request and the response shape are real; so is input_tokens (231: five rows over a 63-token context; one row each for department and refund_requested, three for the levels of frustration). The probabilities are illustrative, not a model's judgement; every value derived from them (choice, confidence, x_p_max, certainty, score, level_fit, fit_mass) was computed by the serving code from those probabilities.

what the model reads for refund_requested
Context:
{"ticket": {"id": "A-104", "messages": [{"from": "customer", "text": "I was charged twice for order A-104. Please refund the duplicate."}]}, "refund_policy": "Duplicate charges are refunded within 5 days."}

Question: Does ticket.messages[0].text ask for a refund?
Options:
(A) no
(B) yes
Answer: (