Decider decision models
A decision model reads a state and answers typed questions about it. It does not generate text: every question is read in its own prompt row, a score question by default in one row per level, and each row takes one forward pass. An answer is the most likely of the options you name, the probability of yes, or the expected level on a scale you describe. Output tokens are 0, and a request bills its input tokens only, with the state counted once.
| Model id | Input | Output | Accuracy vs bf16, in-task / held-out |
|---|---|---|---|
| decider-4b-nvfp4 | $0.04 | free | -0.6 / -0.7 |
| decider-2b-fp8 | $0.03 | free | 0.0 / 0.0 |
| decider-0.8b-fp8 | $0.02 | free | -0.1 / 0.0 |
Accuracy is in points against the bf16 weights on the author's regression set: 95 tasks, 144,226 rows. For decider-2b-fp8 the change is within 0.1 points on both halves. The models are Mapika/decider (Apache 2.0, author on Hugging Face: Mapika), quantised and measured by LLM Tech. Full measurements are on the model cards: decider-4b-nvfp4, decider-2b-fp8, decider-0.8b-fp8.
Request and response fields, limits and error codes are in the docs.
curl https://api.llmtech.eu/v1/systemone \
-H "Authorization: Bearer YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "decider-2b-fp8",
"state": {"ticket": {"id": "A-104", "messages": [
{"from": "customer",
"text": "I was charged twice for order A-104. Please refund the duplicate."}]},
"refund_policy": "Duplicate charges are refunded within 5 days."},
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "charges, invoices, refunds",
"shipping": "delivery and returns", "other": null}},
"refund_requested": {"type": "noul",
"instructions": "Does ticket.messages[0].text ask for a refund?"},
"frustration": {"type": "score", "instructions": "How frustrated is the customer?",
"criteria": ["calm", "frustrated", "very frustrated"]}
}
}'
{
"model": "decider-2b-fp8",
"answers": {
"department": {
"type": "choice",
"choice": "billing",
"confidence": 0.895,
"x_p_max": 0.93,
"certainty": 0.731,
"probabilities": {"billing": 0.93, "shipping": 0.02, "other": 0.05}
},
"refund_requested": {"type": "noul", "noul": 0.97},
"frustration": {
"type": "score",
"score": 1.1,
"confidence": 0.5571,
"x_p_max": 0.7048,
"certainty": 0.2787,
"legend": {"0": "calm", "1": "frustrated", "2": "very frustrated"},
"probabilities": {"0": 0.0952, "1": 0.7048, "2": 0.2},
"level_fit": {"0": 0.1, "1": 0.74, "2": 0.21},
"fit_mass": 1.05
}
},
"usage": {"input_tokens": 231, "output_tokens": 0}
}
The request and the response shape are real; so is input_tokens (231: five rows over a 63-token context; one row each for department and refund_requested, three for the levels of frustration). The probabilities are illustrative, not a model's judgement; every value derived from them was computed by the serving code from those probabilities.
| State length | decider-4b-nvfp4 | decider-2b-fp8 | decider-0.8b-fp8 |
|---|---|---|---|
| 1K tokens | 21 ms | 16 ms | 16 ms |
| 8K tokens | 120 ms | 68 ms | 44 ms |
One request at a time on an NVIDIA RTX PRO 6000 Blackwell with nothing else running on it. A state repeated from the prefix cache came back in about 50 to 60 ms at 29K tokens. These are idle-GPU figures. Latency under production traffic has not been measured yet.
Zero data retention
The same data policy as the rest of the API. The state, the questions and the answers are processed in volatile memory only: never written to disk, logs, analytics or backups, and never used for training. We retain per-request technical metadata only: timestamp, request id, key fingerprint, pricing channel, source IP address, model, token counts, prices and cost, HTTP status, error code and duration. Details are on the security page.
Where does inference run?
The decision models run on the same GPU node as Qwen3.8-27B. All compute is located in the European Union: the GPU node is in Italy, and the edge is in Nuremberg, Germany. No data leaves the EU/EEA and there is no US cloud at any hop. TLS 1.3 to our edge, then an encrypted WireGuard tunnel to the GPU node.