Choose an option.
Route a shipment, select a category, or pick the next action. Your semantic IDs stay yours.
{ "selected": "chilled" }CHOICE ↗LLM TO SYSTEM 1 v0.1
Give your model a question.
Give your code a typed answer.
Turn local GGUF model scores into choices, booleans, and ordered values.
{ "storage_requirement": "chilled", "hours_until_dispatch": 4 }
Output example
01 / THE CONTRACT
Define the options, read their scores, and return a typed value.
Route a shipment, select a category, or pick the next action. Your semantic IDs stay yours.
{ "selected": "chilled" }CHOICE ↗Turn a two-option question into a boolean, with the candidate-relative probability alongside it.
{ "value": true, "p_true": 0.92 }BINARY ↗Use your own ordered levels and numeric values for priority, severity, or another defined scale.
{ "selected": "high" }ORDINAL ↗02 / CLEAR ANSWERS
Choose a decision type and get a value your application can use.
Output example
{
"value": {
"type": "choice",
"selected": "chilled"
}
}03 / YOUR MODEL, YOUR MACHINE
Run your chosen model locally.
Llama · Qwen · Gemma. Local inference on CPU, CUDA and Metal.
Locally verified GGUF checkpoints. Read the verification record ↗
04 / MEASURE THE WORK
Real decisions on a 12 GB RTX 3060. Choose the model that fits your workload.
RTX 3060 · 12 GB CUDACUDA p50 / request
Qwen3 0.6B Q8_0
JevBench · 231 items · raw accuracy 31.60%
Qwen3 0.6B Q8_0 · RTX 3060
Qwen3.5 4B Q8_0 · JevBench
Run inference on your own hardware.
36.1% ↓Qwen3.5 9B Q4 vs Q8 · lower sampled GPU memory peak · both 77.92% JevBench raw accuracy
| Checkpoint | JevBench raw accuracy | Coverage | Accepted accuracy | Accepted correct / all | p50 / p95 | Sampled GPU memory peak | GGUF file size |
|---|---|---|---|---|---|---|---|
| Qwen3.5 4B Q8_0 | 79.65% | 59.74% | 93.48% | 55.84% | 86.9 ms / 1137.3 ms | 4.92 GiB | 4.17 GiB |
| Qwen3.5 9B Q4_K_M | 77.92% | 72.29% | 92.22% | 66.67% | 130.6 ms / 1697.1 ms | 5.50 GiB | 5.29 GiB |
| Qwen3.5 9B Q8_0 | 77.92% | 71.43% | 91.52% | 65.37% | 124.7 ms / 1613.8 ms | 8.61 GiB | 8.87 GiB |
| Gemma 4 E4B it Q4_K_M | 77.49% | 83.98% | 84.54% | 71.00% | 86.6 ms / 1186.5 ms | 3.90 GiB | 4.64 GiB |
| Gemma 4 E4B it Q8_0 | 76.62% | 83.98% | 83.51% | 70.13% | 85.0 ms / 1146.2 ms | 5.91 GiB | 7.63 GiB |
| Qwen3.5 4B Q4_K_M | 76.19% | 60.61% | 92.86% | 56.28% | 88.5 ms / 1175.2 ms | 3.30 GiB | 2.55 GiB |
| Qwen3.8 27B UD-IQ2_XXS | 73.16% | 63.64% | 87.76% | 55.84% | 414.6 ms / 5460.2 ms | 7.64 GiB | 6.77 GiB |
| Qwen3 8B Q8_0 | 71.43% | 93.51% | 74.07% | 69.26% | 121.5 ms / 1817.9 ms | 9.13 GiB | 8.11 GiB |
| Gemma 4 E2B it Q8_0 | 67.97% | 90.48% | 72.25% | 65.37% | 46.8 ms / 701.1 ms | 2.95 GiB | 4.63 GiB |
| Qwen3 4B Q8_0 | 65.80% | 87.45% | 68.81% | 60.17% | 85.3 ms / 1349.3 ms | 5.62 GiB | 3.99 GiB |
| Ministral-3 8B Instruct 2512 Q4_K_M | 65.37% | 59.74% | 83.33% | 49.78% | 441.4 ms / 2399.7 ms | 6.13 GiB | 4.84 GiB |
| gpt-oss 20B Q4_K_M | 64.94% | 64.50% | 79.19% | 51.08% | 181.9 ms / 2262.3 ms | 11.62 GiB | 10.83 GiB |
| Qwen3.5 2B Q8_0 | 62.34% | 42.42% | 84.69% | 35.93% | 41.0 ms / 513.0 ms | 2.41 GiB | 1.87 GiB |
| Gemma 3 4B it Q8_0 | 59.74% | 96.10% | 60.36% | 58.01% | 74.7 ms / 879.9 ms | 5.35 GiB | 3.85 GiB |
| Phi-4-mini Instruct Q8_0 | 56.71% | 50.22% | 74.14% | 37.23% | 70.5 ms / 972.0 ms | 5.21 GiB | 3.80 GiB |
| Qwen3.5 0.8B Q8_0 | 51.95% | 20.78% | 66.67% | 13.85% | 27.9 ms / 350.5 ms | 1.29 GiB | 0.76 GiB |
| SmolLM3 3B Q8_0 | 46.32% | 45.02% | 60.58% | 27.27% | 71.5 ms / 884.8 ms | 3.95 GiB | 3.05 GiB |
| Llama-3.2 3B Instruct Q8_0 | 43.72% | 36.36% | 45.24% | 16.45% | 68.2 ms / 884.8 ms | 4.48 GiB | 3.19 GiB |
| Gemma 3 1B it Q8_0 | 38.96% | 87.01% | 39.80% | 34.63% | 25.1 ms / 309.2 ms | 1.62 GiB | 1.00 GiB |
| Qwen3 0.6B Q8_0 | 31.60% | 82.25% | 32.11% | 26.41% | 34.0 ms / 445.1 ms | 1.81 GiB | 0.60 GiB |
GGUF download and disk space · exact file bytes · 1 GiB = 1,024 MiB. 231 public items per model · context 8192 · direct decisions. Acceptance: top probability ≥ 0.80, candidate mass ≥ 0.05. GPU memory: whole-device samples every 200 ms. GPT-OSS: GGML_CUDA_DISABLE_GRAPHS=1. Protocol and report ↗
Download measured evidence: RTX 3060 summary JSON · Manifest
Typed-decisions · 400 cases / 2000 decisions · RTX 3080 10 GiB, WSL2 · 2026-09-26. Latency per five-question case. Protocol and report ↗
Download measured evidence: Gemma summary JSON · Qwen summary JSON · Gemma scored JSONL · Qwen scored JSONL · Manifest
Run the performance test on your machine05 / START LOCAL
Bring a GGUF checkpoint and a matching llama.cpp build. Define a question, run the CLI, and inspect the JSON.
Download the English examplecargo run --release --features llama -- \
--model /path/to/chat-model.gguf \
--input examples/warehouse.json \
--device cpuBuild with Rust, run your GGUF model, and integrate typed results into your application.