LLM TO SYSTEM 1 v0.1

A small decision.
A clear result.

Give your model a question.
Give your code a typed answer.

Turn local GGUF model scores into choices, booleans, and ordered values.

MIT application codeCPU & CUDABring your own modelllama.cpp · GGUF
decision / storage_zone01
INPUT STATEJSON
{
  "storage_requirement": "chilled",
  "hours_until_dispatch": 4
}
↓LOCAL GGUF MODEL
Aambient—
Bchilled↗
Cfrozen—
↓TYPED RESULT
{"selected": "chilled"}

Output example

01 Your data stays with your runtime02 Scores you can inspect03 Typed answers for your application

01 / THE CONTRACT

Small outputs.
Useful building blocks.

Define the options, read their scores, and return a typed value.

A → B → C

Choose an option.

Route a shipment, select a category, or pick the next action. Your semantic IDs stay yours.

{ "selected": "chilled" }CHOICE ↗
0 / 1

Answer a condition.

Turn a two-option question into a boolean, with the candidate-relative probability alongside it.

{ "value": true, "p_true": 0.92 }BINARY ↗
▂ ▄ ▆

Find the level.

Use your own ordered levels and numeric values for priority, severity, or another defined scale.

{ "selected": "high" }ORDINAL ↗

02 / CLEAR ANSWERS

Return only
clear answers.

Choose a decision type and get a value your application can use.

ⓘ

Output example

Which storage zone?

Selected
Aambient
2%
Bchilled
96%
Cfrozen
2%
RESULT.JSON
{
  "value": {
    "type": "choice",
    "selected": "chilled"
  }
}

03 / YOUR MODEL, YOUR MACHINE

A common contract.
Different models.

Run your chosen model locally.

llama.cpp compatible.

Llama · Qwen · Gemma. Local inference on CPU, CUDA and Metal.

Llama 3.2 3B Q8_0 · RTX 3060 · 231 measured items. ↗

Locally verified GGUF checkpoints. Read the verification record ↗

04 / MEASURE THE WORK

Measured
performance.

Real decisions on a 12 GB RTX 3060. Choose the model that fits your workload.

RTX 3060 · 12 GB CUDA

CUDA p50 / request

34.0ms

Qwen3 0.6B Q8_0

JevBench · 231 items · raw accuracy 31.60%

Sampled GPU memory peak
1.81GiB

Qwen3 0.6B Q8_0 · RTX 3060

JevBench raw accuracy
79.65%

Qwen3.5 4B Q8_0 · JevBench

Shared-state speedup
2.78×

Gemma 4 E2B · RTX 3080 · synthetic 16-question workload

Timing evidence ↗
External inference API fee
0USD

Run inference on your own hardware.

GGUF file size609.8 MiB

Qwen3 0.6B Q8_0 · model download

File size evidence ↗

36.1% ↓Qwen3.5 9B Q4 vs Q8 · lower sampled GPU memory peak · both 77.92% JevBench raw accuracy

Every model. Every metric.

MEASURED DATA
RTX 3060 · 12 GB20 models · 4620 judgments · 2026-09-23
RTX 3060 JevBench public 231 quality, policy, latency, and sampled GPU memory
CheckpointJevBench raw accuracyCoverageAccepted accuracyAccepted correct / allp50 / p95Sampled GPU memory peakGGUF file size
Qwen3.5 4B Q8_079.65%59.74%93.48%55.84%86.9 ms / 1137.3 ms4.92 GiB4.17 GiB
Qwen3.5 9B Q4_K_M77.92%72.29%92.22%66.67%130.6 ms / 1697.1 ms5.50 GiB5.29 GiB
Qwen3.5 9B Q8_077.92%71.43%91.52%65.37%124.7 ms / 1613.8 ms8.61 GiB8.87 GiB
Gemma 4 E4B it Q4_K_M77.49%83.98%84.54%71.00%86.6 ms / 1186.5 ms3.90 GiB4.64 GiB
Gemma 4 E4B it Q8_076.62%83.98%83.51%70.13%85.0 ms / 1146.2 ms5.91 GiB7.63 GiB
Qwen3.5 4B Q4_K_M76.19%60.61%92.86%56.28%88.5 ms / 1175.2 ms3.30 GiB2.55 GiB
Qwen3.8 27B UD-IQ2_XXS73.16%63.64%87.76%55.84%414.6 ms / 5460.2 ms7.64 GiB6.77 GiB
Qwen3 8B Q8_071.43%93.51%74.07%69.26%121.5 ms / 1817.9 ms9.13 GiB8.11 GiB
Gemma 4 E2B it Q8_067.97%90.48%72.25%65.37%46.8 ms / 701.1 ms2.95 GiB4.63 GiB
Qwen3 4B Q8_065.80%87.45%68.81%60.17%85.3 ms / 1349.3 ms5.62 GiB3.99 GiB
Ministral-3 8B Instruct 2512 Q4_K_M65.37%59.74%83.33%49.78%441.4 ms / 2399.7 ms6.13 GiB4.84 GiB
gpt-oss 20B Q4_K_M64.94%64.50%79.19%51.08%181.9 ms / 2262.3 ms11.62 GiB10.83 GiB
Qwen3.5 2B Q8_062.34%42.42%84.69%35.93%41.0 ms / 513.0 ms2.41 GiB1.87 GiB
Gemma 3 4B it Q8_059.74%96.10%60.36%58.01%74.7 ms / 879.9 ms5.35 GiB3.85 GiB
Phi-4-mini Instruct Q8_056.71%50.22%74.14%37.23%70.5 ms / 972.0 ms5.21 GiB3.80 GiB
Qwen3.5 0.8B Q8_051.95%20.78%66.67%13.85%27.9 ms / 350.5 ms1.29 GiB0.76 GiB
SmolLM3 3B Q8_046.32%45.02%60.58%27.27%71.5 ms / 884.8 ms3.95 GiB3.05 GiB
Llama-3.2 3B Instruct Q8_043.72%36.36%45.24%16.45%68.2 ms / 884.8 ms4.48 GiB3.19 GiB
Gemma 3 1B it Q8_038.96%87.01%39.80%34.63%25.1 ms / 309.2 ms1.62 GiB1.00 GiB
Qwen3 0.6B Q8_031.60%82.25%32.11%26.41%34.0 ms / 445.1 ms1.81 GiB0.60 GiB

GGUF download and disk space · exact file bytes · 1 GiB = 1,024 MiB. 231 public items per model · context 8192 · direct decisions. Acceptance: top probability ≥ 0.80, candidate mass ≥ 0.05. GPU memory: whole-device samples every 200 ms. GPT-OSS: GGML_CUDA_DISABLE_GRAPHS=1. Protocol and report ↗

Download measured evidence: RTX 3060 summary JSON · Manifest

Typed-decisions · 400 cases / 2000 decisions · RTX 3080 10 GiB, WSL2 · 2026-09-26. Latency per five-question case. Protocol and report ↗

Download measured evidence: Gemma summary JSON · Qwen summary JSON · Gemma scored JSONL · Qwen scored JSONL · Manifest

Review all measured models ↗

Run the performance test on your machine

05 / START LOCAL

Your next decision
starts here.

Bring a GGUF checkpoint and a matching llama.cpp build. Define a question, run the CLI, and inspect the JSON.

Download the English example
TERMINAL / QUICK START
cargo run --release --features llama -- \
  --model /path/to/chat-model.gguf \
  --input examples/warehouse.json \
  --device cpu

Open source. MIT licensed.

Build with Rust, run your GGUF model, and integrate typed results into your application.

License details ↗