▚ LOCAL LLM BENCHMARK SUITE // LFU CACHE & ACID AUDIT

// BENCHMARK_RESULTS .json

8 models graded on a strict 5-pillar / 100-pt rubric · O(1) LFU + ACID transactions · M3 Max · LM Studio  ·  TOP: DeepSeek V4 Flash (CLOUD baseline)
Models Tested
8
Top Score
91
Average
66.9
Prod-Ready
1/8

▮ Score vs Throughput (tok/sec)

Local models only — cloud baseline (DeepSeek) excluded from speed axis. Gemma 4 bars flagged ⚠ (GPU-offload suspect).

▮ 5-Pillar Radar — Top 3

Each pillar scored 0–20. Outer = stronger.

▮ LEADERBOARD

#ModelSpeedScoreVerdictBest For
#1
DeepSeek V4 Flash (CLOUD baseline)CLOUD
n/a (cloud)
⚠ speed suspect
N/A t/s
91
PROD Reference-quality baseline (91/100) — the bar the local models are mea… DECODE ▸
#2
Qwen 3.6 35B-A3B
6-bit MLX
68.9 t/s
82
FLAWS Solid daily-driver scaffolding for ACID/async patterns — produces runn… DECODE ▸
#3
Gemma 4 31B
GGUF
⚠ speed suspect
10.1 t/s
78
FLAWS Clean, correct, runnable code with solid O(1) structure and good concu… DECODE ▸
#4
Gemma 4 31B QAT
QAT GGUF
⚠ speed suspect
15.0 t/s
70
CRIT Runnable and structurally sound, but the LFU eviction has a stale-min_… DECODE ▸
#5
KAT-Coder v2.5 Dev XL
MLX
65.3 t/s
65
CRIT Promising code-design instincts (cleanest abstractions and best transa… DECODE ▸
#6
Qwen 3.6 35B-A3B
4-bit MLX
83.3 t/s
57
CRIT Not recommended for systems code as-is. The 4-bit quant degrades logic… DECODE ▸
#7
Qwen 3.6 35B-A3B (uncensored hauhaucs aggressive)
GGUF
62.5 t/s
49
CRIT Not recommended for production code. Reasonable API shape and correctl… DECODE ▸
#8
Gemma 4 12B Coder (fable5-composer2.5-v1-uncensored-heretic merge)
mxfp8 MLX
25.3 t/s
43
CRIT Not usable as-is — the cache cannot store its first key and the backgr… DECODE ▸