Protocol v2.1 · October 9–10, 2026

Seventeen configurations.
One 16GB GPU.

Qwen3.8-27B quants tested on an RTX 5070 Ti. GSQ-RCO IQ3_S MTP leads the measured practical/coding index at 97.70 and HumanEval+ at 148/164. Speed, memory, and long-context answers tell the rest of the story.

17 / 17Configurations scored
1,224Practical responses · 24 tasks × 3
2,788HumanEval+ candidates · 164 each
14h 20mFull benchmark + queue wall time
A fast empty answer is still an empty answer. HauhauCS Q2_K_P scored 97.61 on the practical/coding index, but returned no final answer in all six long-context requests. Its 6.90s adaptive median is retained in the data and excluded from successful-speed rankings. The index does not include long-context quality.
What to try first

Quality, speed, and headroom are different choices.

Measured quality leader

GSQ-RCO IQ3_S MTP

97.70 index · 90.24% HumanEval+ · 31.51s adaptive median. Minimum sampled free VRAM: 1.59 GiB. A strong first choice when this headroom is sufficient.

Fastest successful dossier output

GSQ-RCO IQ2_XS MTP

23.07s adaptive median · 95.00 index · 143/164 HumanEval+. About 4.09 GiB minimum free VRAM. Useful when latency and room for other workloads matter.

Higher quality with more headroom

Pi transfer IQ3_XXS

97.58 index · 143/164 HumanEval+ · 47.07s adaptive median. About 4.18 GiB minimum free VRAM. Pi control ties its index at displayed precision, with 136/164 coding passes.

These are workload-specific choices. The leader's gap to HauhauCS Q2_K_P is only 0.09 points, and the runner-up has a long-context failure. Fixed-seed repeats do not establish statistically significant ranking differences.

Explore the field

One ranking cannot describe every workload.

Bars use an absolute zero baseline. MTP entries are violet; the leading value is green. Timing is median request wall time with prompt-cache reuse, and output lengths vary.

All seventeen configurations

Full results

Balanced index = 90% practical categories + 10% HumanEval+. Practical-only score renormalizes those categories to 100. Ties retain the same rank. Repository links resolve to the exact tested revision.

RankProvider / quant / MTPIndex /100Practical /100HumanEval+Adaptive medianMin. free VRAMPractical length finishes
1GSQ-RCO · IQ3_S · MTP97.7098.53148/164 90.24%31.51s1.59 GiB3/72
2HauhauCS Aggressive · Q2_K_P · MTP97.6198.90141/164 85.98%6.90s empty final2.96 GiB0/72
3Pi control · IQ2_M97.5899.21136/164 82.93%39.53s4.62 GiB0/72
3Pi transfer · IQ3_XXS97.5898.74143/164 87.20%47.07s4.18 GiB0/72
5GSQ-RCO · IQ3_XXS · MTP97.3798.50143/164 87.20%30.52s3.30 GiB0/72
6GSQ-RCO · IQ3_XXS97.3198.50142/164 86.59%55.26s4.18 GiB0/72
7Unsloth · UD-Q2_K_XL97.0998.39140/164 85.37%42.32s4.84 GiB0/72
8OrcaRouter · IQ3_XXS · MTP97.0098.08143/164 87.20%26.28s2.38 GiB0/72
9HauhauCS Aggressive · IQ3_XS · MTP96.7797.91142/164 86.59%28.07s1.40 GiB3/72
10Swift · IQ2_S96.5298.50129/164 78.66%35.89s5.09 GiB0/72
11GSQ-RCO · IQ3_S96.4497.33145/164 88.41%53.75s2.62 GiB0/72
12GSQ-RCO · IQ2_S · MTP96.4397.66140/164 85.37%26.93s3.57 GiB3/72
13GSQ-RCO · IQ2_S95.5496.68140/164 85.37%35.96s4.96 GiB0/72
14HauhauCS Aggressive · IQ2_M · MTP95.2695.95146/164 89.02%25.19s2.97 GiB3/72
15GSQ-RCO · IQ2_XS · MTP95.0095.87143/164 87.20%23.07s4.09 GiB0/72
16Unsloth · UD-IQ2_S94.9297.34120/164 73.17%48.49s5.80 GiB3/72
17GSQ-RCO · IQ2_XS93.4694.30141/164 85.98%43.35s5.62 GiB3/72

All 1,224 practical responses contained final text. Length finishes flag incomplete API completion, not necessarily a content failure. Six configurations had three length finishes each, for 18/1,224 responses. The revised rubric caps unfinished responses at 80.

Matched GSQ-RCO published pairs

MTP reduced measured response time in all four pairs.

QuantBase → MTP medianMeasured speedupIndex deltaExtra sampled VRAM
IQ2_XS43.35s → 23.07s1.88×+1.541.53 GiB
IQ2_S35.96s → 26.93s1.34×+0.891.39 GiB
IQ3_XXS55.26s → 30.52s1.81×+0.060.89 GiB
IQ3_S53.75s → 31.51s1.71×+1.261.03 GiB

These are base/MTP artifact comparisons with adaptive settings and differing output lengths, not identical-file on/off experiments. No HauhauCS custom FastMTP patch or sidecar was used.

Personalized routing

Role leaders under the saved weighting

C-Suite / Analyst

99.64

Pi control · IQ2_M

Researcher / Scout

99.29

Pi control · IQ2_M

Hermes Coordinator

98.13

HauhauCS Aggressive · IQ3_XS · MTP

Librarian / Knowledge

98.93

Pi control · IQ2_M

Forge / Coder

94.12

GSQ-RCO · IQ3_S · MTP

Every tied leader is listed. These are weighted practical/coding preferences, not separate evaluations. A role score does not override a failed long-context response.

What was actually measured

Same local hardware. Versioned protocol.

One NVIDIA RTX 5070 Ti 16GB (16,303 MiB reported capacity), AMD Ryzen 5 5600G, llama.cpp 0.6.0 at commit d81235049384534c167caea52b85a694f6103d14. Each configuration ran three fixed and three adaptive dossier requests; 24 practical tasks three times; and one candidate on each of 164 HumanEval+ problems. Academic benchmark suites were not included.

Practical and HumanEval completion caps were 2,048 tokens, with a 512-token reasoning budget. The core dossier cap remained 3,072. Practical scoring revision 2.1 checks numeric equivalents, structured tool calls, schedule feasibility, negation, and selected robustness requirements. Legacy literal scores stay preserved as diagnostics and are not this leaderboard.

Context: the dossier exercised 57,135 prompt tokens at a 73,728-token allocation. Separate short-prompt allocation probes reached the per-model values below. The 14-marker dossier check is limited evidence; it is not validation of recall across 131K tokens.

Timing: three-request medians include cached prompt reuse. Request one processes the full dossier; requests two and three reuse roughly 57,131 prompt tokens. Faster wall time can reflect shorter output as well as faster generation. Empty final answers are excluded from successful-speed rankings.

Memory: headroom is the minimum sampled free GPU memory across adaptive repetitions; it is GPU-wide and includes runtime/KV allocations. File size is a separate metric.

Cost: $0 in cloud GPU rental for this local study. Electricity, hardware amortization, bandwidth and historic quant creation were not metered. This is not a claim of zero total project cost. The full study took 14h 20m wall time; the 14-entry queue took 11h 53m, including download, collection, scoring and the recovered interruption. Earlier audit/review work is excluded.

Category scores and exact weights

Business 20% · Research 20% · Knowledge 20% · Agent / tools 15% · Sysadmin 5% · HumanEval+ 10% · Executive / content 10%

ConfigurationBusinessResearchKnowledgeAgent / toolsSysadminExecutive / contentHumanEval+
GSQ-RCO · IQ3_S · MTP100.00100.00100.00100.0095.0089.2890.24
HauhauCS Aggressive · Q2_K_P · MTP100.00100.00100.00100.00100.0090.0885.98
Pi control · IQ2_M100.00100.00100.00100.00100.0092.8682.93
Pi transfer · IQ3_XXS100.0097.50100.00100.00100.0093.6587.20
GSQ-RCO · IQ3_XXS · MTP100.00100.00100.00100.00100.0086.5187.20
GSQ-RCO · IQ3_XXS100.00100.00100.00100.00100.0086.5186.59
Unsloth · UD-Q2_K_XL100.00100.00100.0097.50100.0089.2885.37
OrcaRouter · IQ3_XXS · MTP100.00100.00100.0097.50100.0086.5187.20
HauhauCS Aggressive · IQ3_XS · MTP95.00100.00100.00100.0095.0093.6586.59
Swift · IQ2_S100.00100.0098.22100.00100.0090.0878.66
GSQ-RCO · IQ3_S100.00100.00100.0090.62100.0090.0888.41
GSQ-RCO · IQ2_S · MTP95.0097.50100.00100.0095.0096.4385.37
GSQ-RCO · IQ2_S95.0095.00100.00100.00100.0090.0885.37
HauhauCS Aggressive · IQ2_M · MTP91.67100.00100.0097.5095.0086.5189.02
GSQ-RCO · IQ2_XS · MTP95.0095.23100.0090.62100.0096.4387.20
Unsloth · UD-IQ2_S100.00100.0095.00100.0097.5087.3073.17
GSQ-RCO · IQ2_XS86.6797.73100.0090.6295.0096.4385.98
Profiles, timing, generation, and context detail

Profile cells show ubatch / MTP depth. Core median generation tokens/s is API decode throughput, separate from request wall time. Literal marker hits are listed for the three adaptive responses.

ConfigurationGGUF filePractical ubatch / MTPCore adaptive ubatch / MTPFirst dossier requestMedian generation tok/sAdaptive markersMax allocation only
GSQ-RCO · IQ3_S · MTP11.29 GiB256 / 2256 / 384.35s73.7714, 14, 14/14131,072
HauhauCS Aggressive · Q2_K_P · MTP9.94 GiB256 / 2512 / 356.98s83.370, 0, 0/14131,072
Pi control · IQ2_M9.32 GiB1024 / 0512 / 093.80s41.5514, 14, 14/14131,072
Pi transfer · IQ3_XXS9.40 GiB512 / 0512 / 097.66s41.5814, 14, 14/14131,072
GSQ-RCO · IQ3_XXS · MTP9.73 GiB1024 / 3256 / 285.88s67.3714, 14, 14/14131,072
GSQ-RCO · IQ3_XXS9.40 GiB1024 / 0512 / 0105.85s41.5714, 14, 14/14131,072
Unsloth · UD-Q2_K_XL9.15 GiB1024 / 0512 / 092.59s42.4114, 14, 14/14131,072
OrcaRouter · IQ3_XXS · MTP10.84 GiB256 / 2256 / 378.76s72.3514, 14, 14/14131,072
HauhauCS Aggressive · IQ3_XS · MTP11.34 GiB256 / 21024 / 377.66s69.4414, 14, 14/14131,072
Swift · IQ2_S9.22 GiB1024 / 0512 / 089.36s42.9014, 14, 14/14131,072
GSQ-RCO · IQ3_S10.96 GiB1024 / 0512 / 0101.91s39.1113, 13, 13/14131,072
GSQ-RCO · IQ2_S · MTP8.95 GiB1024 / 21024 / 279.88s67.3814, 14, 14/14131,072
GSQ-RCO · IQ2_S8.62 GiB1024 / 0512 / 087.16s43.0014, 14, 14/14131,072
HauhauCS Aggressive · IQ2_M · MTP9.61 GiB256 / 21024 / 375.50s74.2714, 14, 14/14131,072
GSQ-RCO · IQ2_XS · MTP8.17 GiB512 / 31024 / 375.89s72.7214, 14, 14/14131,072
Unsloth · UD-IQ2_S7.80 GiB1024 / 0512 / 098.79s44.2114, 14, 14/14131,072
GSQ-RCO · IQ2_XS7.84 GiB1024 / 0512 / 095.17s44.2514, 14, 14/14131,072
Limits and reproducibility

Keep the conclusions inside the evidence.

All generation archives and evaluation manifests were checked before building this report. The original three cached GGUFs remain preserved; all fourteen temporary quants were removed after verified collection and scoring. The scorer-image reconstruction is recorded per affected configuration. No previous scores were overwritten.