GSQ-RCO IQ3_S MTP
97.70 index · 90.24% HumanEval+ · 31.51s adaptive median. Minimum sampled free VRAM: 1.59 GiB. A strong first choice when this headroom is sufficient.
Qwen3.8-27B quants tested on an RTX 5070 Ti. GSQ-RCO IQ3_S MTP leads the measured practical/coding index at 97.70 and HumanEval+ at 148/164. Speed, memory, and long-context answers tell the rest of the story.
97.70 index · 90.24% HumanEval+ · 31.51s adaptive median. Minimum sampled free VRAM: 1.59 GiB. A strong first choice when this headroom is sufficient.
23.07s adaptive median · 95.00 index · 143/164 HumanEval+. About 4.09 GiB minimum free VRAM. Useful when latency and room for other workloads matter.
97.58 index · 143/164 HumanEval+ · 47.07s adaptive median. About 4.18 GiB minimum free VRAM. Pi control ties its index at displayed precision, with 136/164 coding passes.
These are workload-specific choices. The leader's gap to HauhauCS Q2_K_P is only 0.09 points, and the runner-up has a long-context failure. Fixed-seed repeats do not establish statistically significant ranking differences.
Bars use an absolute zero baseline. MTP entries are violet; the leading value is green. Timing is median request wall time with prompt-cache reuse, and output lengths vary.
Balanced index = 90% practical categories + 10% HumanEval+. Practical-only score renormalizes those categories to 100. Ties retain the same rank. Repository links resolve to the exact tested revision.
| Rank | Provider / quant / MTP | Index /100 | Practical /100 | HumanEval+ | Adaptive median | Min. free VRAM | Practical length finishes |
|---|---|---|---|---|---|---|---|
| 1 | GSQ-RCO · IQ3_S · MTP | 97.70 | 98.53 | 148/164 90.24% | 31.51s | 1.59 GiB | 3/72 |
| 2 | HauhauCS Aggressive · Q2_K_P · MTP | 97.61 | 98.90 | 141/164 85.98% | 6.90s empty final | 2.96 GiB | 0/72 |
| 3 | Pi control · IQ2_M | 97.58 | 99.21 | 136/164 82.93% | 39.53s | 4.62 GiB | 0/72 |
| 3 | Pi transfer · IQ3_XXS | 97.58 | 98.74 | 143/164 87.20% | 47.07s | 4.18 GiB | 0/72 |
| 5 | GSQ-RCO · IQ3_XXS · MTP | 97.37 | 98.50 | 143/164 87.20% | 30.52s | 3.30 GiB | 0/72 |
| 6 | GSQ-RCO · IQ3_XXS | 97.31 | 98.50 | 142/164 86.59% | 55.26s | 4.18 GiB | 0/72 |
| 7 | Unsloth · UD-Q2_K_XL | 97.09 | 98.39 | 140/164 85.37% | 42.32s | 4.84 GiB | 0/72 |
| 8 | OrcaRouter · IQ3_XXS · MTP | 97.00 | 98.08 | 143/164 87.20% | 26.28s | 2.38 GiB | 0/72 |
| 9 | HauhauCS Aggressive · IQ3_XS · MTP | 96.77 | 97.91 | 142/164 86.59% | 28.07s | 1.40 GiB | 3/72 |
| 10 | Swift · IQ2_S | 96.52 | 98.50 | 129/164 78.66% | 35.89s | 5.09 GiB | 0/72 |
| 11 | GSQ-RCO · IQ3_S | 96.44 | 97.33 | 145/164 88.41% | 53.75s | 2.62 GiB | 0/72 |
| 12 | GSQ-RCO · IQ2_S · MTP | 96.43 | 97.66 | 140/164 85.37% | 26.93s | 3.57 GiB | 3/72 |
| 13 | GSQ-RCO · IQ2_S | 95.54 | 96.68 | 140/164 85.37% | 35.96s | 4.96 GiB | 0/72 |
| 14 | HauhauCS Aggressive · IQ2_M · MTP | 95.26 | 95.95 | 146/164 89.02% | 25.19s | 2.97 GiB | 3/72 |
| 15 | GSQ-RCO · IQ2_XS · MTP | 95.00 | 95.87 | 143/164 87.20% | 23.07s | 4.09 GiB | 0/72 |
| 16 | Unsloth · UD-IQ2_S | 94.92 | 97.34 | 120/164 73.17% | 48.49s | 5.80 GiB | 3/72 |
| 17 | GSQ-RCO · IQ2_XS | 93.46 | 94.30 | 141/164 85.98% | 43.35s | 5.62 GiB | 3/72 |
All 1,224 practical responses contained final text. Length finishes flag incomplete API completion, not necessarily a content failure. Six configurations had three length finishes each, for 18/1,224 responses. The revised rubric caps unfinished responses at 80.
| Quant | Base → MTP median | Measured speedup | Index delta | Extra sampled VRAM |
|---|---|---|---|---|
| IQ2_XS | 43.35s → 23.07s | 1.88× | +1.54 | 1.53 GiB |
| IQ2_S | 35.96s → 26.93s | 1.34× | +0.89 | 1.39 GiB |
| IQ3_XXS | 55.26s → 30.52s | 1.81× | +0.06 | 0.89 GiB |
| IQ3_S | 53.75s → 31.51s | 1.71× | +1.26 | 1.03 GiB |
These are base/MTP artifact comparisons with adaptive settings and differing output lengths, not identical-file on/off experiments. No HauhauCS custom FastMTP patch or sidecar was used.
Pi control · IQ2_M
Pi control · IQ2_M
HauhauCS Aggressive · IQ3_XS · MTP
Pi control · IQ2_M
GSQ-RCO · IQ3_S · MTP
Every tied leader is listed. These are weighted practical/coding preferences, not separate evaluations. A role score does not override a failed long-context response.
One NVIDIA RTX 5070 Ti 16GB (16,303 MiB reported capacity), AMD Ryzen 5 5600G, llama.cpp 0.6.0 at commit d81235049384534c167caea52b85a694f6103d14. Each configuration ran three fixed and three adaptive dossier requests; 24 practical tasks three times; and one candidate on each of 164 HumanEval+ problems. Academic benchmark suites were not included.
Practical and HumanEval completion caps were 2,048 tokens, with a 512-token reasoning budget. The core dossier cap remained 3,072. Practical scoring revision 2.1 checks numeric equivalents, structured tool calls, schedule feasibility, negation, and selected robustness requirements. Legacy literal scores stay preserved as diagnostics and are not this leaderboard.
Context: the dossier exercised 57,135 prompt tokens at a 73,728-token allocation. Separate short-prompt allocation probes reached the per-model values below. The 14-marker dossier check is limited evidence; it is not validation of recall across 131K tokens.
Timing: three-request medians include cached prompt reuse. Request one processes the full dossier; requests two and three reuse roughly 57,131 prompt tokens. Faster wall time can reflect shorter output as well as faster generation. Empty final answers are excluded from successful-speed rankings.
Memory: headroom is the minimum sampled free GPU memory across adaptive repetitions; it is GPU-wide and includes runtime/KV allocations. File size is a separate metric.
Cost: $0 in cloud GPU rental for this local study. Electricity, hardware amortization, bandwidth and historic quant creation were not metered. This is not a claim of zero total project cost. The full study took 14h 20m wall time; the 14-entry queue took 11h 53m, including download, collection, scoring and the recovered interruption. Earlier audit/review work is excluded.
Business 20% · Research 20% · Knowledge 20% · Agent / tools 15% · Sysadmin 5% · HumanEval+ 10% · Executive / content 10%
| Configuration | Business | Research | Knowledge | Agent / tools | Sysadmin | Executive / content | HumanEval+ |
|---|---|---|---|---|---|---|---|
| GSQ-RCO · IQ3_S · MTP | 100.00 | 100.00 | 100.00 | 100.00 | 95.00 | 89.28 | 90.24 |
| HauhauCS Aggressive · Q2_K_P · MTP | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 90.08 | 85.98 |
| Pi control · IQ2_M | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 92.86 | 82.93 |
| Pi transfer · IQ3_XXS | 100.00 | 97.50 | 100.00 | 100.00 | 100.00 | 93.65 | 87.20 |
| GSQ-RCO · IQ3_XXS · MTP | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 86.51 | 87.20 |
| GSQ-RCO · IQ3_XXS | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 86.51 | 86.59 |
| Unsloth · UD-Q2_K_XL | 100.00 | 100.00 | 100.00 | 97.50 | 100.00 | 89.28 | 85.37 |
| OrcaRouter · IQ3_XXS · MTP | 100.00 | 100.00 | 100.00 | 97.50 | 100.00 | 86.51 | 87.20 |
| HauhauCS Aggressive · IQ3_XS · MTP | 95.00 | 100.00 | 100.00 | 100.00 | 95.00 | 93.65 | 86.59 |
| Swift · IQ2_S | 100.00 | 100.00 | 98.22 | 100.00 | 100.00 | 90.08 | 78.66 |
| GSQ-RCO · IQ3_S | 100.00 | 100.00 | 100.00 | 90.62 | 100.00 | 90.08 | 88.41 |
| GSQ-RCO · IQ2_S · MTP | 95.00 | 97.50 | 100.00 | 100.00 | 95.00 | 96.43 | 85.37 |
| GSQ-RCO · IQ2_S | 95.00 | 95.00 | 100.00 | 100.00 | 100.00 | 90.08 | 85.37 |
| HauhauCS Aggressive · IQ2_M · MTP | 91.67 | 100.00 | 100.00 | 97.50 | 95.00 | 86.51 | 89.02 |
| GSQ-RCO · IQ2_XS · MTP | 95.00 | 95.23 | 100.00 | 90.62 | 100.00 | 96.43 | 87.20 |
| Unsloth · UD-IQ2_S | 100.00 | 100.00 | 95.00 | 100.00 | 97.50 | 87.30 | 73.17 |
| GSQ-RCO · IQ2_XS | 86.67 | 97.73 | 100.00 | 90.62 | 95.00 | 96.43 | 85.98 |
Profile cells show ubatch / MTP depth. Core median generation tokens/s is API decode throughput, separate from request wall time. Literal marker hits are listed for the three adaptive responses.
| Configuration | GGUF file | Practical ubatch / MTP | Core adaptive ubatch / MTP | First dossier request | Median generation tok/s | Adaptive markers | Max allocation only |
|---|---|---|---|---|---|---|---|
| GSQ-RCO · IQ3_S · MTP | 11.29 GiB | 256 / 2 | 256 / 3 | 84.35s | 73.77 | 14, 14, 14/14 | 131,072 |
| HauhauCS Aggressive · Q2_K_P · MTP | 9.94 GiB | 256 / 2 | 512 / 3 | 56.98s | 83.37 | 0, 0, 0/14 | 131,072 |
| Pi control · IQ2_M | 9.32 GiB | 1024 / 0 | 512 / 0 | 93.80s | 41.55 | 14, 14, 14/14 | 131,072 |
| Pi transfer · IQ3_XXS | 9.40 GiB | 512 / 0 | 512 / 0 | 97.66s | 41.58 | 14, 14, 14/14 | 131,072 |
| GSQ-RCO · IQ3_XXS · MTP | 9.73 GiB | 1024 / 3 | 256 / 2 | 85.88s | 67.37 | 14, 14, 14/14 | 131,072 |
| GSQ-RCO · IQ3_XXS | 9.40 GiB | 1024 / 0 | 512 / 0 | 105.85s | 41.57 | 14, 14, 14/14 | 131,072 |
| Unsloth · UD-Q2_K_XL | 9.15 GiB | 1024 / 0 | 512 / 0 | 92.59s | 42.41 | 14, 14, 14/14 | 131,072 |
| OrcaRouter · IQ3_XXS · MTP | 10.84 GiB | 256 / 2 | 256 / 3 | 78.76s | 72.35 | 14, 14, 14/14 | 131,072 |
| HauhauCS Aggressive · IQ3_XS · MTP | 11.34 GiB | 256 / 2 | 1024 / 3 | 77.66s | 69.44 | 14, 14, 14/14 | 131,072 |
| Swift · IQ2_S | 9.22 GiB | 1024 / 0 | 512 / 0 | 89.36s | 42.90 | 14, 14, 14/14 | 131,072 |
| GSQ-RCO · IQ3_S | 10.96 GiB | 1024 / 0 | 512 / 0 | 101.91s | 39.11 | 13, 13, 13/14 | 131,072 |
| GSQ-RCO · IQ2_S · MTP | 8.95 GiB | 1024 / 2 | 1024 / 2 | 79.88s | 67.38 | 14, 14, 14/14 | 131,072 |
| GSQ-RCO · IQ2_S | 8.62 GiB | 1024 / 0 | 512 / 0 | 87.16s | 43.00 | 14, 14, 14/14 | 131,072 |
| HauhauCS Aggressive · IQ2_M · MTP | 9.61 GiB | 256 / 2 | 1024 / 3 | 75.50s | 74.27 | 14, 14, 14/14 | 131,072 |
| GSQ-RCO · IQ2_XS · MTP | 8.17 GiB | 512 / 3 | 1024 / 3 | 75.89s | 72.72 | 14, 14, 14/14 | 131,072 |
| Unsloth · UD-IQ2_S | 7.80 GiB | 1024 / 0 | 512 / 0 | 98.79s | 44.21 | 14, 14, 14/14 | 131,072 |
| GSQ-RCO · IQ2_XS | 7.84 GiB | 1024 / 0 | 512 / 0 | 95.17s | 44.25 | 14, 14, 14/14 | 131,072 |
All generation archives and evaluation manifests were checked before building this report. The original three cached GGUFs remain preserved; all fourteen temporary quants were removed after verified collection and scoring. The scorer-image reconstruction is recorded per affected configuration. No previous scores were overwritten.