Skip to content

Benchmark results

Full test split of KRLabsOrg/tool-output-extraction-swebench (618 tool outputs from SWE-bench agent runs, 27 tool types), measured with benchmarks/compare.py on a consumer GPU in float32.

Metrics follow Squeez's own evaluation: predicted and gold lines are compared as sets of stripped lines. Recall is the share of gold lines kept, token reduction is measured with Headroom's own token estimator (markers included), and mandatory lost counts error / failure / traceback lines that were dropped.

Head to head at equal compression

Each method ranks lines with its own score and keeps the top share, so recall compares ranking quality directly. This is the project's go/no-go gate.

Lines dropped relevance_split (BM25) relevance_split (hybrid) Squeez 150M highlighter Squeez 32M pooled
50% 0.711 0.729 0.874 0.886
70% 0.597 0.629 0.847 0.842
90% 0.444 0.459 0.731 0.681

As shipped

The compressor with its safety rules (failure lines, whole tracebacks, ±2 context, first/last lines) and gates.

Configuration Recall Token reduction Mandatory lost p50 p95
relevance_split, BM25 (tail dropped)¹ 0.725 58.6% 7,953 1 ms 4 ms
relevance_split, hybrid (tail dropped)¹ 0.753 57.6% 7,771 2.2 s 8.6 s
Winnow, 150M highlighter, unbounded 0.836 39.7% 0 1.5 s 9.0 s
Winnow, 150M highlighter, 2,048-token budget 0.905 9.0% 0 1 ms² 0.8 s
Winnow, 32M pooled 0.847 41.4% 0 0.29 s 0.92 s

¹ Inside Headroom the low-relevance tail is Kompressed rather than dropped, so these token reductions are an upper bound. ² Most outputs exceed the budget and pass straight through to Headroom.

Takeaways

  • Squeez ranks lines far better than lexical or embedding similarity: 15-27 recall points over Headroom's built-in relevance_split at the same compression.
  • The safety rules work: zero failure or traceback lines lost, against ~7,800 for relevance_split.
  • The 32M pooled model makes it practical: ~5x faster than the 150M highlighter at the median and ~10x at p95, with the same recall and slightly better compression, so it can score every output instead of only small ones.

Reproduce

# Headroom baselines + 150M highlighter
python benchmarks/compare.py --device cuda

# 32M pooled model (rbk4209/winnow-pooled-32m, trained with training/kaggle_train_pooled.ipynb)
HEADROOM_WINNOW_MAX_TOKENS=1000000000 python benchmarks/compare.py --device cuda \
    --backend pooled --model-path rbk4209/winnow-pooled-32m \
    --methods squeez-raw,headroom-winnow