Skip to content

Backends

backends

Span-model backends: lazy, pinned, and isolated from the request path.

The compressor only needs one thing from a model — character spans of content that matter for query — so backends implement the tiny :class:SpanBackend protocol. That keeps the pruning logic testable with a fake backend and lets models be swapped without touching the compressor.

Two backends ship, selected by $HEADROOM_WINNOW_BACKEND (see :func:make_backend):

  • :class:PooledBackend (default) loads a Squeez pooled line classifier. The default model, rbk4209/winnow-pooled-32m, is a 32M encoder trained for this project (training/).
  • :class:HighlighterBackend wraps KRLabsOrg/verbatim-rag-modern-bert-v2, the ~150M ModernBERT span model Squeez itself uses for its extractive path.

Three rules shape both:

  • Lazy. Headroom instantiates entry-point classes during discovery, for every process that builds a router, selected or not. Loading a model there would tax every Headroom user who merely has this package installed, so nothing heavy happens until the first find_spans.
  • Pinned. The models ship their own code (trust_remote_code=True), so weights and tokenizer load at a fixed commit. The highlighter's process() would otherwise fetch its tokenizer from the moving main branch, so we load it ourselves at the same revision and hand it over.
  • Loud once, then quiet. Any load failure (no torch, no network, bad revision) is raised as :class:BackendUnavailableError; the compressor turns that into a single warning and passes through for the rest of the process.

Long outputs need no chunking here: both models window long inputs themselves and map results back onto content. Shrinking the highlighter's window to go faster is not an option: on the Squeez test split, gold-line recall at 90% compression falls from 0.78 (8192) to 0.47 (2048) and 0.31 (512), because the model needs the whole output in view.

What bounds latency instead is max_input_tokens: the largest input a backend can score within the latency budget on its device. Each backend class carries its own GPU and CPU budgets, sized from measured forward times. The 150M highlighter needs ~0.5 s for 2,048 tokens on a consumer GPU and ~4 s for 8,192, so it stays small; the 32M pooled classifier scores the whole Squeez test split at 0.29 s p50 / 0.92 s p95 on the same GPU, so its GPU budget covers practically any tool output. The compressor passes anything larger straight to Headroom's own path.

BackendUnavailableError

Bases: RuntimeError

The backend cannot run in this process (missing deps, download failed, ...).

SpanBackend

Bases: Protocol

Anything that can point at the parts of content relevant to query.

find_spans(query, content)

Return [start, end) character spans of content relevant to query.

Raises:

Type Description
BackendUnavailableError

If the backend can never run in this process.

HighlighterBackend(model_id=None, revision=None, *, device=None, dtype=None, threshold=0.1, min_span_chars=10, merge_gap_chars=20, max_length=8192, doc_stride=256, max_input_tokens=None)

Bases: _TransformersBackend

Verbatim-RAG ModernBERT highlighter, loaded on first use.

Parameters:

Name Type Description Default
model_id str | None

Hugging Face repo id. Defaults to $HEADROOM_WINNOW_MODEL or :data:DEFAULT_HIGHLIGHTER_MODEL.

None
revision str | None

Commit to load. Defaults to $HEADROOM_WINNOW_REVISION, or :data:DEFAULT_HIGHLIGHTER_REVISION when the default model is used. A custom model without a revision loads main.

None
device str | None

"cpu", "cuda", ... or "auto" (CUDA when available, else CPU). Defaults to $HEADROOM_WINNOW_DEVICE or "auto".

None
dtype str | None

"float32" or "float16". Defaults to $HEADROOM_WINNOW_DTYPE or "float32": float16 is not a safe default, since GPUs without tensor cores (e.g. GTX 16xx) run it several times slower than float32.

None
threshold float

Per-token probability for a token to join a span. Squeez's recall-tuned value for technical content.

0.1
min_span_chars int

Spans shorter than this are discarded by the model.

10
merge_gap_chars int

Spans closer than this are merged by the model.

20
max_length int

Token window per forward pass.

8192
doc_stride int

Token overlap between consecutive windows.

256
max_input_tokens int | None

Largest input to score. Defaults to $HEADROOM_WINNOW_MAX_TOKENS, else :attr:cuda_token_budget or :attr:cpu_token_budget for the resolved device.

None

find_spans(query, content)

Run the highlighter and return its character spans over content.

PooledBackend(model_id=None, revision=None, *, device=None, dtype=None, threshold=0.5, max_input_tokens=None)

Bases: _TransformersBackend

Squeez pooled line classifier (squeez.encoder.train --classifier-type pooled).

The model scores whole lines, so each kept line becomes one span covering it. It is trained on the same data as the highlighter but can sit on a much smaller encoder (e.g. jhu-clsp/ettin-encoder-32m), which is what lets the token budget grow.

Parameters:

Name Type Description Default
model_id str | None

Local directory or Hugging Face repo with the trained model (it ships modeling_squeez_pooled.py for trust_remote_code). Defaults to $HEADROOM_WINNOW_MODEL or :data:DEFAULT_POOLED_MODEL.

None
revision str | None

Commit to load for a hub repo. Defaults to $HEADROOM_WINNOW_REVISION, or :data:DEFAULT_POOLED_REVISION when the default model is used.

None
device str | None

As for :class:HighlighterBackend.

None
dtype str | None

As for :class:HighlighterBackend.

None
threshold float

Line probability at or above which a line is kept.

0.5
max_input_tokens int | None

As for :class:HighlighterBackend.

None

line_probabilities(query, content)

Return per-line relevance over content.split("\n").

find_spans(query, content)

Return one span per line the classifier keeps.

lines_to_spans(content, probs, threshold)

Map per-line probabilities over content.split("\n") to character spans.

Parameters:

Name Type Description Default
content str

The scored text.

required
probs list[float]

One probability per "\n"-separated line.

required
threshold float

Lines at or above it (and not blank) become spans.

required

Returns:

Type Description
list[tuple[int, int]]

[start, end) spans, one per kept line, newline excluded.

make_backend()

Build the backend named by $HEADROOM_WINNOW_BACKEND (default pooled).

Constructing a backend loads nothing, so this is safe during discovery.