Backends¶
backends
¶
Span-model backends: lazy, pinned, and isolated from the request path.
The compressor only needs one thing from a model — character spans of
content that matter for query — so backends implement the tiny
:class:SpanBackend protocol. That keeps the pruning logic testable with a
fake backend and lets models be swapped without touching the compressor.
Two backends ship, selected by $HEADROOM_WINNOW_BACKEND (see
:func:make_backend):
- :class:
PooledBackend(default) loads a Squeez pooled line classifier. The default model,rbk4209/winnow-pooled-32m, is a 32M encoder trained for this project (training/). - :class:
HighlighterBackendwrapsKRLabsOrg/verbatim-rag-modern-bert-v2, the ~150M ModernBERT span model Squeez itself uses for its extractive path.
Three rules shape both:
- Lazy. Headroom instantiates entry-point classes during discovery, for
every process that builds a router, selected or not. Loading a model there
would tax every Headroom user who merely has this package installed, so
nothing heavy happens until the first
find_spans. - Pinned. The models ship their own code (
trust_remote_code=True), so weights and tokenizer load at a fixed commit. The highlighter'sprocess()would otherwise fetch its tokenizer from the movingmainbranch, so we load it ourselves at the same revision and hand it over. - Loud once, then quiet. Any load failure (no torch, no network, bad
revision) is raised as :class:
BackendUnavailableError; the compressor turns that into a single warning and passes through for the rest of the process.
Long outputs need no chunking here: both models window long inputs themselves
and map results back onto content. Shrinking the highlighter's window to go
faster is not an option: on the Squeez test split, gold-line recall at 90%
compression falls from 0.78 (8192) to 0.47 (2048) and 0.31 (512), because the
model needs the whole output in view.
What bounds latency instead is max_input_tokens: the largest input a
backend can score within the latency budget on its device. Each backend class
carries its own GPU and CPU budgets, sized from measured forward times. The
150M highlighter needs ~0.5 s for 2,048 tokens on a consumer GPU and ~4 s for
8,192, so it stays small; the 32M pooled classifier scores the whole Squeez test
split at 0.29 s p50 / 0.92 s p95 on the same GPU, so its GPU budget covers
practically any tool output. The compressor passes anything larger straight to
Headroom's own path.
BackendUnavailableError
¶
Bases: RuntimeError
The backend cannot run in this process (missing deps, download failed, ...).
SpanBackend
¶
Bases: Protocol
Anything that can point at the parts of content relevant to query.
find_spans(query, content)
¶
Return [start, end) character spans of content relevant to query.
Raises:
| Type | Description |
|---|---|
BackendUnavailableError
|
If the backend can never run in this process. |
HighlighterBackend(model_id=None, revision=None, *, device=None, dtype=None, threshold=0.1, min_span_chars=10, merge_gap_chars=20, max_length=8192, doc_stride=256, max_input_tokens=None)
¶
Bases: _TransformersBackend
Verbatim-RAG ModernBERT highlighter, loaded on first use.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_id
|
str | None
|
Hugging Face repo id. Defaults to |
None
|
revision
|
str | None
|
Commit to load. Defaults to |
None
|
device
|
str | None
|
|
None
|
dtype
|
str | None
|
|
None
|
threshold
|
float
|
Per-token probability for a token to join a span. Squeez's recall-tuned value for technical content. |
0.1
|
min_span_chars
|
int
|
Spans shorter than this are discarded by the model. |
10
|
merge_gap_chars
|
int
|
Spans closer than this are merged by the model. |
20
|
max_length
|
int
|
Token window per forward pass. |
8192
|
doc_stride
|
int
|
Token overlap between consecutive windows. |
256
|
max_input_tokens
|
int | None
|
Largest input to score. Defaults to
|
None
|
find_spans(query, content)
¶
Run the highlighter and return its character spans over content.
PooledBackend(model_id=None, revision=None, *, device=None, dtype=None, threshold=0.5, max_input_tokens=None)
¶
Bases: _TransformersBackend
Squeez pooled line classifier (squeez.encoder.train --classifier-type pooled).
The model scores whole lines, so each kept line becomes one span covering
it. It is trained on the same data as the highlighter but can sit on a much
smaller encoder (e.g. jhu-clsp/ettin-encoder-32m), which is what lets
the token budget grow.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_id
|
str | None
|
Local directory or Hugging Face repo with the trained model
(it ships |
None
|
revision
|
str | None
|
Commit to load for a hub repo. Defaults to
|
None
|
device
|
str | None
|
As for :class: |
None
|
dtype
|
str | None
|
As for :class: |
None
|
threshold
|
float
|
Line probability at or above which a line is kept. |
0.5
|
max_input_tokens
|
int | None
|
As for :class: |
None
|
lines_to_spans(content, probs, threshold)
¶
Map per-line probabilities over content.split("\n") to character spans.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
content
|
str
|
The scored text. |
required |
probs
|
list[float]
|
One probability per |
required |
threshold
|
float
|
Lines at or above it (and not blank) become spans. |
required |
Returns:
| Type | Description |
|---|---|
list[tuple[int, int]]
|
|
make_backend()
¶
Build the backend named by $HEADROOM_WINNOW_BACKEND (default pooled).
Constructing a backend loads nothing, so this is safe during discovery.