Skip to content

Configuration

Winnow is configured through environment variables, so it works the same under the Headroom proxy and in your own code. Every variable is optional.

Variable Default Meaning
HEADROOM_WINNOW_BACKEND pooled pooled or highlighter
HEADROOM_WINNOW_MODEL the backend's published model Hub id or local path
HEADROOM_WINNOW_REVISION pinned commit of the default model Model revision to load
HEADROOM_WINNOW_DEVICE auto auto (CUDA when available), cuda or cpu
HEADROOM_WINNOW_DTYPE float32 float16 is faster only on GPUs with tensor cores
HEADROOM_WINNOW_MAX_TOKENS per backend and device Largest output the model scores; larger ones go to Headroom

Token budgets

Each backend knows how much it can score within a reasonable latency on each kind of device. Outputs above the budget skip the model entirely and go to Headroom's own path, so Winnow never stalls the agent.

Backend GPU budget CPU budget
pooled 32,768 tokens 2,048 tokens
highlighter 2,048 tokens 512 tokens

Override with HEADROOM_WINNOW_MAX_TOKENS if your hardware is faster or slower.

Pruning settings

The pruning rules can be tuned in code through WinnowSettings:

from headroom_winnow import WinnowCompressor, WinnowSettings

winnow = WinnowCompressor(
    settings=WinnowSettings(
        min_lines=40,  # shorter outputs pass through
        context_lines=2,  # neighbours kept around every kept line
        edge_lines=2,  # first/last lines always kept
        min_savings=0.2,  # pass through if pruning saves less than this
        max_tokens=None,  # None = use the backend's budget
    )
)

The defaults are recall-first: dropping a line the agent needed costs a retrieval round trip or a re-run, which is far more expensive than a few spare tokens.