Training a model¶
The recommended pooled model was trained with a single notebook on a free Kaggle
GPU: training/kaggle_train_pooled.ipynb.
What it does¶
- Clones Squeez at a pinned commit and uses its own training code, so the recipe is reproducible.
- Downloads
KRLabsOrg/tool-output-extraction-swebench: 10,508 training, 240 dev and 618 test examples across 27 tool types. - Fine-tunes
jhu-clsp/ettin-encoder-32mas a pooled line classifier: 4,096-token windows, effective batch 96, learning rate 5e-5, 3 epochs, fp16. - Evaluates on the held-out test split and zips the model.
It takes about six hours on one T4. Validation loss fell from 0.63 to 0.31.
Running it¶
- Create a Kaggle notebook from the file (or push it with the Kaggle CLI using
training/kernel-metadata.json). - Settings → Accelerator → GPU T4, and Settings → Internet on.
- Run All, then download
squeez_pooled_ettin32m.zipfrom the Output tab.
Two fixes the notebook applies
- Training is pinned to one GPU. Kaggle's T4 ×2 makes the Hugging Face
Trainer wrap the model in
DataParallel, which hides an attribute Squeez's loss function reads. - Squeez's checkpoint hook passes a
metricsargument thattransformers5.2 no longer accepts. The notebook patches the call so it forwards whatever the Trainer passes.
Trying other encoders¶
Change --base-model in the training cell. Larger encoders
(answerdotai/ModernBERT-base) trade speed for capacity; smaller ones
(jhu-clsp/ettin-encoder-17m) go the other way. Measure any new model with: