Presented at CPSA (Ottawa, June 2-4, 2026), this presentation introduces LLM Tool: a hybrid annotation pipeline built for computational social science. It pairs the interpretive quality of large language models with the speed of a distilled BERT classifier, so that a few thousand LLM-annotated sentences can train a classifier that labels millions, while evaluation stays anchored in human judgment. The talk walks through the distillation idea, the validation design, what the benchmark shows, and two databases built with these methods (CCF and YouPol).
Knowledge distillation, quality at the speed of a small model

A teacher LLM reads a codebook and annotates ~38,000 sentences; a student XLM-RoBERTa classifier learns the labeling logic and then annotates the full corpus at roughly 50 ms per sentence, a 109-395x speedup over LLM inference, with human-coded labels kept at the evaluative center.

Validated against human consensus, task complexity drives the gap

Five LLM annotators (27-120B) were benchmarked on a 4-task coding schema against consensus labels from four human experts on an isolated test set. Performance splits into two tiers, and task complexity, not annotator size, drives the gap: on the 21-label policy themes the open-weight GPT-OSS even overtakes GPT-5.

LLM Tool - CPSA 2026
PDF Open
Scroll Navigate between slides Click Interact with data Fullscreen for a better experience
Use arrow keys or click to navigate through the slides.