CCF & LLM Tool: machine learning for climate research and social science
The CCF database is the first openly released, machine-learning-annotated, pan-Canadian climate media corpus, with 47 years of historical depth and bilingual EN/FR coverage. Frames, actors, events and locations are annotated at the sentence level, so questions about thematic frame evolution since 1978, regional patterns of fossil-fuel dependence, who structures climate politics (the 110-policymaker network of 2024) and how a media cascade propagates (the Health cascade of Summer 2023) can all be answered from the same resource — quantitatively or qualitatively, through a web explorer and an open API.
Behind the corpus, LLM Tool turns prompting, cleaning, training and inference into one logged local workflow, with five modes (Annotator, Annotator Factory, Training Arena, BERT Annotation Studio, Validation Lab). Distilled XLM-RoBERTa classifiers match GPT-5-class accuracy on the human-consensus test set while running 109–395× faster per sentence. Prompts, schemas, model versions, hyperparameters, failures and outputs are recorded as part of the method, so the workflow can be inspected, adapted and reused beyond climate for any sentence-level annotation task.