Invited at the University of Lausanne by Dato Gogishvili and Martin Müller (Team M3, Institute of Geography and Sustainability, 29 June 2026, 16:00–17:30), this guest talk presents the Canadian Climate Framing (CCF) database — 266,000 articles, 9.2 million sentences, 20 newspapers, 1978–2024, EN + FR — and the LLM Tool open pipeline that produced its 65 annotation categories. One corpus, one tool, two angles on the same methodological commitment: large-scale qualitative data that stays auditable.
A first-of-its-kind, open climate media database

The CCF database is the first openly released, machine-learning-annotated, pan-Canadian climate media corpus, with 47 years of historical depth and bilingual EN/FR coverage. Frames, actors, events and locations are annotated at the sentence level, so questions about thematic frame evolution since 1978, regional patterns of fossil-fuel dependence, who structures climate politics (the 110-policymaker network of 2024) and how a media cascade propagates (the Health cascade of Summer 2023) can all be answered from the same resource — quantitatively or qualitatively, through a web explorer and an open API.

A local, auditable pipeline that transfers beyond climate

Behind the corpus, LLM Tool turns prompting, cleaning, training and inference into one logged local workflow, with five modes (Annotator, Annotator Factory, Training Arena, BERT Annotation Studio, Validation Lab). Distilled XLM-RoBERTa classifiers match GPT-5-class accuracy on the human-consensus test set while running 109–395× faster per sentence. Prompts, schemas, model versions, hyperparameters, failures and outputs are recorded as part of the method, so the workflow can be inspected, adapted and reused beyond climate for any sentence-level annotation task.

CCF & LLM Tool - UNIL Lausanne 2026
PDF Open
Scroll Navigate between slides Click Interact with data Fullscreen for a better experience
Use arrow keys or click to navigate through the slides.