Antisemitism in the Canadian Hansard · Orders of worth · Coding framework
How this corpus was assembled, what was checked, and what it cannot tell you. Every count on this page comes from the scripts named at the end, and each can be re-run.
Canadian parliamentary debate is not published as one dataset. Four sources cover different chambers and periods, and they differ in how they segment and transcribe speech.
| Source | Covers | Form | Scanned |
|---|---|---|---|
| Linked Parliamentary Data Project | House, 1901–2019 | Daily CSV, OCR of scanned Hansard | 4,316,216 speeches 606.2M words |
| openparliament.ca | House 1994–2026, committees 2006–2026 | PostgreSQL dump, native digital | 3,869,693 statements |
| sencanada.ca | Senate, 1996–2026 | HTML pages, scraped | 256,898 statements 2,149 sitting days |
| Canadiana (parl.canadiana.ca) | Senate to 1996; House to 1900 | Requested, not yet held | — |
Joining these gives the House of Commons continuously from 1901 to August 2026, a hundred and twenty-five years in one corpus. Records to 1993 come from the Linked Parliamentary Data Project and from 1994 onward from openparliament. LiPaD also covers 1994 to 2019, and that overlap is deliberately excluded so no statement is counted twice. The openparliament dump used here is dated 1 August 2026.
The statements begin in 1934, and the thirty-three years before that are worth stating as a result rather than a gap. Between 1901 and 1933 the House sat on more than two thousand days; 1930 to 1933 alone carry some 114,000 speeches, and the search returns nothing. The same holds inside the war years: 1942 has 121 sitting days and 32,824 speeches, the fullest year of the decade, and not one of them uses the word. That is a finding about parliamentary vocabulary, not a hole in the data.
Because the absence carries weight, the pattern was tested rather than assumed. It matches the hyphenated anti-Semitism, the spaced anti Semitism and the closed antisemitism alike, in either language, case-insensitively. Re-scanning 1930 to 1960 for any word containing the Semitic root, whatever its spelling, turns up four strings the corpus does not hold, and all four are semitropical.
The two sources agree closely. Over their twenty-five overlapping years they identify the same 320 sitting days, and differ by 7% on statement counts only because they divide long speeches differently: one intervention in the earlier archive may be two statements in the later one. Sitting-day counts are therefore the safer unit for any series crossing 1994, and the chart on the front page offers both. This is a footnote about resolution, not a warning about the data: the same debates, the same members and the same words appear in both.
The Canadiana collection holds the debates, journals and committee
documentation of both chambers from 1867 to 1996. Three things this project lacks
sit there: Senate debates before 1996, House debates before 1901, and House
committee evidence before 2006. Its operator publishes a robots.txt asking automated
agents not to retrieve from the site, so nothing here was scraped from it. Bulk
access for research is arranged directly with the operator, and a request is
pending. The coverage panel on the front page marks these years as outstanding.
A statement enters the corpus when its text matches one of four patterns, applied case-insensitively. The patterns are deliberately stem-based so that every inflection is caught.
antisemitism anti[\s\-]?s[ée]mit\w* jew_hatred jew[\s\-]?hatred | hatred of (the )?jews judeophobia jud[ée]ophob\w* antijuif anti[\s\-]?jui[fv]\w* | haine des juifs
The first pattern matches antisemitism, anti-Semitic, antisémitisme and the rest of the family. Matching is done on text, not on judgment: a statement appears because it contains one of these words, not because anyone has decided what the speaker meant.
A small number of statements enter because the debate was titled for
antisemitism, even though the member never used the word. These carry the tag
where:heading and can be isolated or excluded. They matter more than
their number suggests, and section 5 explains why.
Speakers' formulae such as "Motion agreed to" and "Agreed" are flagged and hidden by default, with a toggle to show them. In the Senate the collective interjection "Hon. Senators" is treated the same way, since it is a chorus rather than a speaker.
Both official languages are covered. Members speak in English or French, and Hansard publishes both versions of every intervention, so French speech is present throughout. Because an English rendering accompanies almost every statement, French speech is usually caught through it. French vocabulary is also searched directly, which adds statements the English search misses: 20 in the House and 24 in committees.
LiPaD publishes English text only, translating French interventions rather than dropping them, so Francophone members are represented in the historical period but not in their own words.
LiPaD and openparliament both cover the House from 1994 to 2019, which allows the two pipelines to be checked against each other rather than merely joined.
| Unit | LiPaD | openparliament | Disagreement |
|---|---|---|---|
| Matching statements | 635 | 685 | 7.3% |
| Sitting days with a match | 320 | 328 | 2.4% |
The two agree exactly from 1994 to 2006. Day-level agreement is 97.6%, and every day LiPaD identifies also appears in openparliament. Most of the statement level gap falls in 2015, where openparliament exposes debate headings that LiPaD's flat text cannot see.
Statements cluster heavily. The median sitting day with any mention has exactly one; the maximum has 91. A single take-note debate titled "Rise in anti-Semitism" on 24 February 2015 contributes 91 statements, roughly 94% of that year's House total.
A yearly count therefore reflects whether Parliament held a debate named for antisemitism at least as much as how often members used the word. Both measures are reported because they say different things:
| Period (House) | Statements | Sitting days | Per day |
|---|---|---|---|
| 1994–2005 | 124 | 86 | 1.4 |
| 2006–2014 | 240 | 145 | 1.7 |
| 2015–2020 | 343 | 112 | 3.1 |
| 2021–2026 | 526 | 206 | 2.6 |
Attention in 2015 to 2020 concentrated into fewer days at higher intensity. The most recent period spreads across more sitting days than any earlier one. Neither pattern is visible from a single measure.
For any statistical claim, the sitting day rather than the statement should be treated as the independent unit.
The 1901 to 1993 material was recognised from scanned pages, so a strict search could silently miss corrupted spellings, and would miss more of them the further back it looked. Rather than guessing at tolerant patterns, every token in the corpus whose shape could hide a corrupted semit core was harvested and read.
The harvest returned 47 candidate tokens. Genuine corruptions
(antisemiti, antisemitist, epresemitaibion,
senit) account for eight occurrences, against roughly
3,500 clean matches. The remainder are unrelated words the loose probe caught,
such as housewife, biscuits and the semi- prefixes.
Recognition error is therefore below 1% and does not explain the trend.
A stratified sample of records was fetched from the live source and compared against what this corpus holds: whether the link resolves, whether the term appears on that page, whether the extracted passage is present, and whether the speaker is correctly attributed.
| Sample | Clean |
|---|---|
| House records, one per year | 33 / 33 |
| Committee records | 19 / 19 |
| Official Hansard sitting links | 12 / 12 |
One House record initially failed and was found on inspection to be correct: it matched French text, while the source page renders English by default. The check, not the record, was at fault.
The corpus is defined by a word. A statement is included because it contains antisemitism or a close variant. That decision makes the corpus reproducible and keeps the sampling frame from pre-judging the contested question of what counts. It also creates one blind spot, and it is a large one.
Antisemitism is often expressed without the word Jew appearing at all. A speaker uses a substitute term that the audience understands while leaving the speaker room to deny the meaning: globalists, international bankers, cosmopolitan elites, cultural Marxism, rootless influence, the money power, or a named individual such as Soros standing in for a group. Scholars call this coded or dog-whistle language. Whether any particular use of these terms is antisemitic is contested, and often genuinely ambiguous, which is part of what makes the device effective.
Twenty-six statements in the corpus do use that register, but they are present only because they also happen to name antisemitism, usually for some other reason. They are a by-catch rather than a sample.
Two of them are recorded under a distinct value, coded, meaning the substitute vocabulary is present but is not applied to Jews anywhere in the statement. In both, the register appears but is not applied to Jews within the statement: in one, “globalists” sits in an anti-internationalist passage while the antisemitism reference is the member objecting to having been accused of a conspiracy theory. Whether register alone is antisemitic is genuinely contested in the literature, so these are held apart rather than folded into the count of antisemitic statements. Merging them would assert an answer; excluding them entirely would lose the observation.
Capturing coded speech would require a second corpus built on a different principle: search on the register itself across the whole of Hansard, then judge each hit. That is a much larger and much more contestable undertaking, since most uses of most of those terms are not antisemitic. It is not part of this project as it stands.
Not yet run. The scheme below is settled; the results are not in the tool.
Each statement will receive six tags. The first records what the speaker is doing with the term: condemns, accuses, disputes, enacts, or reports. Further tags record who is accused, whether the antisemitism described is located in Canada or abroad, what prompted the statement, and whether another form of hatred is named alongside.
A sixth tag records the grounds on which the claim is justified, following Boltanski and Thévenot's orders of worth, with the projective order from Boltanski and Chiapello and the green order from Lafaye and Thévenot. Up to two orders are recorded per statement, and statements combining two are flagged as compromises.
Coding is done by a language model. Agreement figures, the full prompt and the codebook will be published here so the labels can be assessed rather than trusted.
| Script | Does |
|---|---|
count_hits.py | Counts matches and denominators in the openparliament dump |
extract_instances.py | Builds House and committee records |
scan_lipad.py | Streams the LiPaD archive; counts, denominators, OCR harvest |
build_lipad_site_data.py | Builds pre-1994 House records |
scrape_senate.py | Scrapes sencanada.ca, one request per second |
normalise_senate_speakers.py | Collapses Senate speaker labels |
validate_extraction.py | Checks records against the live sources |
Every passage on this site originates in the official record of the Parliament of Canada: the Debates of the House of Commons (Hansard), the Evidence of its standing committees, and the Debates of the Senate (Hansard). The projects listed below made that record machine-readable, but the words are Parliament's.
Proceedings of the House of Commons are covered by the Speaker's Permission, which allows reproduction in whole or in part and in any medium provided the reproduction is accurate and is not presented as official. This site is offered on those terms.
Levi, R. (2026). Antisemitism in the Canadian Hansard. Lab for the Global Study of Antisemitism, Anne Tanenbaum Centre for Jewish Studies, University of Toronto. hansard-antisemitism.pages.dev. Derived from the Debates of the House of Commons and the Debates of the Senate of Canada.
This site is a project of the Lab for the Global Study of Antisemitism, housed in the Anne Tanenbaum Centre for Jewish Studies at the University of Toronto. The Lab supports research, teaching and events on the causes, manifestations and effects of antisemitism across the humanities, social sciences and professional fields.
Ron Levi. Distinguished Professor of Global Justice, Munk School of
Global Affairs and Public Policy and Department of Sociology, University of Toronto.
ron.levi@utoronto.ca
Correspondence about this corpus, corrections, and requests to reuse the data are welcome.