Over the past several months, Rasmus Rosenberg Larsen — an Assistant Professor of Forensic Epistemology and Philosophy of Science at the University of Toronto Mississauga (UTM), cross-appointed in UTM's Institute of Forensic Sciences and the University of Toronto's Department of Philosophy — has been pressing a sustained academic argument: that the psychological assessment most commonly used to label someone a "psychopath" in Canadian courtrooms rests on a shakier scientific foundation than the legal system generally treats it as having.
This is not a single press release or a newly dropped experiment. It is the latest stage of a multi-year research and public-communication effort. Larsen is the author of Psychopathy Unmasked: The Rise and Fall of a Dangerous Diagnosis (MIT Press, 2025), a book-length critical review of psychopathy research and its courtroom use. He discussed that book on CBC's Ideas with Nahlah Ayed in an episode that aired August 3, 2026, telling a national audience that the diagnosis of psychopathy should no longer be used in the criminal justice system. That interview followed a peer-reviewed study Larsen led — published in Psychology, Public Policy, and Law and widely reported in early February 2026 — which examined 3,315 Canadian court cases between 1980 and 2023 in which psychopathy assessments were introduced as evidence.
Stripped to its core, Larsen's published position rests on three connected claims, each traceable to specific research:
1. The instrument's reliability in real courtroom use is inconsistent. Larsen's 2026 study of Canadian case law found that scores on the Hare Psychopathy Checklist–Revised (PCL-R) — the tool used in the large majority of the cases reviewed — differed depending on whether the evaluating expert had been retained by the prosecution or by the defence. Robert Hare, the Canadian forensic psychologist who created the PCL-R, has himself acknowledged and cautioned about this pattern: in a 1998 paper, he noted that clinician scores can be influenced by which side retained the evaluator, and he stipulated that reliable use requires a highly trained clinician, high-quality case information, and ideally two independent raters — conditions Larsen's study argues are difficult to guarantee inside an adversarial courtroom process.
2. Courtroom testimony has, in a substantial number of cases, overstated what the science supports. The same study found that expert testimony frequently characterized psychopathy as effectively untreatable — a claim Larsen's earlier 2020 systematic review (co-authored with Jarkko Jalava and Stephanie Griffiths, published in Psychology, Public Policy, and Law) had already concluded was not well supported. That review examined decades of research behind three foundational claims associated with the PCL-R's originator, Robert Hare: that high scorers are especially dangerous, that they are more likely to reoffend after treatment or become worse, and that they are meaningfully impaired in moral understanding compared with other offenders. Larsen and colleagues reported finding no solid evidentiary basis for the "untreatable, and may get worse with treatment" claim specifically, and pointed instead to a growing body of research showing conventional treatment programs can have measurable positive effects even among high scorers.
3. The underlying instrument has not kept pace with the research literature that critiques it. Larsen has noted that the PCL-R manual, first published in 1980, has not been formally revised since 2003, while the empirical literature examining its predictive value — particularly in populations outside the original development samples — has continued to grow and, in his account, to raise questions.
It is worth being precise about what these are and are not. They are the arguments of one researcher and his collaborators, published in peer-reviewed venues and expressed in public broadcast commentary — not a court ruling, not a professional body's revised practice guideline, and not, on the evidence available, a finding that has displaced the PCL-R's continued use in Canadian or other courts.
| Confirmed argument | Scientific context | Forensic implication | |
|---|---|---|---|
| Bias | PCL-R scores in Canadian case law differed by which side retained the expert | Hare (1998) documented this risk and specified conditions to guard against it | Adversarial retainer structure can be difficult to fully insulate from |
| Untreatability claim | Courtroom testimony often described psychopathy as untreatable | 2020 systematic review found no strong evidence for that specific claim; some research shows treatment benefit | Overstated testimony could shape sentencing, parole, or treatment-eligibility decisions on a premise the literature does not clearly support |
| Instrument currency | PCL-R manual last revised in 2003 | Predictive-validity research has continued to accumulate since, with mixed findings across populations | Raises the question of how forensic practice should account for research published after a tool's last formal update |
The Hare Psychopathy Checklist–Revised (PCL-R) is a structured professional assessment instrument developed by Canadian psychologist Robert Hare, first published in 1980 and revised in 1991 and 2003. It is administered by a qualified, trained clinician through a semi-structured interview combined with a review of collateral file information — institutional records, criminal history, and similar documentation — rather than through a self-report questionnaire a person fills out alone.
The instrument scores an individual across 20 items related to interpersonal style, affective functioning, and behavioural history — traits such as glibness, lack of remorse, impulsivity, and a pattern of criminal versatility are represented among the item set, though the full item wording and scoring criteria are proprietary clinical material that this article does not reproduce. Each item is rated on a scale, and the item scores are summed into a total score. In North American forensic contexts, a score at or above a defined threshold has conventionally been used to classify someone as meeting criteria for psychopathy, with lower scores used in some contexts (including the UK) to indicate elevated but sub-threshold traits.
The PCL-R is widely described in the psychological literature as the field's most established and frequently used measure of psychopathic traits in forensic and correctional settings, and it has been the subject of an extensive research literature spanning several decades. That same body of research is not uniform: it includes studies supporting the tool's usefulness as a structured risk-assessment aid and studies — including Larsen's — raising concerns about its reliability and interpretation when applied outside carefully controlled research conditions.
Importantly, a PCL-R score is not itself a legal finding, a diagnosis in the sense of a formal psychiatric disorder listed in diagnostic manuals, or a prediction rendered with certainty. It is a structured clinical judgment about the degree to which an individual displays a defined set of traits, intended to be interpreted by trained professionals alongside other information — not treated as a standalone verdict on a person's character or future conduct.
Forensic psychological assessments, including psychopathy-related evidence, can potentially be introduced or considered at multiple points in the justice process: during sentencing submissions, in some bail or dangerous-offender hearings, in security classification decisions within correctional institutions, and in parole board deliberations about risk and readiness for release.
Whether and how such evidence is actually used varies substantially by jurisdiction, by the specific legal proceeding, by applicable admissibility rules, and by the facts of an individual case. Not every court accepts psychopathy-related testimony, and where it is accepted, the weight a judge or parole board assigns it is a matter of legal judgment, not a fixed formula. Larsen's own research focuses specifically on Canadian case law and should be read in that context; it does not establish how courts in other jurisdictions treat comparable evidence.
What his research does document, within the Canadian cases reviewed, is that psychopathy evidence has been introduced across a range of these high-stakes decision points over more than four decades — which is precisely why questions about its reliability and about how confidently its conclusions are communicated in court carry legal significance beyond the academic debate itself.
In forensic psychology, "reliability" refers to consistency — would two qualified evaluators, working from the same file and interview material, arrive at similar scores? Would the same evaluator score the same case similarly on a different occasion? Reliability is a precondition for validity: an instrument that produces inconsistent results cannot be a dependable predictor of anything, however sound its underlying theory.
Research on the PCL-R's reliability is mixed and depends heavily on context. Some studies — generally those conducted under research conditions with trained, blinded raters — report solid inter-rater agreement. Other "field reliability" research, examining how the instrument performs in actual clinical and forensic practice rather than controlled study settings, has reported considerably weaker agreement between raters. A study published in Law and Human Behavior, for instance, found poor rater agreement and concluded the instrument demonstrated inadequate field reliability and validity for recidivism prediction in the hospital and prison samples it examined, a gap the researchers attributed in part to differences between raters' training backgrounds and the complexity of real-world cases compared with research samples.
This distinction between reliability under controlled study conditions and reliability in day-to-day forensic practice is central to why researchers like Larsen argue that court use deserves separate scrutiny from laboratory validation. An instrument can be reasonably consistent in the hands of specialist researchers following a strict protocol and considerably less consistent when applied by a wider range of practitioners, under adversarial retainer pressures, with variable-quality file information — the conditions Larsen's Canadian case-law study specifically examined.
Predictive validity asks a different question from reliability: even if scored consistently, does a PCL-R score actually forecast future behaviour, particularly violent recidivism?
The published meta-analytic literature does show a statistical association between PCL-R scores and recidivism. Multiple meta-analyses spanning several decades of research have reported average correlations in the range of roughly r = .20 to .30 between PCL-R scores and general or violent recidivism — a modest-to-moderate effect size in statistical terms. A recent updated meta-analysis (2025) similarly reported moderate effect sizes for general and violent recidivism and for institutional misconduct and violence, while finding the "affective/interpersonal" component of the scale (Factor 1) consistently weaker as a predictor than the "antisocial/lifestyle" component (Factor 2).
It is critical to distinguish what this kind of statistical association can and cannot tell a courtroom about one individual. A correlation observed across a large group of people describes an average tendency across that group — it is not the same as a reliable prediction about a specific person's future conduct. Several strands of the literature underline this gap directly. Some analyses have reported false-positive rates for violence prediction exceeding 50 percent in certain samples — meaning that when using elevated scores to flag individuals as future violent recidivists, a majority of those flagged did not go on to commit the predicted act. Base rates matter enormously here: when the behaviour being predicted (serious violence) is relatively uncommon, even a moderately accurate instrument will misclassify a substantial number of individuals, in both directions.
Context and individual circumstances — release conditions, social supports, treatment engagement, age, and many other factors — also shape actual outcomes in ways a single score cannot capture. This is why professional guidance in the field, including recommendations published by researchers such as DeMatteo and Olver (2022), has called for reporting PCL-R results using percentile ranks and confidence ranges rather than a single number, specifically to better communicate measurement uncertainty to legal decision-makers who may otherwise read a score as a precise, individualized forecast.
In short: an association observed across groups is real and worth taking seriously in risk-assessment frameworks; treating that same association as a certain prediction about one person in front of a court is a separate, and considerably less defensible, inferential step.
Terminology carries weight beyond its technical meaning. The word "psychopath" arrives in a courtroom, or in a parole hearing, freighted with decades of cultural association — much of it shaped by fictional and sensationalized portrayals rather than by the clinical research literature. Larsen's broadcast comments and published work both raise the concern that this gap between public perception and research findings can influence how an assessment label is received, independent of the specific evidence underlying an individual score.
This is not a claim that judges, prosecutors, or parole board members are acting in bad faith or are automatically biased. It is a more specific concern about the potential for terminology to carry connotations — of permanence, untreatability, and heightened dangerousness — that Larsen's research argues are not consistently supported by the empirical literature. How an expert communicates a finding — whether framed with appropriate caveats about uncertainty and treatability, or framed in more categorical terms — may matter as much to a legal outcome as the underlying score itself. This is a communication and interpretation question as much as a scientific one, and it is a recurring theme across critiques of psychological expert evidence generally, not unique to psychopathy assessment.
A recurring theme across this literature — and arguably the strongest, most broadly applicable lesson from the debate — is the risk of overinterpretation: treating a probabilistic research finding as though it were a deterministic fact about an individual.
This can happen in several identifiable ways. A moderate statistical association between a score and an outcome can be presented, in testimony or in a report, with more confidence than the underlying research supports. A finding about typical outcomes across a study sample can be applied to a specific defendant as though it settled the question of that person's likely future conduct. A correlation — such as the association between certain traits and recidivism — can be discussed in language that implies a causal mechanism the research has not established. And uncertainty itself — confidence intervals, the possibility of measurement error, the existence of contradictory studies — can be omitted from a summary presented to a non-specialist decision-maker who has no independent way of knowing what was left out.
None of this requires bad faith on the part of any individual evaluator. It can result from the ordinary pressures of an adversarial legal process, from time constraints on report writing, or simply from the difficulty of conveying statistical nuance in a courtroom setting built around definitive answers. But the effect, researchers in this field argue, is the same regardless of intent: a legal decision-maker may end up relying on a degree of certainty the science does not actually offer.
The corrective, as reflected across this literature, is not to abandon structured assessment in favour of unstructured judgment — research generally suggests structured tools outperform pure clinical intuition — but for forensic experts to communicate the limitations, confidence levels, and scope of their conclusions as explicitly as the substantive findings themselves.
Scientific balance requires acknowledging that psychopathy research is an active, substantial field, and that the PCL-R remains, by most accounts in the literature, the most extensively validated and widely used structured instrument of its kind in forensic and correctional settings.
The meta-analytic literature described above does show a real, replicated statistical association between PCL-R scores and recidivism outcomes — an association some researchers describe as a meaningful improvement over unstructured clinical judgment alone, which historically has performed poorly at predicting violence. Ongoing research continues to refine understanding of which components of the instrument (interpersonal/affective traits versus antisocial/lifestyle traits) carry more predictive weight for which outcomes, and professional training standards emphasize that reliable use depends on qualified administration, high-quality source information, and — per Hare's own guidance — ideally more than one independent rater.
The debate, properly understood, is not "psychopathy assessment works perfectly" versus "psychopathy assessment is worthless." It is a narrower and more consequential question: given a real but moderate, context-dependent statistical association, how should its conclusions be validated, communicated, and weighed when the outcome is a person's liberty, sentence, or correctional pathway? Larsen's position is that current courtroom practice frequently gets this balance wrong; the continued professional use of structured psychopathy assessment across forensic psychology reflects that many practitioners and researchers believe, with appropriate safeguards, it still offers value that unstructured judgment does not.
The broader legal questions this debate raises extend beyond psychopathy assessment specifically to the treatment of expert psychological evidence generally: standards for scientific admissibility, the transparency with which methodological limitations are disclosed to the court, the effectiveness of cross-examination in surfacing those limitations for a lay decision-maker, and the degree to which judges and juries can be expected to independently evaluate technical claims about reliability and predictive validity.
Larsen's research is specific to Canadian case law and does not purport to establish uniform findings across other legal systems; jurisdictions vary in their rules of evidence, their standards for admitting expert testimony, and the procedural context in which psychological assessments are introduced. Readers in other jurisdictions should treat the Canadian findings as a case study illustrating a set of concerns that may or may not generalize to their own legal system, rather than as a direct statement about local practice.
The most important forensic question raised by this debate may not simply be whether the PCL-R is a scientifically useful instrument in the abstract. It may instead be whether its findings, in specific cases, are being communicated and interpreted in a manner proportionate to the strength and limitations of the underlying evidence.
A structured psychological instrument with a moderate, well-documented statistical association to an outcome can still be misused — not through fraud, but through the ordinary human tendency to compress nuance into certainty when a definitive answer is what a legal process demands. Larsen's Canadian case-law findings, if their pattern holds up to further scrutiny and replication, point to exactly this kind of gap: a tool developed and validated under research conditions being applied, described, and weighed in an adversarial courtroom context that was not built to accommodate the statistical humility the underlying science actually calls for.
This suggests the most productive response to this debate is unlikely to be a binary choice between banning psychopathy assessment outright or leaving current practice unexamined. It is more likely to lie in the direction several researchers in this field — including those who continue to defend the PCL-R's underlying validity — already point toward: clearer professional standards for how conclusions are worded in reports and testimony, more consistent disclosure of measurement uncertainty (through percentile ranges rather than single scores, for example), safeguards against the specific bias risk Hare himself identified around retainer source, and ongoing judicial and legal-professional education about what a moderate statistical association can and cannot support. Scientific humility, transparency about limitations, and proportionality between the strength of evidence and the weight given to it in a legal decision are not concessions that weaken forensic psychology's contribution to criminal justice — they are what allow that contribution to be trusted.