Despite decades of refinement to diagnostic manuals, questions persist about the consistency and scientific basis of psychiatric diagnosis. A new article published in Frontiers in Psychiatry finds that psychiatric diagnoses are often inconsistent and unreliable, with clinicians commonly disagreeing on what diagnosis a set of symptoms should receive. This study, led my Mateo Boberg from the University Hospital of Copenhagen in Denmark, also reports that agreement between clinicians was especially low for schizophrenia diagnosis while being higher for OCD, bipolar disorder, and major depressive disorder. The authors write:
“We found a poor overall reliability of psychiatric diagnoses … corresponding to a situation in which two clinicians, assessing the same written clinical case, agree on the diagnosis approximately half of the time … From a clinical perspective, the level of diagnostic agreement is not merely a statistical concern but has direct implications for patient care. If two clinicians are likely to assign different diagnoses to the same patient, this raises questions about the stability of treatment decisions, prognostic expectations, and communication across clinical settings.”

Validity and Reliability Problems of Psychiatric Diagnosis
Psychiatric diagnosis has faced criticism in the past for lacking validity. This means experts have questioned whether the diagnostic categories actually line up with reality, which includes well documented problems with diagnostic tests producing false positives.
Allen Frances, the chair of the DSM-IV taskforce, has criticized broad diagnostic criteria and the DSM-5, saying “It is simply impossible, given available knowledge, to create criteria sets that will be specific enough to avoid also identifying a large pool of false positives.” Thomas Insel, former director of the United States Institute of Mental Health, has also questioned the validity of psychiatric diagnosis, remarking that the strength of DSM-style diagnostic criteria is not validity but reliability: “The strength of each of the editions of DSM has been ‘reliability’—each edition has ensured that clinicians used the same terms the same ways … The weakness is its lack of validity.”
However, DSM-style diagnosis also has reliability issues, with the same tests sometimes producing different results. Similar to the current study, past research has shown issues with inter-rater reliability, meaning different evaluators sometimes arrive at different diagnoses using data from a single interview, evaluation, or vignette. In order for a diagnostic system to be sound, it must demonstrate both validity and reliability. DSM-style diagnosis has demonstrated issues in both areas.
Psychiatric survivors have detailed the harms of diagnosis, including stigma, discrimination, warped self-concept, worse quality of life, and increased risk of preventable medical errors. Peter Gøtzsche has also written about how diagnosis can become a self-fulfilling prophecy, obscure the real causes of suffering, and create power imbalances.
Study Details
The goal of this study was to evaluate the inter-rater reliability of psychiatric diagnosis using diagnostic criteria from the International Classification of Diseases 10 (ICD-10). In other words, the authors wanted to measure how often doctors agreed on a psychiatric diagnosis given a set of symptoms.
The authors recruited participants from 19 countries in Europe and South America to examine clinical cases and assign a psychiatric diagnosis based on ICD-10 criteria. Participants had to be medical doctors working in adult psychiatry. This included psychiatrists, residents, and other medical doctors. Participants were recruited through local collaborators’ professional networks. Each participant responded to a survey that asked about demographic information and presented one or two (out of nine total) clinical cases for diagnosis. The participants chose a diagnosis for each clinical case from a list of 30 common psychiatric diagnoses. In total, the authors analyzed 1,902 responses from 1,038 participants.
The nine included clinical cases were assigned a best-fit diagnosis by a team of senior psychiatrists. As determined by the expert team, these clinical vignettes included three cases of schizophrenia, and one each of schizotypal disorder, bipolar disorder, depressive disorders, anxiety disorders, OCD, and personality disorders.
The majority of participants were male (55.5%), psychiatrists (61.6%), and based in Europe (79.2%). Thirty-two percent of participants had between 0 – 4 years of experience and 30% had fifteen or more years of experience.
The authors describe the overall reliability of psychiatric diagnosis in the current study as “modest.” They used a metric called Krippendorff’s alpha to measure overall reliability. This metric reflects how much agreement there is between raters after accounting for the agreement that could happen by chance. A Krippendorff’s alpha coefficient between 0.80 – 1.00 indicates strong agreement and a reliable rating system. Anything less than 0.67 in considered unreliable while 0.00 would indicate random chance. The current study reported a Krippendorff’s alpha coefficient of 0.48, indicating poor agreement and unreliable diagnostic criteria. This finding means that clinicians agreed with each other on a diagnosis a little over half of the time after accounting for chance.
The authors note that the doctors’ level of agreement was similar regardless of whether they were experienced psychiatrists, psychiatry residents, or other physicians working in psychiatry. It also did not vary based on how many years of clinical experience they had. The authors write:
“Diagnostic agreement did not differ significantly among the doctors, neither when stratified into position (psychiatrist, residents, and other medical doctor) nor when stratified into years of clinical experience. Thus, the observed diagnostic variability cannot be explained simply by varying levels of clinical experience.”
Overall, 66% of participants selected the diagnosis assigned to the clinical cases by the expert team of senior psychiatrists. The bipolar disorder (94% of participants agreed with the expert team), OCD (92%), and depressive disorders (88%) cases saw the highest rates of agreement between participants and the expert team, followed by anxiety disorders (73%), and personality disorders (72%). The three schizophrenia cases (37% – 54%) and the schizotypal disorder case (38%) showed substantial disagreement between the participants and the expert psychiatrists.
While the scores for OCD, depression, and bipolar disorder seem to indicate strong agreement, it is important to note that false positives were not counted as disagreement in this measure. OCD and bipolar disorder in particular were mistakenly chosen by many participants as a diagnosis for the schizophrenia clinical case vignettes, indicating that agreement around these diagnoses is not as strong as the above numbers would make it seem.
This pattern suggests that clinicians were not simply making careless errors, but were interpreting the same symptoms through different lenses. In other words, the symptoms presented in these cases may overlap across multiple psychiatric categories, making it difficult to draw clear boundaries between diagnoses.
These findings raise questions about whether psychiatric diagnoses represent clearly defined disorders or whether they impose boundaries on experiences that are often complex, overlapping, and difficult to separate.
The authors acknowledge several limitations to this research. The clinical cases were written vignettes rather than real patient interviews. This means participants could not ask follow-up questions or assess non-verbal behaviors. The reference diagnoses determined by the expert team was merely a consensus and based on the clinical judgment of the experts. Participants were given a list of 30 diagnoses to choose from. This does not reflect the full range of diagnoses (over 200 in the ICD-10) or everyday diagnostic practices. Participants were recruited through professional networks and took the survey anonymously. This means the response rate and representativeness of the participant sample was unknown. The authors conclude:
“Despite decades of methodological refinement and the introduction of diagnostic criteria and rule-based classification, the inter-rater reliability of psychiatric diagnoses remains modest, even under the standardized conditions of written clinical cases. The present findings could suggest that refinement of diagnostic criteria alone may be insufficient for achieving high levels of diagnostic agreement.”