multi-item-scale · dataset v0.9.0 contested: Structural validity, Convergent and discriminant validity

Patient Health Questionnaire-9 (PHQ-9)

Licence verified: 2026-07-12 · literature last reviewed: 2026-09-01 · grades last confirmed: 2026-09-03 · rubric v1.6

In plain English

Patient Health Questionnaire-9 (PHQ-9) is a multi-item scale (9 items). What it claims to measure: Severity of depressive symptoms over the preceding two weeks and, secondarily, provisional detection of major depressive disorder. Licence for an employer or vendor: free, no permission needed (verified 2026-07-12).

The published evidence is strongest for structural validity and internal consistency; moderate for convergent and discriminant validity, criterion validity against a reference standard, test-retest reliability and responsiveness to change; weak for measurement invariance. No published evidence was located for criterion validity against organisational outcomes (absence, turnover, performance): the registry searched and found none in the sweep to date, so if you need that property evidenced, this instrument does not yet carry it. The evidence base for structural validity and convergent and discriminant validity is contested in the published literature. 2 of the 7 graded properties rest on evidence from clinical, student or otherwise non-working samples; where the property is sensitive to population that evidence cannot carry a High grade, and the reason is stated on each cell below. 5 of the 7 graded properties rest on adult general-population samples rather than samples of working adults; the flag says so on each cell.

Structural validity: HighConvergent validity: ModerateCriterion (reference standard): ModerateCriterion (organisational): AbsentInternal consistency: HighTest-retest: ModerateInvariance: LowResponsiveness: ModerateLicence: free-no-permission

Grades summarise the published evidence for this instrument on its own terms; they are not comparable across instruments, and this page makes no recommendation. Whether an instrument fits your workforce is a judgement this registry informs but cannot make. Full evidence, with citations, below. This summary is generated from the record's data, not written by hand.

Who graded this. Every grade in this registry was assigned by one rater employed by the steward (1 rater, 0 independent), with AI assistance in literature retrieval and drafting. Grades and statuses are single-rater and frozen from first publication; they will not move until two named psychometric raters who are not employees of the steward have joined. Until then, automated sweeps add citations and flag cells for review; corrections of fact are made in public; no grade changes. Why, and how to volunteer.

Identity

Version: PHQ-9 (2001); abbreviated variants PHQ-8 (omits item 9) and PHQ-2 (first two items) are in wide use

Structure: 9 items, each scored 0 to 3 (total 0 to 27); severity bands 5, 10, 15, 20 map to mild, moderate, moderately severe and severe

Original citation: Kroenke K, Spitzer RL, Williams JB (2001) The PHQ-9: validity of a brief depression severity measure. Journal of General Internal Medicine. 10.1046/j.1525-1497.2001.016009606.x

Steward / publisher: Developed by Spitzer, Williams and Kroenke as the depression module of the PRIME-MD PHQ, under an educational grant from Pfizer Inc; the PHQ suite was subsequently released by Pfizer for free public access.

free-no-permissionfree, no permission needed for employer or vendor use. Steward terms exempt the screener from general copyright restrictions; free to download and use as stated on the steward site.
Licence status (verified 2026-07-12): CONFIRMED and now verified, with a wording refinement. The official steward terms state: 'Content found at the PHQ Screeners site is expressly exempted from Pfizer's general copyright restrictions; content found on the PHQ Screeners site is free for download and use as stated within the PHQ Screeners site'. This confirms free download and use with no permission or fee, but is a copyright exemption / free-use grant rather than a formal public-domain dedication.
Source: https://www.phqscreeners.com/terms (official PHQ Screeners terms, Pfizer)
Steward page as read, Internet Archive: archived copy 1 · archived copy 2
The registry records the licence class and the steward's page, never a price: fees change without notice and the archived page is the record of what was read.

Constructs claimed

Severity of depressive symptoms over the preceding two weeks and, secondarily, provisional detection of major depressive disorder. The nine items map directly onto the nine DSM-IV/DSM-5 symptom criteria for a major depressive episode, so the instrument claims both a continuous severity construct and a criterion-referenced screening function (Kroenke 2001). It is a depression-specific measure, not a general wellbeing or distress index, and was designed for primary-care and clinical use rather than for occupational or workplace-wellbeing measurement.

Evidence

Deployment context caveat. PHQ-9 is a depression screener whose criterion validity was earned in primary-care and specialist diagnostic settings; workplace deployment is a different context, and the word clinical is not applied affirmatively to any workplace use. (applies to every property below)

Structural validity Highcontestedevidence form: canonical

general nationally representative data; also university students, traumatic brain injury, stroke and coronary heart disease samples (flag basis: Korean nationally representative data; Brazilian university students; traumatic brain injury patients; stroke sample; Italian coronary heart disease sample)

state: assessed · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03 · under rubric v1.3: precondition evidence 5 entries → 3 entries (re-read before first publication, correction C-0007) · under rubric v1.1: indirectness direct → general (re-read before first publication, correction C-0005)

The dimensionality of the PHQ-9 is genuinely contested, but the practical reading across studies is that it behaves as an essentially unidimensional depression severity scale even where a two-factor model fits marginally better. The original development treated the total score as a single severity dimension (Kroenke 2001). Within the individual participant data (IPD) programme, a comparison of unidimensional, two-dimensional (cognitive/affective versus somatic) and bifactor latent models found that all fitted reasonably and that scoring by the more complex latent models improved sensitivity only marginally (by about 0.04 to 0.05) while reducing specificity, relative to the sum score at a cut-off of 10 or above (Fischer et al 2021). Several population studies reach the same conclusion from different angles: in Korean nationally representative data a single-factor model fitted well (CFI 0.944) with satisfactory internal consistency (Lee et al 2023); in Brazilian university students both one- and two-dimensional models fitted, with the authors favouring the unidimensional solution (Rufino et al 2024); and in traumatic brain injury patients a general factor explained about 85% of the variance despite a slight bifactor improvement (Teymoori et al 2020). Where a two-factor (somatic versus cognitive-affective) structure is reported, the factors tend to be very highly correlated: in a stroke sample the two-factor model fitted slightly better than the one-factor model (CFI 0.984 versus 0.974) but the inter-factor correlation of 0.866 pointed back to unidimensionality (Blake et al 2024), and an Italian coronary heart disease sample reported a bi-dimensional somatic/cognitive solution (Di Matteo et al 2025). The disagreement is therefore about whether the somatic items form a distinguishable subfactor, not about whether a single severity score is defensible; the weight of evidence supports scoring and interpreting a single total. Sample sizes were 6000 primary care and obstetrics-gynaecology patients in the development study (Kroenke 2001), 24 calibration studies with 4378 participants and 17 validation studies with 4252 participants in the IPD comparison, where sensitivity gains were 0.036, 0.050 and 0.049 and specificity losses 0.017, 0.026 and 0.019 for the unidimensional, two-dimensional and bifactor models (Fischer et al 2021), 5347 Korean adults, with SRMR 0.040, RMSEA 0.076 and omega 0.812 for the one-factor model (Lee et al 2023), 3163 Brazilian university students (Rufino et al 2024), 2137 patients at six months after traumatic brain injury (Teymoori et al 2020), 787 stroke and 12,016 non-stroke participants with a propensity-matched comparison subsample of 1574 (Blake et al 2024) and 427 Italian coronary heart disease patients, with model-based reliability 0.80 (Di Matteo et al 2025).

High precondition (rubric 1.6): the cited studies that meet the High precondition, each with a sample size and a statistic of this property.

StudySampleStatisticDOI
Blake et al 2024787 stroke and 12,016 non-stroke; propensity-matched comparison subsample 1574two-factor CFI 0.984 versus one-factor CFI 0.974; inter-factor r 0.866; configural invariance CFI 0.983, RMSEA 0.080; strong invariance violated (delta CFI -0.003)10.1016/j.jpsychores.2024.111983
Lee et al 20235347one-factor: chi-square 770.765, CFI 0.944, SRMR 0.040, RMSEA 0.076 (90% CI 0.072 to 0.081); omega 0.81210.3389/fpsyg.2023.1217038
Teymoori et al 20202137 at six months (1922 longitudinal)bifactor slight fit improvement; general factor explained 85.0% of variance (omega hierarchical)10.3390/jcm9030873

Confidence note (legacy, first pass; not the basis of the grade): High: many good-quality factor-analytic studies across large and varied samples, converging on essentially unidimensional use despite a recurring, well-characterised somatic/cognitive two-factor debate.

Convergent and discriminant validity Moderatecontestedevidence form: canonical

general general population; also students, traumatic brain injury and coronary heart disease samples (flag basis: Chinese general population; students; traumatic brain injury patients; coronary heart disease patients)

state: assessed · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03 · under rubric v1.1: indirectness direct → general (re-read before first publication, correction C-0005)

Convergent evidence is consistent in direction but uneven in magnitude, and discriminant separation from anxiety is imperfect. In development, higher PHQ-9 scores tracked substantial decrements across all six SF-20 functional-status subscales and greater symptom-related difficulty, supporting construct validity (Kroenke 2001). In the Chinese general population the PHQ-9 correlated negatively with SF-36 subscales (r from -0.11 to -0.47) as expected, but its correlation with the Zung Self-Rating Depression Scale was unexpectedly weak (r = 0.29), an inconsistency worth noting against the usual assumption of strong convergence with other depression measures (Wang et al 2014). Convergent relationships with sleep quality, alcohol use and physical activity have also been reported in students (Rufino et al 2024). On discriminant validity, the PHQ-9 and the GAD-7 anxiety scale share a large common factor: in traumatic brain injury patients a general distress factor dominated both scales (about 85% of variance) even though the instruments related differently to SF-36 subscales (Teymoori et al 2020), and in coronary heart disease patients PHQ-9 and GAD-7 scores were significantly positively correlated (Di Matteo et al 2025). Depression and anxiety symptoms measured this way are statistically distinguishable but strongly overlapping.

Confidence note (legacy, first pass; not the basis of the grade): Moderate: multiple studies establish expected convergent and functional-status correlations, but magnitudes vary (including one weak convergent correlation) and discriminant separation from anxiety is only partial.

Criterion validity: reference standard Moderatewell-establishedevidence form: canonical

indirect clinic patients and diagnostic meta-analysis samples; no working-adult or general-population sample (flag basis: primary-care and obstetric-gynaecology clinic patients; major depression cases)

state: assessed · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03 · under rubric v1.2: grade High → Moderate; precondition evidence 5 entries → none (re-read before first publication, correction C-0006)

Criterion validity against a diagnostic interview for major depression is the PHQ-9's best-evidenced property, and it is strong. In the original validation in primary-care and obstetric-gynaecology clinic patients a cut-off of 10 or above gave 88% sensitivity and 88% specificity against an independent mental-health-professional interview (Kroenke 2001). The individual participant data meta-analysis programme is the anchor here. The 2019 IPD meta-analysis (58 studies, n = 17,357, 2,312 major depression cases) found combined sensitivity and specificity maximised at a cut-off of 10 or above against semistructured interviews (sensitivity 0.88, 95% CI 0.83 to 0.92; specificity 0.85, 0.82 to 0.88) (Levis et al 2019). The 2021 update (100 studies, n = 44,503) reproduced this at the same cut-off (sensitivity 0.85, 0.79 to 0.89; specificity 0.85, 0.82 to 0.87) (Negeri et al 2021). A crucial nuance is reference-standard dependence: sensitivity against semistructured clinician interviews ran higher than against fully structured lay interviews (5 to 22 percentage points) or the MINI (2 to 15 points) in the 2019 synthesis (Levis et al 2019), with median differences of around 21% and 11% in the 2021 update, while specificity was similar across standards (Negeri et al 2021). The diagnostic algorithm approach performed worse than the cut-off (algorithm sensitivity around 0.57 to 0.61 versus 0.88 for the cut-off against semistructured interviews) (He et al 2019), and the PHQ-2 with cut-off 3 or above (sensitivity 0.72, specificity 0.85) is a reasonable first-stage screen followed by the PHQ-9 (Levis et al 2020). The original criterion sample was 580 patients drawn from the 6,000 who completed the PHQ-9 across 8 primary care and 7 obstetrics-gynaecology clinics (Kroenke 2001); the cut-off 10 estimate in the 2019 synthesis rests on 29 semistructured-interview studies with 6,725 participants (Levis et al 2019); the algorithm meta-analysis pooled 54 studies (n=16,688, 2,091 cases) and found algorithm specificity of 0.95 against 0.86 for the cut-off (He et al 2019); and the PHQ-2 analysis drew on 100 studies (44,318 participants, 4,572 with major depression), with an area under the curve of 0.88 for semistructured interviews and a PHQ-2 (2 or above) then PHQ-9 (10 or above) sequence giving sensitivity 0.82 and specificity 0.87 (Levis et al 2020).

Confidence note (legacy, first pass; not the basis of the grade): High for diagnostic criterion validity against clinical interviews (large, consistent IPD evidence); Absent for criterion validity against organisational/workplace outcomes (no such studies retrieved). The overall grade is split by criterion.

Criterion validity: organisational Absent (searched; none found in the sweep to date)untestedevidence form: canonical

absence type: population-general searched; none found in the sweep to date

state: assessed_absent · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03

No study retrieved in this pass tested PHQ-9 scores against measured sickness absence, staff turnover, productivity loss or other workplace outcomes as criteria. The closest health-related association is in the development study, where self-reported sick days and health-care utilisation rose with PHQ-9 severity (Kroenke 2001). Criterion validity against organisational endpoints is therefore unestablished.

Internal consistency Highwell-establishedevidence form: canonical

general general population and a nationally representative sample; the pooled studies name no population (flag basis: Chinese general population; Korean nationally representative sample)

state: assessed · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03 · under rubric v1.1: indirectness direct → general (re-read before first publication, correction C-0005)

Internal consistency is consistently high across populations. A reliability generalisation meta-analysis of 60 studies (232,147 participants) estimated a pooled Cronbach's alpha of 0.86 (95% CI 0.85 to 0.87), with self-administered formats slightly higher (alpha 0.87) than face-to-face interview administration (alpha 0.80), though between-study heterogeneity was very large (I-squared 99.3%) (Ajele 2025). Single-study estimates sit in the same range: alpha 0.86 in the Chinese general population (Wang et al 2014) and omega 0.812 in a Korean nationally representative sample (Lee et al 2023). A systematic review of the wider PHQ family likewise reported good internal consistency for the PHQ-9 (Kroenke 2010), and a meta-analysis of the Spanish-language versions evaluated internal consistency alongside accuracy (Martinez et al 2023). Reported coefficients of omega tend to align with alpha where both are given, but omega is reported far less often than alpha. Single-study sample sizes were n = 1045 Shanghai community residents (Wang et al 2014) and n = 5347 Korean adults over 19 (Lee et al 2023); the Spanish-language meta-analysis pooled 10 studies with 5164 adults and reported PHQ-9 Cronbach's alpha of 0.78 to 0.90 and McDonald's coefficient of 0.79 to 0.90 (Martinez et al 2023), and the PHQ family review drew on four multisite studies of 9740 patients but its abstract gives no internal consistency coefficient (Kroenke 2010).

High precondition (rubric 1.6): the cited studies that meet the High precondition, each with a sample size and a statistic of this property.

StudySampleStatisticDOI
Martinez et al 202310 studies, 5164 Spanish-speaking adults (mean age range 34.1 to 71.8); 8 in primary carePHQ-9 Cronbach alpha 0.78 to 0.90; McDonald's coefficient (printed as ψ) 0.79 to 0.9010.1001/jamanetworkopen.2023.36529
Ajele 202560 studies, 232,147 participantsPooled alpha 0.86 (95% CI 0.85 to 0.87); self-administered alpha 0.87; face-to-face alpha 0.80; I-squared 99.3%; test-retest 0.82 (8 studies)10.1007/s44192-025-00181-x
Wang et al 20141045 (100 retested at 2 weeks)Cronbach's alpha 0.86; test-retest 0.8610.1016/j.genhosppsych.2014.05.021
Lee et al 20235,347Omega 0.81210.3389/fpsyg.2023.1217038

Confidence note (legacy, first pass; not the basis of the grade): High: a large reliability-generalisation meta-analysis plus multiple primary studies converge on alpha around 0.86, albeit with substantial heterogeneity and sparse omega reporting.

Test-retest reliability Moderatethinevidence form: canonical

general general population; also a late-life depression treatment sample and primary care patients (flag basis: Chinese general population; late-life depression treatment sample; Spanish primary care patients)

state: assessed · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03 · under rubric v1.1: indirectness direct → general (re-read before first publication, correction C-0005)

CoefficientTypeIntervalSamplePopulationEvidence form
0.82 (95% CI 0.74 to 0.90)pooled test-retest estimatenot stated (pooled across 8 studies)8 studiesmixed, as pooled in the reliability generalisation meta-analysiscanonical
0.86rtwo weeksnot stated in the summaryChinese general populationcanonical
described as excellent; no coefficient in the abstractnot statedseven daysnot stated in the summarylate-life depression treatment samplecanonical
intraclass agreement, value not stated in the summaryICCself- versus telephone-administrationnot stated in the summarySpanish primary care patientscanonical

Test-retest reliability exists but is markedly under-studied relative to the vast diagnostic-accuracy literature, and this asymmetry is itself the finding. The reliability generalisation meta-analysis could pool a test-retest estimate from only 8 of its 60 studies, yielding 0.82 (95% CI 0.74 to 0.90) (Ajele 2025). Primary estimates are favourable where they exist: a two-week retest correlation of 0.86 in the Chinese general population (Wang et al 2014), test-retest described as excellent over a seven-day interval in a late-life depression treatment sample (Lowe et al 2004), and intraclass agreement assessed for self- versus telephone-administration in Spanish primary care (Pinto-Meza et al 2005). No test-retest study in an occupational or workplace sample was located in this pass. The property is present and reassuring in the settings studied, but the evidence base is thin and skewed toward clinical and general-population samples over short intervals.

Confidence note (legacy, first pass; not the basis of the grade): Moderate: several favourable estimates (roughly 0.82 to 0.86) exist and a meta-analytic pooled value is available, but from few studies (n = 8 pooled), short intervals, and no workplace data. Not absent, but comparatively neglected.

Measurement invariance Lowthinevidence form: canonical

general nationally representative sample; also coronary heart disease and implantable-defibrillator cohorts (flag basis: Korean nationally representative sample; Italian coronary heart disease cohort; Danish implantable-defibrillator cohort; clinical or general-population samples)

state: assessed · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03 · under rubric v1.1: indirectness direct → general (re-read before first publication, correction C-0005)

Invariance has been tested piecemeal and mostly reaches metric to scalar level within the groups examined, but coverage of occupational populations is absent. In a Korean nationally representative sample the one-factor PHQ-9 showed equivalent structure, factor loadings and item intercepts across age groups, i.e. up to scalar invariance across ages (Lee et al 2023). Invariance across sex and across age (65 and over versus under 65) was tested in an Italian coronary heart disease cohort (Di Matteo et al 2025). Item response theory analysis in a large Danish implantable-defibrillator cohort found no differential item functioning across educational level, age, clinical indication or heart-failure severity, with only a single item showing DIF by gender (Pedersen et al 2016). This points to broadly stable measurement across sex and age in clinical and general populations. However, the studies are confined to clinical or general-population samples; invariance across occupations, across employed versus unemployed status, and over repeated workplace administrations was not established in retrieved evidence. Longitudinal (over-time) invariance is likewise weakly evidenced for the PHQ-9 specifically.

Confidence note (legacy, first pass; not the basis of the grade): Low to Moderate: scalar invariance is demonstrated across age (one strong study) and DIF is minimal in another, but evidence is scattered across clinical populations, occupational invariance is untested, and over-time invariance is weak.

Responsiveness and MIC Moderatethinevidence form: canonical

indirect late-life depression trial participants; no working-adult or general-population sample (flag basis: IMPACT late-life depression trial)

state: assessed · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03

Responsiveness to change is well supported, and a minimal important change has been proposed, though the MIC rests largely on one study. In the IMPACT late-life depression trial (n = 434 intervention participants) the PHQ-9 was responsive to treatment, with an effect size (about -1.3 at three months) exceeding the SCL-20 depression scale at three months and matching it at six months; change scores discriminated persistent depression, partial remission and full remission against structured diagnostic interviews (Lowe et al 2004). That study estimated a minimal clinically important difference for individual change of about 5 points on the 0 to 27 scale, derived as two standard errors of measurement (Lowe et al 2004). A systematic review of the PHQ family concluded that sensitivity to change is well established for the PHQ-9 (Kroenke 2010). Routine-outcome use in stepped-care services also relies implicitly on responsiveness, with large pre-post effect sizes reported for depression (Richards 2009). The MIC of 5 points is widely cited but should be understood as a single-derivation estimate rather than a triangulated consensus value.

Confidence note (legacy, first pass; not the basis of the grade): Moderate: responsiveness is demonstrated in a good treatment study and endorsed by a review, but the 5-point MIC derives essentially from one study and one method (2 SEM), and no MIC has been established in a workplace context.

Populations, languages and norms

The PHQ-9 has been validated across a wide range of clinical, general-population and disease-specific samples and in many languages, but formal population norms in the classical sense are scarce; the UK benchmark is embedded in a national service dataset rather than a norm table. The IPD meta-analyses aggregate around 100 primary studies from many countries (Negeri et al 2021, Levis et al 2020). Validated non-English versions retrieved in the latest review pass include Chinese (Wang et al 2014) and Spanish, the latter via a dedicated systematic review and meta-analysis of the Spanish-language PHQ-2 and PHQ-9 (Martinez et al 2023), with further evaluations in Korean (Lee et al 2023), Danish (Pedersen et al 2016) and Italian (Di Matteo et al 2025) samples. For the UK specifically, the PHQ-9 is the routine depression outcome measure of the English NHS Talking Therapies programme (formerly Improving Access to Psychological Therapies, IAPT), where a score of 10 or above defines 'caseness' and movement below it underpins recovery reporting; large UK service cohorts and a UK randomised trial have used it in exactly this way (Richards 2009, Barkham et al 2021). UK reference values therefore live in national service reporting rather than in a published normative sample. General-population norms exist chiefly through large national surveys used in the invariance literature rather than as a dedicated UK norm set.

Criticisms and controversies

Several substantive controversies recur in the literature. First, cut-point and accuracy estimation. Studies that selectively report only well-performing cut-offs bias meta-analytic accuracy: for the PHQ-9, published results underestimated sensitivity below a cut-off of 10 (median difference about -0.06) and overestimated it above 10 (median about +0.07) (Neupane et al 2021). Relatedly, using small datasets to simultaneously choose an optimal cut-off and estimate its accuracy is biased; in resampling from the IPD database the population-level optimal cut-off was actually 8 or above, and small studies scattered widely around it (only about 17% of 100-participant studies recovered the true optimum) (Levis et al 2024). This questions the near-universal reliance on a cut-off of exactly 10. Second, reference-standard dependence: reported sensitivity is substantially higher against semistructured clinician interviews than against fully structured lay interviews or the MINI, so headline accuracy figures are conditional on the comparator (Levis et al 2019, Negeri et al 2021, He et al 2019). Third, somatic-item confounding in physically ill populations. In systemic sclerosis, somatic items accounted for a larger share of the total score than in matched healthy respondents, inflating scores by roughly 1.0 to 1.4 points (Hedges g 0.38 to 0.55) (Leavens et al 2012); in stroke, summed scoring moderately overestimated depression relative to a comparison sample (Cohen d about 0.43), driven by items such as tiredness and appetite (Blake et al 2024). This matters wherever respondents have physical illness or fatigue. Fourth, item 9 (thoughts of death or self-harm) is often misread as a suicidality measure, yet most positive responses are not associated with suicidality; the PHQ-8, which omits item 9, correlates almost perfectly with the PHQ-9 (r = 0.996) and performs almost identically for detecting depression (Wu et al 2019). Fifth, complexity does not pay: latent-variable and machine-learning scoring add negligible accuracy over the simple sum score with a cut-off (Fischer et al 2021, Hong et al 2022). Finally, and central to the OWHS use case, the PHQ-9 is a clinical depression screener whose criterion validity was earned in primary-care and clinical populations against diagnostic interviews. Workplace deployment is a different context: screening in a lower-prevalence, largely non-help-seeking workforce reduces positive predictive value, the somatic-confounding problem is relevant to occupational groups with physical demands or illness, and no retrieved study validates the PHQ-9 against workplace criteria. The instrument's clinical provenance must not be read as endorsement of clinical-grade performance in a workplace-wellbeing programme; deployments that used it in worker samples treated it as an off-the-shelf symptom measure rather than validating it there (Doki et al 2024).

References (26)

  1. Kroenke K, Spitzer RL, Williams JB (2001). The PHQ-9: validity of a brief depression severity measure https://doi.org/10.1046/j.1525-1497.2001.016009606.x
  2. Kroenke K, Spitzer RL, Williams JB, Lowe B (2010). The Patient Health Questionnaire Somatic, Anxiety, and Depressive Symptom Scales: a systematic review https://doi.org/10.1016/j.genhosppsych.2010.03.006
  3. Levis B, Benedetti A, Thombs BD, et al (2019). Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: individual participant data meta-analysis https://doi.org/10.1136/bmj.l1476
  4. Negeri ZF, Levis B, Sun Y, et al (2021). Accuracy of the Patient Health Questionnaire-9 for screening to detect major depression: updated systematic review and individual participant data meta-analysis https://doi.org/10.1136/bmj.n2183
  5. Wu Y, Levis B, Riehm KE, et al (2019). Equivalency of the diagnostic accuracy of the PHQ-8 and PHQ-9: a systematic review and individual participant data meta-analysis https://doi.org/10.1017/S0033291719001314
  6. He C, Levis B, Riehm KE, et al (2019). The Accuracy of the Patient Health Questionnaire-9 Algorithm for Screening to Detect Major Depression: An Individual Participant Data Meta-Analysis https://doi.org/10.1159/000502294
  7. Levis B, Sun Y, He C, et al (2020). Accuracy of the PHQ-2 Alone and in Combination With the PHQ-9 for Screening to Detect Major Depression https://doi.org/10.1001/jama.2020.6504
  8. Neupane D, Levis B, Bhandari PM, et al (2021). Selective cutoff reporting in studies of the accuracy of the Patient Health Questionnaire-9 and Edinburgh Postnatal Depression Scale https://doi.org/10.1002/mpr.1873
  9. Levis B, Bhandari PM, Neupane D, et al (2024). Data-Driven Cutoff Selection for the Patient Health Questionnaire-9 Depression Screening Tool https://doi.org/10.1001/jamanetworkopen.2024.29630
  10. Fischer F, Levis B, Falk C, et al (2021). Comparison of different scoring methods based on latent variable models of the PHQ-9: an individual participant data meta-analysis https://doi.org/10.1017/S0033291721000131
  11. Hong ZM, Williams J, Bulloch A, et al (2022). Alternative scoring of the Patient Health Questionnaire-9 in neurological populations: an approach based on a predictive algorithm deriving from individual item scores https://doi.org/10.1016/j.genhosppsych.2022.04.011
  12. Rufino JV, Rodrigues R, Birolim MM, et al (2024). Analysis of the dimensional structure of the Patient Health Questionnaire-9 (PHQ-9) in undergraduate students at a public university in Brazil https://doi.org/10.1016/j.jad.2024.01.051
  13. Blake JJ, Munyombwe T, Fischer F, et al (2024). The factor structure of the Patient Health Questionnaire-9 in stroke: A comparison with a non-stroke population https://doi.org/10.1016/j.jpsychores.2024.111983
  14. Teymoori A, Gorbunova A, Haghish FE, et al (2020). Factorial Structure and Validity of Depression (PHQ-9) and Anxiety (GAD-7) Scales after Traumatic Brain Injury https://doi.org/10.3390/jcm9030873
  15. Wang W, Bian Q, Zhao Y, et al (2014). Reliability and validity of the Chinese version of the Patient Health Questionnaire (PHQ-9) in the general population https://doi.org/10.1016/j.genhosppsych.2014.05.021
  16. Lowe B, Unutzer J, Callahan CM, et al (2004). Monitoring depression treatment outcomes with the Patient Health Questionnaire-9 https://doi.org/10.1097/00005650-200412000-00006
  17. Leavens A, Patten SB, Hudson M, et al (2012). Influence of somatic symptoms on Patient Health Questionnaire-9 depression scores among patients with systemic sclerosis compared to a healthy general population sample https://doi.org/10.1002/acr.21675
  18. Lee EH, Kang EH, Kang HJ, et al (2023). Measurement invariance of the patient health questionnaire-9 depression scale in a nationally representative population-based sample https://doi.org/10.3389/fpsyg.2023.1217038
  19. Di Matteo R, Bolgeo T, Simonelli N, et al (2025). Psychometric Properties and Measurement Invariance of the Patient Health Questionnaire 9 in an Italian Coronary Heart Disease Population https://doi.org/10.1097/JCN.0000000000001178
  20. Pedersen SS, Mathiasen K, Christensen KB, et al (2016). Psychometric analysis of the Patient Health Questionnaire in Danish patients with an implantable cardioverter defibrillator (The DEFIB-WOMEN study) https://doi.org/10.1016/j.jpsychores.2016.09.010
  21. Ajele KW, Idemudia ES (2025). Charting the course of depression care: a meta-analysis of reliability generalization of the Patient Health Questionnaire (PHQ-9) as the measure https://doi.org/10.1007/s44192-025-00181-x
  22. Martinez A, Teklu SM, Tahir P, et al (2023). Validity of the Spanish-Language Patient Health Questionnaires 2 and 9: A Systematic Review and Meta-Analysis https://doi.org/10.1001/jamanetworkopen.2023.36529
  23. Pinto-Meza A, Serrano-Blanco A, Penarrubia MT, et al (2005). Assessing depression in primary care with the PHQ-9: can it be carried out over the telephone? https://doi.org/10.1111/j.1525-1497.2005.0144.x
  24. Barkham M, Saxon D, Hardy GE, et al (2021). Person-centred experiential therapy versus cognitive behavioural therapy delivered in the English Improving Access to Psychological Therapies service for the treatment of moderate or severe depression (PRaCTICED) https://doi.org/10.1016/S2215-0366(21)00083-3
  25. Richards DA, Suckling R (2009). Improving access to psychological therapies: phase IV prospective cohort study https://doi.org/10.1348/014466509X405178
  26. Doki S, Hori D, Takahashi T, et al (2024). Designing a test battery for workers' well-being: the first wave of the Tsukuba Salutogenic Occupational Cohort Study https://doi.org/10.1265/ehpm.23-00372

Record notes

[Upgraded from v0.1 to v0.2 structure in pass two; criterion field split, licence re-verified 2026-07-12.] Overall confidence: the PHQ-9 is the confidence-grading high-water mark. Structural validity, diagnostic criterion validity and internal consistency are High, anchored on the Levis/Thombs individual participant data programme and a reliability-generalisation meta-analysis; responsiveness and the somatic-confounding and cut-point critiques are well evidenced. The genuinely weaker or absent cells are test-retest (present but from few studies, none occupational), measurement invariance (scattered, no occupational coverage, weak over-time evidence), criterion validity against organisational outcomes (absent), and workplace-specific psychometrics (absent). Schema stress-test notes: (1) The single 'criterion_validity' field conflates two very different evidence bases, diagnostic criterion validity (High) versus criterion validity against organisational outcomes (Absent). The maintainers graded it as split and said so in the findings, but a schema that forced one grade would misrepresent the instrument; the field should ideally be divided. (2) 'confidence' is a single scalar per property, yet for several properties the honest grade differs by sub-question (e.g. invariance is scalar-level across age but untested across occupation). The maintainers encoded the dominant grade and qualified it in the justification. (3) The clinical-origin versus workplace-deployment gap is the most important caveat for this registry and does not have a dedicated field; The maintainers carried it in criticisms_controversies and flagged it in constructs_claimed and criterion_validity, but it risks being lost if a reader scans only the property grades. (4) Licence status is factually clear (public domain, Pfizer-released) but the maintainers could not attach a DOI-bearing primary source to the licensing act itself, only to the original validation paper; The maintainers flagged this rather than attach a non-resolvable citation. (5) 'Absent' was used strictly for organisational criterion validity and workplace psychometrics; test-retest was deliberately NOT graded Absent because evidence exists, only sparsely, which the scale's wording ('barely-studied') made a close call between Low and Moderate.