GAD-7 (Generalised Anxiety Disorder-7)
In plain English
GAD-7 (Generalised Anxiety Disorder-7) is a multi-item scale (7 items). What it claims to measure: Severity of generalised anxiety disorder symptoms over the preceding two weeks. Licence for an employer or vendor: free, no permission needed (verified 2026-07-12).
The published evidence is strongest for structural validity and internal consistency; moderate for convergent and discriminant validity, criterion validity against a reference standard, measurement invariance and responsiveness to change; weak for test-retest reliability. No published evidence was located for criterion validity against organisational outcomes (absence, turnover, performance): the registry searched and found none in the sweep to date, so if you need that property evidenced, this instrument does not yet carry it. The evidence base for structural validity is contested in the published literature. 2 of the 7 graded properties rest on evidence from clinical, student or otherwise non-working samples; where the property is sensitive to population that evidence cannot carry a High grade, and the reason is stated on each cell below. 5 of the 7 graded properties rest on adult general-population samples rather than samples of working adults; the flag says so on each cell.
Grades summarise the published evidence for this instrument on its own terms; they are not comparable across instruments, and this page makes no recommendation. Whether an instrument fits your workforce is a judgement this registry informs but cannot make. Full evidence, with citations, below. This summary is generated from the record's data, not written by hand.
Identity
Version: Original 7-item GAD-7 (2006); a 2-item short form (GAD-2, first two items) is also distributed
Structure: 7
Original citation: Spitzer RL, Kroenke K, Williams JBW, Lowe B. A brief measure for assessing generalized anxiety disorder: the GAD-7. Arch Intern Med. 2006;166(10):1092-1097. doi:10.1001/archinte.166.10.1092
Steward / publisher: Developed by Robert L. Spitzer, Janet B.W. Williams, Kurt Kroenke and colleagues with an educational grant from Pfizer Inc.; distributed by the steward site phqscreeners.com (copyright Pfizer Inc.)
Licence status (verified 2026-07-12): Free to use, no permission required. The steward's currently distributed GAD-7 English PDF (phqscreeners.com) carries the footer 'No permission required to reproduce, translate, display or distribute', developed with an educational grant from Pfizer Inc.; the steward instrument entry also carries 'Copyright (c) Pfizer Inc. All rights reserved.' This is a no-permission-required grant with copyright retained by Pfizer, NOT a formal open or Creative Commons licence. Confirmed at the last verification attempt from the steward's distributed PDF footer text (body content, not merely page titles) and the LOINC mirror of the steward entry; direct HTML fetches of phqscreeners.com returned HTTP 403.
Constructs claimed
Severity of generalised anxiety disorder symptoms over the preceding two weeks. Functions both as a screener for probable GAD and as a continuous self-report measure of anxiety-symptom severity.
Evidence
Structural validity Highcontestedevidence form: canonical
general nationally representative general-population samples; also perinatal women, coronary heart disease, primary-care patients, college students, young adults and psychotherapy samples (flag basis: nationally representative German general-population confirmatory factor analysis; four nationally representative European samples including the UK; Spanish primary-care patients; US college-student study)
GAD-7 was designed as, and is most often confirmed as, a single-factor (unidimensional) scale. The original development study confirmed anxiety and depression as distinct dimensions (Spitzer 2006), and a nationally representative German general-population confirmatory factor analysis (N=5030) substantiated a one-dimensional structure with factorial invariance for gender and age (Lowe 2008). Unidimensional fit has been replicated across many settings: Cypriot perinatal women (Vogazianos 2022), an Italian coronary heart disease sample (Bolgeo 2023), four nationally representative European samples including the UK (Shevlin 2022), Canadian young adults where a one-factor model fit best (Riglea 2025), and 20 Czech samples (N=5529) supporting a unidimensional structure (Cigler 2026). However, a competing two-factor structure separating cognitive-emotional from somatic items recurs: a network analysis of Spanish primary-care patients revealed a two-factor solution (Moriana 2021), a US college-student study fit both one- and two-factor models (White 2023), a large psychotherapy sample found the scale technically multidimensional though sum scores remained justifiable (Stochl 2020), and a Malaysian study preferred a six-item second-order model (Pheh 2023). The one-factor model dominates practice and fits well in most samples, but the dimensionality question is genuinely unsettled. Sample sizes were 2740 primary care patients, 965 of them interviewed, in the development study (Spitzer 2006), 457 Cypriot perinatal women, with CFI 0.999 and SRMR 0.027 for the single-factor model (Vogazianos 2022), 398 Italian coronary heart disease patients (Bolgeo 2023), 6054 combined across the UK, Ireland, Spain and Italy (Shevlin 2022), 799 Canadian young adults at age 30 (Riglea 2025), 1704 primary care patients (Moriana 2021), 582 US college students, where the one-factor model gave CFI .994 and RMSEA .098, the two-factor model CFI .997 and RMSEA .069, and the inter-factor correlation was .91 (White 2023), 22,362 psychotherapy patients (Stochl 2020) and 1272 Malaysian adults (Pheh 2023).
Sub-grades (evidence differs by subgroup):
- {"subgroup": "one-factor (unidimensional) model", "grade": "High", "note": "Well replicated across languages and settings; the dominant and recommended scoring model"}
- {"subgroup": "cognitive-emotional vs somatic two-factor model", "grade": "Moderate", "note": "Recurs in network analyses and some CFAs; keeps dimensionality contested rather than settled"}
High precondition (rubric 1.6): the cited studies that meet the High precondition, each with a sample size and a statistic of this property.
| Study | Sample | Statistic | DOI |
|---|---|---|---|
| White 2023 | 582 | one-factor RMSEA .098, CFI .994; two-factor RMSEA .069, CFI .997; inter-factor r .91; alpha and omega 0.91 total | 10.1037/tps0000382 |
| Vogazianos 2022 | 457 (222 pregnant, 235 postpartum) | single factor by parallel analysis; CFA chi-square 21.207 (p 0.096), CFI 0.999, SRMR 0.027; alpha 0.907, omega 0.909 | 10.1186/s12884-022-05127-7 |
Convergent and discriminant validity Moderatewell-establishedevidence form: canonical
general general-population sample; also college students (flag basis: German general population; US college students)
Convergent validity is consistently supported. In the German general population GAD-7 correlated r=0.64 with the PHQ-2 depression module and r=-0.43 with the Rosenberg Self-Esteem Scale (Lowe 2008). In US college students the total score correlated r=0.70 with the trait scale of the State-Trait Anxiety Inventory (convergent) and only r=-0.04 with a behavioural activation reward subscale (discriminant) (White 2023). Increasing scores were strongly associated with multiple domains of functional impairment in the original study (Spitzer 2006). The recurring discriminant concern is the strong overlap with depression: GAD-7 and PHQ-9 anxiety and depression factors are highly correlated and frequently co-occur (Stochl 2020; Bolgeo 2023), so the scale distinguishes anxiety from unrelated constructs well but discriminates anxiety from depression less cleanly.
Criterion validity: reference standard Moderatewell-establishedevidence form: canonical
indirect primary-care and epilepsy patient samples; no working-adult or general-population sample (flag basis: primary-care patient study; Taiwanese epilepsy patients)
Against a structured or semi-structured clinical interview, GAD-7 has been extensively evaluated. The original primary-care patient study identified a cut-off of 10 with sensitivity 89% and specificity 82% for GAD (Spitzer 2006). An early diagnostic meta-analysis (12 samples, 5223 participants) found pooled sensitivity 0.83 and specificity 0.84 at a cut-off of 8, with cut-offs 7 to 10 performing similarly (Plummer 2016). The most comprehensive synthesis, a 2025 Cochrane review of 48 studies (19,228 participants, 27 countries, 24 languages), reported that at the recommended cut-off of 10 or higher the GAD-7 summary sensitivity was 0.64 (95% CI 0.56 to 0.72) and specificity 0.91 (95% CI 0.87 to 0.93) for detecting GAD, with an area under the curve of 0.86; for detecting any anxiety disorder sensitivity fell to 0.48 (specificity 0.91) (Akturk 2025). The pooled sensitivity at cut-off 10 is therefore markedly lower than the original single-study estimate, with pronounced heterogeneity, and the Cochrane authors caution that the summary estimates are rough averages that may deviate substantially in specific situations. Setting-specific validations against interview (for example Taiwanese epilepsy patients, optimal cut-off 7) add further threshold variability (Shih 2022). The original study interviewed 965 of 2,740 adult patients across 15 US primary care clinics (Spitzer 2006); the Cochrane estimate for generalised anxiety disorder drew on 35 studies at a median prevalence of 12%, with the GAD-2 at 3 or higher giving sensitivity 0.68 and specificity 0.86 (Akturk 2025); the earlier pooled estimate was based on 11 samples for GAD, with 95% CIs of 0.71 to 0.91 for sensitivity and 0.70 to 0.92 for specificity (Plummer 2016); and the Taiwanese epilepsy validation involved 109 patients, 17 (15.9%) with GAD on the MINI, without reporting sensitivity or specificity in its abstract (Shih 2022).
Criterion validity: organisational Absent (searched; none found in the sweep to date)untestedevidence form: canonical
absence type: population-general searched; none found in the sweep to date
No study validating the full GAD-7 against objective organisational outcomes (recorded sickness absence, turnover, or measured job performance) was located in the latest review pass. The nearest evidence is adjacent rather than direct: a study of 4953 working Australians linked probable anxiety to worse self-reported presenteeism and absenteeism on the WHO Health and Work Performance Questionnaire, but it used the 2-item GAD-2, not the full GAD-7, and relied on self-reported rather than employer-recorded outcomes (Deady 2021). A Polish validation among employees related GAD-7 to professional burnout and psychological distress, again self-reported constructs rather than organisational records (Basinska 2023). The original study related GAD-7 to self-reported disability days and functional impairment, not to work-context criterion outcomes (Spitzer 2006). Organisational criterion validity for the fielded GAD-7 is therefore essentially untested.
Internal consistency Highwell-establishedevidence form: canonical
general general-population sample; also college students, perinatal women, cardiac, epilepsy and HIV samples (flag basis: German general population; US college students; Cypriot perinatal women; Taiwanese epilepsy patients; Swahili HIV sample)
Internal consistency is uniformly high across populations and languages. Cronbach's alpha was 0.89, identical across all gender and age subgroups, in the German general population (Lowe 2008); 0.91 in US college students (White 2023); alpha 0.907 with McDonald's omega 0.909 in Cypriot perinatal women (Vogazianos 2022); alpha 0.89 with composite reliability 0.90 in an Italian cardiac sample (Bolgeo 2023); and 0.928 in Taiwanese epilepsy patients (Shih 2022). Lower but still acceptable values appear in some translations, for example alpha 0.82 (95% CI 0.78 to 0.85) in a Swahili HIV sample (Nyongesa 2020) and a median alpha of 0.86 across 20 Czech samples (Cigler 2026). Sample sizes were n = 5030 in the German household survey (Lowe 2008), n = 582 undergraduates, who also gave omega = 0.91 (White 2023), n = 457 perinatal women (222 pregnant, 235 postpartum) (Vogazianos 2022), n = 398 coronary heart disease inpatients (Bolgeo 2023), n = 109 epilepsy patients (Shih 2022), n = 450 adults living with HIV (Nyongesa 2020) and N = 5529 across the 20 Czech samples (Cigler 2026).
High precondition (rubric 1.6): the cited studies that meet the High precondition, each with a sample size and a statistic of this property.
| Study | Sample | Statistic | DOI |
|---|---|---|---|
| Bolgeo 2023 | 398 (mean age 64.7; 78.9% male) | Cronbach's alpha 0.89; composite reliability 0.90 | 10.1016/j.jad.2023.04.140 |
| Cigler 2026 | 5529 across 20 adult samples | Median alpha 0.86 | 10.1016/j.janxdis.2026.103149 |
| Shih 2022 | 109 (17 with GAD) | Cronbach's alpha 0.928 | 10.1016/j.jfma.2022.04.018 |
| White 2023 | 582 (mean age 19.0; 79.4% women; 81.6% White) | Total score alpha 0.91 and omega 0.91; cognitive-emotional factor alpha 0.90, omega 0.91; somatic tension alpha 0.76, omega 0.77 | 10.1037/tps0000382 |
| Lowe 2008 | 5030 (53.6% female; mean age 48.4) | Alpha 0.89, identical across all subgroups | 10.1097/MLR.0b013e318160d093 |
| Vogazianos 2022 | 457 (222 pregnant, 235 postpartum) | Alpha 0.907; omega 0.909 | 10.1186/s12884-022-05127-7 |
| Nyongesa 2020 | 450 | Alpha 0.82 (95% CI 0.78 to 0.85); test-retest ICC 0.70 | 10.1186/s12991-020-00312-4 |
Test-retest reliability Lowthinevidence form: canonical
general general-population adults; also primary care, adults living with HIV and psychiatric adults (flag basis: Czech general-population and psychiatric adults; US primary-care patients; Adults living with HIV, Kilifi, Kenya)
| Coefficient | Type | Interval | Sample | Population | Evidence form |
|---|---|---|---|---|---|
| 0.83 | ICC | not reported in retrieved sources | not reported in retrieved sources | US primary-care patients (original validation) | canonical |
| 0.59 | ICC | 2 weeks | 60 | Adults living with HIV, Kilifi, Kenya (Swahili version) | canonical |
| 0.46 to 0.53 | r | repeated measures over 2 to 4 time points spanning major life events | subset of 5529 | Czech general-population and psychiatric adults | canonical |
Dedicated short-interval test-retest studies of GAD-7 are sparse. The original validation reportedly gave ICC 0.83, but its interval and sample size were not captured in sources retrieved in the latest review pass (the value is cited second-hand in Nyongesa 2020, doi:10.1186/s12991-020-00312-4). A 2-week ICC of 0.59 was found in a Swahili HIV sample (Nyongesa 2020), and the Czech multi-sample study found moderate stability of r=0.46 to 0.53, but over intervals spanning major life events rather than a fixed short retest window (Cigler 2026, doi:10.1016/j.janxdis.2026.103149). Test-retest reliability in a working-adult population is untested. The absence of a consistent, short-interval, well-sampled retest estimate is itself the finding.
Measurement invariance Moderatewell-establishedevidence form: canonical
general nationally representative samples; also treatment-seeking samples, adolescents, working-age and older adults, traumatic brain injury, psychotherapy and partial-hospital samples (flag basis: four European nationally representative samples; very large UK treatment-seeking samples; Canadian adolescents; UK working-age and older adults; partial-hospital sample)
Measurement invariance is one of the most heavily studied properties of the GAD-7, and it is largely supported across groups but contested over time. Across sex/gender, invariance or an absence of differential item functioning is repeatedly demonstrated, including in very large UK treatment-seeking samples (N=165,872) (Saunders 2023), Canadian adolescents where strict invariance held by sex and grade (Romano 2021), and four European nationally representative samples with no DIF across sex, age or country (Shevlin 2022). Across age, invariance held between UK working-age and older adults with only limited DIF (Delamain 2024). Across language and country, invariance is supported across English and French (Riglea 2025), across 18 countries after traumatic brain injury (Teymoori 2020), and with full scalar invariance in rural India (De Man 2021). Longitudinal invariance is where the literature conflicts: strict temporal invariance was established across 10 psychotherapy sessions (Stochl 2020) and across time in Czech and Canadian samples (Cigler 2026; Riglea 2025), yet longitudinal invariance was NOT established in a partial-hospital sample, implying that raw pre-post change scores may be unreliable in that setting (Ong 2021). Further sample sizes and invariance levels: n = 2137 patients six months after traumatic brain injury, with negligible DIF (Teymoori 2020); n = 166,816 IAPT patients (159,325 working age, 7491 older) plus a propensity-matched n = 5868 (Delamain 2024); n = 799 Canadian young adults at age 30 with partial strong invariance across sex and full strong invariance across language, and n = 633 with strong invariance across ages 30, 34 and 35 (Riglea 2025); N = 5529 across 20 Czech samples with invariance to the residual level across gender, clinical status and time (Cigler 2026); n = 59,052 adolescents with strict invariance (Romano 2021); N = 22,362 psychotherapy patients (Stochl 2020); n = 4,323 partial hospital patients (Ong 2021); N = 6,054 across the UK, Ireland, Spain and Italy (Shevlin 2022); and n = 1,209 rural Indian adults (1,007 at risk of diabetes, 202 with diabetes) with full scalar and full or partial residual invariance across age, gender, education, diabetes status and time, and a hierarchical omega of 0.76 (De Man 2021).
Sub-grades (evidence differs by subgroup):
- {"subgroup": "sex / gender", "grade": "High", "note": "Strongly supported across many samples including large UK datasets"}
- {"subgroup": "age", "grade": "High", "note": "Supported including UK working-age versus older adults (Delamain 2024)"}
- {"subgroup": "language / country", "grade": "Moderate", "note": "Supported across several languages and a UK-inclusive four-country study, but by translation rather than exhaustive"}
- {"subgroup": "longitudinal / over time", "grade": "Low", "note": "Contested: strict temporal invariance in some large samples but not established in a partial-hospital sample, complicating change-score use (Ong 2021)"}
- {"subgroup": "occupation / industry", "grade": "Absent", "note": "No study testing invariance across occupational groups located in the latest review pass"}
Responsiveness and MIC Moderatethinevidence form: canonical
indirect chronic-depression trial patients and psychotherapy patients; no working-adult or general-population sample (flag basis: multisite chronic-depression trial; psychotherapy patients)
Responsiveness (sensitivity to change) is demonstrated, but formal minimal important change estimates are sparse. In a multisite chronic-depression trial (N=261), GAD-7 scores fell significantly in patients who improved on the Hamilton depression rating (effect size -0.51 at 12 weeks, -1.0 at 48 weeks) and rose in those who worsened, supporting sensitivity to change (Toussaint 2020). The Czech multi-sample study found clear sensitivity to change in psychotherapy patients and derived a reliable-change threshold of about plus or minus 5.2 points (Cigler 2026). A single, well-anchored minimal important change value validated for a working-adult population was not located in the latest review pass; responsiveness is established while the MIC anchor remains setting-dependent.
Populations, languages and norms
GAD-7 has been fielded and psychometrically evaluated in a very wide range of populations and languages: the 2025 Cochrane review alone drew on 27 countries and 24 languages (Akturk 2025). General-population normative data exist, including German norms by sex and age where roughly 5% scored 10 or higher (Lowe 2008). Validations span primary care, adolescents and older adults, perinatal women, and disease-specific groups (HIV, epilepsy, coronary heart disease, traumatic brain injury), plus students and employees. UK-relevant data come from nationally representative European samples and large UK treatment-seeking (IAPT) cohorts (Shevlin 2022; Saunders 2023; Delamain 2024). Dedicated UK working-population norms were not located in the latest review pass. The Cochrane review pooled 48 studies with 19,228 participants and reported a summary sensitivity of 0.64 and specificity of 0.91 for the GAD-7 at a cut-off of 10 or more for generalised anxiety disorder (Akturk 2025); the German norm sample comprised 5030 subjects with alpha 0.89 (Lowe 2008), the four-country European sample combined 6054 respondents (Shevlin 2022), and the IAPT age-invariance study used 166,816 patients, 159,325 of working age and 7491 aged 65 or over, from eight services (Delamain 2024).
Criticisms and controversies
Recurring criticisms cluster around five points. First, dimensionality is unsettled: the one-factor model dominates but a cognitive-emotional versus somatic two-factor structure recurs in network analyses and some CFAs (Moriana 2021, doi:10.1002/jclp.23217; Pheh 2023, doi:10.1371/journal.pone.0285435). Second, discriminant validity against depression is weak in the sense that GAD-7 and PHQ-9 factors are highly correlated and co-occur, so the scale separates anxiety from depression less cleanly than from unrelated constructs (Stochl 2020, doi:10.1177/1073191120976863). Third, the 2025 Cochrane meta-analysis found only modest pooled sensitivity (0.64) at the standard cut-off of 10 for GAD, with pronounced heterogeneity, meaning the scale misses a meaningful share of cases at that threshold and performs worse for any anxiety disorder (sensitivity 0.48) (Akturk 2025, doi:10.1002/14651858.CD015455). Fourth, longitudinal measurement invariance is not guaranteed, which complicates the common practice of interpreting raw pre-post change scores (Ong 2021, doi:10.1177/10731911211035833). Fifth, the scale targets GAD specifically rather than the full anxiety-disorder spectrum, and its clinical/primary-care origin means any workplace deployment is outside its validation context. Organisational criterion validity is effectively untested.
References (24)
- Spitzer RL, Kroenke K, Williams JBW, Lowe B (2006). A brief measure for assessing generalized anxiety disorder: the GAD-7 https://doi.org/10.1001/archinte.166.10.1092
- Lowe B, Decker O, Muller S, et al. (2008). Validation and standardization of the Generalized Anxiety Disorder Screener (GAD-7) in the general population https://doi.org/10.1097/MLR.0b013e318160d093
- Plummer F, Manea L, Trepel D, McMillan D (2016). Screening for anxiety disorders with the GAD-7 and GAD-2: a systematic review and diagnostic metaanalysis https://doi.org/10.1016/j.genhosppsych.2015.11.005
- Akturk Z, Hapfelmeier A, Fomenko A, et al. (2025). Generalized Anxiety Disorder 7-item (GAD-7) and 2-item (GAD-2) scales for detecting anxiety disorders in adults https://doi.org/10.1002/14651858.CD015455
- White EJ, Karr JE (2025). Psychometric properties of the GAD-7 among college students: reliability, validity, factor structure, and measurement invariance https://doi.org/10.1037/tps0000382
- Cigler H, Patkova Dansova P, Javurkova A, et al. (2026). Validity and factor structure of the Czech GAD-7 across twenty samples and four independent translations https://doi.org/10.1016/j.janxdis.2026.103149
- Moriana JA, Jurado-Gonzalez FJ, Garcia-Torres F, et al. (2021). Exploring the structure of the GAD-7 scale in primary care patients with emotional disorders: a network analysis approach https://doi.org/10.1002/jclp.23217
- Riglea T, Wellman RJ, Sylvestre MP, et al. (2025). Factor structure and measurement invariance of the GAD-7 across time, sex, and language in young adults https://doi.org/10.1016/j.jad.2025.01.117
- Pheh KS, Tan CS, Lee KW, et al. (2023). Factorial structure, reliability, and construct validity of the Generalized Anxiety Disorder 7-item (GAD-7): evidence from Malaysia https://doi.org/10.1371/journal.pone.0285435
- Vogazianos P, Motrico E, Dominguez-Salas S, et al. (2022). Validation of the generalized anxiety disorder screener (GAD-7) in Cypriot pregnant and postpartum women https://doi.org/10.1186/s12884-022-05127-7
- De Man J, Absetz P, Sathish T, et al. (2021). Are the PHQ-9 and GAD-7 suitable for use in India? A psychometric analysis https://doi.org/10.3389/fpsyg.2021.676398
- Ong CW, Pierce BG, Klein KP, et al. (2021). Longitudinal measurement invariance of the PHQ-9 and GAD-7 https://doi.org/10.1177/10731911211035833
- Romano I, Ferro MA, Patte KA, et al. (2021). Measurement invariance of the GAD-7 and CESD-R-10 among adolescents in Canada https://doi.org/10.1093/jpepsy/jsab119
- Saunders R, Moinian D, Stott J, et al. (2023). Measurement invariance of the PHQ-9 and GAD-7 across males and females seeking treatment for common mental health disorders https://doi.org/10.1186/s12888-023-04804-x
- Delamain H, Buckman JEJ, Stott J, et al. (2024). Measurement invariance and differential item functioning of the PHQ-9 and GAD-7 between working age and older adults https://doi.org/10.1016/j.jad.2023.11.048
- Bolgeo T, Di Matteo R, Simonelli N, et al. (2023). Psychometric properties and measurement invariance of the 7-item General Anxiety Disorder scale (GAD-7) in an Italian coronary heart disease sample https://doi.org/10.1016/j.jad.2023.04.140
- Shevlin M, Butter S, McBride O, et al. (2022). Measurement invariance of the Patient Health Questionnaire (PHQ-9) and Generalized Anxiety Disorder (GAD-7) across four European countries https://doi.org/10.1186/s12888-022-03787-5
- Stochl J, Fried EI, Fritz J, et al. (2020). On dimensionality, measurement invariance, and suitability of sum scores for the PHQ-9 and the GAD-7 https://doi.org/10.1177/1073191120976863
- Teymoori A, Real R, Gorbunova A, et al. (2020). Measurement invariance of assessments of depression (PHQ-9) and anxiety (GAD-7) across sex, strata and linguistic backgrounds https://doi.org/10.1016/j.jad.2019.10.035
- Shih YC, Chou CC, Lu YJ, et al. (2022). Reliability and validity of the traditional Chinese version of the GAD-7 in Taiwanese patients with epilepsy https://doi.org/10.1016/j.jfma.2022.04.018
- Toussaint A, Husing P, Gumz A, et al. (2020). Sensitivity to change and minimal clinically important difference of the 7-item Generalized Anxiety Disorder Questionnaire (GAD-7) https://doi.org/10.1016/j.jad.2020.01.032
- Deady M, Collins DAJ, Johnston DA, et al. (2021). The impact of depression, anxiety and comorbidity on occupational outcomes https://doi.org/10.1093/occmed/kqab142
- Nyongesa MK, Mwangi P, Koot HM, et al. (2020). The reliability, validity and factorial structure of the Swahili version of the 7-item generalized anxiety disorder scale (GAD-7) among adults living with HIV from Kilifi, Kenya https://doi.org/10.1186/s12991-020-00312-4
- Basinska MA, Kwissa-Gajewska Z (2023). Psychometric properties of the Polish version of the Generalized Anxiety Disorder scale (GAD-7) in a non-clinical sample of employees https://doi.org/10.13075/ijomeh.1896.02104
Record notes
Overall confidence: GAD-7 has a deep, high-quality evidence base for internal consistency (High), diagnostic accuracy against clinical interview (High, though the best synthesis shows modest pooled sensitivity at cut-off 10 with high heterogeneity), a predominantly unidimensional structure (High but with a genuinely contested two-factor alternative), and cross-group measurement invariance by sex and age (High). That evidence was earned overwhelmingly in clinical, primary-care and general-population samples, making it indirect for working adults; UK general-population and IAPT data exist and are recorded as fact. The honest gaps that most matter for a workplace registry: organisational criterion validity against objective work outcomes is Absent (the nearest evidence uses the derivative GAD-2 and self-reported outcomes); dedicated short-interval test-retest studies are sparse and mixed (Low), with the original ICC of 0.83 available only second-hand without interval or n; longitudinal invariance is contested, which bears directly on using the scale to track change; and a working-population minimal important change was not located in the latest review pass. Licence verified on 2026-07-12 from the body text of the steward's currently distributed GAD-7 English PDF (phqscreeners.com footer) and the LOINC mirror of the steward entry, not from the founding paper: free to use, Pfizer copyright retained, no permission required, which is a no-permission-required grant rather than a formal open or Creative Commons licence. Direct HTML fetches of phqscreeners.com returned HTTP 403, so confirmation rests on the steward's own distributed PDF footer and the LOINC mirror. Schema v0.2 recorded these honestly; the main tension was that most invariance/structure evidence is shared with the PHQ-9 in joint studies, which the maintainers have tagged canonical for the GAD-7 factor specifically where the analysis modelled the GAD-7 items separately.