WHO-5 Well-Being Index
In plain English
WHO-5 Well-Being Index is a multi-item scale (5 items). What it claims to measure: The WHO-5 is presented as a unidimensional measure of subjective psychological (hedonic) well-being over the preceding two weeks, tapping positive mood, vitality and general interest. Licence for an employer or vendor: free for non-commercial use only (verified 2026-07-12).
The published evidence is strongest for structural validity, convergent and discriminant validity and internal consistency; moderate for criterion validity against a reference standard and measurement invariance; weak for test-retest reliability and responsiveness to change. No published evidence was located for criterion validity against organisational outcomes (absence, turnover, performance): the registry searched and found none in the sweep to date, so if you need that property evidenced, this instrument does not yet carry it. 2 of the 7 graded properties rest on evidence from clinical, student or otherwise non-working samples; where the property is sensitive to population that evidence cannot carry a High grade, and the reason is stated on each cell below. 3 of the 7 graded properties rest on adult general-population samples rather than samples of working adults; the flag says so on each cell.
Grades summarise the published evidence for this instrument on its own terms; they are not comparable across instruments, and this page makes no recommendation. Whether an instrument fits your workforce is a judgement this registry informs but cannot make. Full evidence, with citations, below. This summary is generated from the record's data, not written by hand.
Identity
Version: WHO-5 (1998 version), the current standard short form derived from earlier WHO-10 and WHO-6/28-item well-being schedules
Structure: 5 items, each positively worded, referring to the previous two weeks, rated 0 (at no time) to 5 (all of the time); raw score 0 to 25 conventionally multiplied by 4 to give 0 to 100
Original citation: Bech P, Olsen LR, Kjoller M, Rasmussen NK (2003) Measuring well-being rather than the absence of distress symptoms: a comparison of the SF-36 Mental Health subscale and the WHO-Five Well-Being Scale. International Journal of Methods in Psychiatric Research 12(2):85-91. https://doi.org/10.1002/mpr.145 (the 1998 five-item version; consolidated evidence reviewed by Topp et al. 2015, https://doi.org/10.1159/000376585)
Steward / publisher: World Health Organization; the scale was developed under the WHO Regional Office for Europe and historically maintained by the Psychiatric Research Unit, Mental Health Centre North Zealand, Denmark (Bech and colleagues). Master versions and translations are distributed by WHO.
Licence status (verified 2026-07-12): Copyright World Health Organization 2024; the WHO-5 master version is available under the Creative Commons Attribution-NonCommercial-ShareAlike 3.0 IGO licence (CC BY-NC-SA 3.0 IGO). Document ref WHO/UCN/MSD/MHE/2024.1. Non-commercial use and adaptation permitted with attribution and share-alike; commercial (including employer/vendor) use requires WHO permission.
Constructs claimed
The WHO-5 is presented as a unidimensional measure of subjective psychological (hedonic) well-being over the preceding two weeks, tapping positive mood, vitality and general interest. It is a measure of the presence of positive well-being rather than of symptom burden, although its origin and much of its validation lie in depression screening. It is a generic well-being index, not a workplace-specific or organisational instrument; workplace studies adopt it unchanged as a general well-being outcome.
Relations to other records
- embedded in Eurofound EWCS/EQLS wellbeing and working-conditions item sets. The EWCS and EQLS questionnaires field the five WHO-5 items as the mental wellbeing index; documented in the Eurofound questionnaire and technical reports for each wave. The WHO-5 evidence on this record is inherited from the WHO-5 record and flagged as parent-form evidence where it is not workplace evidence.
Evidence
Structural validity Highwell-establishedevidence form: canonical
general community sample and general sample; also adolescents with type 1 diabetes, diabetes outpatients, diabetes patients, university sample, bipolar patients, schizophrenia spectrum disorders, post-discharge sample, adolescents and children (flag basis: Sinhala community sample; Bangla general sample; type 1 and type 2 diabetes outpatients; Chinese university sample; adolescent study)
The WHO-5 is predominantly unidimensional: a single well-being factor is recovered across a wide range of populations and languages. The anchoring systematic review concluded high clinimetric validity and treated the scale as a coherent single dimension (Topp 2015). One-factor structures have since been confirmed by confirmatory factor analysis in adolescents with type 1 diabetes (de Wit 2007), in type 1 and type 2 diabetes outpatients (Hajos 2013), in a Chinese university sample (Fung 2022), in Chinese type 2 diabetes patients (Du 2023), in a Sinhala community sample (Perera 2020), in a Bangla general sample (Faruk 2021) and in Portuguese adolescents (Carvalho 2025); principal component analysis in euthymic bipolar patients returned a single factor explaining 59.7% of variance (Bonnín 2017). Item response and Rasch analyses concur: the scale was unidimensional with no differential item functioning by age, sex or inpatient/outpatient status in schizophrenia spectrum disorders, although initial category disordering was only resolved by merging two middle response options (Nielsen 2023), and a nationwide Norwegian post-discharge sample found a one-factor solution explaining 71.7% of variance (Iversen 2025). A five-country clinimetric study (Italy, Poland, Denmark, China, Japan) using Mokken and Rasch analyses reported scalability coefficients ranging from 0.42 to 0.84 with most language versions fitting Rasch expectations and unidimensionality shown explicitly for the Italian and Danish versions (Carrozzino 2022). Two qualifications recur. First, in five-item models the RMSEA is frequently elevated (for example 0.23 in an Australian diabetes sample despite CFI 0.98 and loadings 0.78 to 0.92, Halliday 2017; elevated again in the Norwegian sample, Iversen 2025), an artefact of very low degrees of freedom rather than clear misfit. Second, and more substantively, in a 15-country adolescent study the full five-item WHO-5 did not achieve a good measurement model; a four-item version dropping the first item ('cheerful and in good spirits') was needed (Cosma 2022), while a Japanese child study supported both the five-item and a four-item form and set that item aside on cultural grounds (Adachi 2025). Sample sizes behind these solutions were n=2,310 discharged Norwegian inpatients (Iversen 2025), N=3,249 Australian adults with diabetes (Halliday 2017), 104 patients with bipolar disorder and 40 healthy controls (Bonnín 2017), 3,762 adults across the five countries (Carrozzino 2022), 399 Danish patients with schizophrenia spectrum disorders (Nielsen 2023), 269 Bangladeshi adults, where the one-factor CFA gave RMSEA 0.062, CFI 0.986, TLI 0.964 and SRMR 0.0255 (Faruk 2021), 384 type 1 and 549 type 2 Dutch outpatients (Hajos 2013), 1,916 Portuguese students (Carvalho 2025), 200 Chinese type 2 diabetes patients (Du 2023), 267 Sri Lankan community members, where CFA gave RMSEA 0.03, SRMR 0.02, CFI 0.99 and loadings 0.55 to 0.89 (Perera 2020), 91 adolescents with type 1 diabetes (de Wit 2007), 1,414 Chinese university students across two studies (n=903 and n=511) (Fung 2022), 6,983 Japanese schoolchildren aged 10 to 15 (Adachi 2025) and N=74,071 adolescents from 15 European countries and regions (Cosma 2022); the systematic review covered 213 articles but its abstract gives no sample size or fit statistic (Topp 2015).
High precondition (rubric 1.6): the cited studies that meet the High precondition, each with a sample size and a statistic of this property.
| Study | Sample | Statistic | DOI |
|---|---|---|---|
| Iversen 2025 | 2310 patients | EFA one-factor solution explaining 71.7% of variance, confirmed by CFA; CFI/TLI/SRMR strong, RMSEA elevated; configural, metric and scalar invariance across sex, age, education | 10.1007/s11136-025-04104-9 |
| Halliday 2017 | N=3249 | CFA chi2(5)=834.94, RMSEA=0.23 (90% CI 0.21-0.24), CFI=0.98, TLI=0.96, loadings 0.78-0.92 | 10.1016/j.diabres.2017.07.005 |
| Bonnín 2017 | 104 patients with bipolar disorder and 40 healthy controls | PCA single factor accounting for 59.74% of variance | 10.1016/j.jad.2017.12.006 |
| Carrozzino 2022 | 3762 adult participants | Mokken scalability coefficients 0.42 to 0.84; majority of WHO-5 versions fitted Rasch model; Italian and Danish versions unidimensional on paired t-tests | 10.1016/j.jad.2022.05.111 |
| Faruk 2021 | 269 participants | EFA one factor explaining 38.68% of variance; CFA chi2=295.852, chi2/df=2.017, RMSEA=0.062, CFI=0.986, TLI=0.964, SRMR=0.0255 | 10.1017/gmh.2021.26 |
| Perera 2020 | 267 persons (219 paper, 48 online) | CFA chi2=4.99 (p=0.28), RMSEA=0.03 (90% CI 0.00-0.10), SRMR=0.02, TLI=0.99, CFI=0.99; loadings 0.55-0.89 | 10.1186/s12955-020-01532-8 |
Confidence note (legacy, first pass; not the basis of the grade): High: many good-quality CFA/IRT studies, large total N, consistently unidimensional in adults, with a documented five-item-model RMSEA artefact and a contested first item in some adolescent/cross-cultural samples.
Convergent and discriminant validity Highwell-establishedevidence form: canonical
general Bangladeshi adults and Sinhala-speaking community members; also diabetes, adolescent, depressive-episode, post-discharge, older-adult and university samples (flag basis: Bangladeshi adults; Sinhala-speaking community members)
The WHO-5 correlates strongly and negatively with depression measures and positively with other well-being measures, as expected for a well-being index that shares variance with low mood. Against depression screeners and symptom scales it shows large negative correlations: r = -0.73 with the PHQ-9 in Australian adults with diabetes (Halliday 2017), r = -0.694 with the PHQ-9 and r = -0.610 with the Hamilton Depression Rating Scale in Chinese type 2 diabetes patients (Du 2023), and r = -0.67 with the CES-D in adolescents with type 1 diabetes (de Wit 2007); moderate-to-strong correlations of 0.55 to 0.69 with PHQ, diabetes-distress and SF-12 mental component scores were reported in diabetes outpatients (Hajos 2013). It relates very strongly to depression severity even after controlling for anxiety, supporting some discriminant separation from anxiety (Krieger 2013). Convergent evidence with positive constructs includes r = 0.542 with the Warwick-Edinburgh Mental Well-Being Scale alongside a smaller divergent r = -0.443 with perceived stress (Faruk 2021), and strong associations with life satisfaction and meaning in life against weaker links to physical health (Iversen 2025). Correlations are more modest against broader distress or family-function measures: r = -0.45 with the PHQ-9 and -0.56 with the Kessler K10 in a Sinhala sample (Perera 2020), and r = 0.35 with a family-function measure in older Peruvian adults (Díaz Gamarra 2025). Convergent construct validity with well-being, self-efficacy and self-esteem measures was also supported in a Chinese sample (Fung 2022). Sample sizes behind these correlations were n = 3249 Australian adults with diabetes (Halliday 2017), n = 200 Chinese type 2 diabetes patients (Du 2023), n = 91 adolescents aged 13 to 17 with type 1 diabetes (de Wit 2007), n = 384 Type 1 and n = 549 Type 2 Dutch outpatients (Hajos 2013), n = 414 of whom 207 had a current major depressive episode (Krieger 2013), n = 269 Bangladeshi adults (Faruk 2021), n = 2310 Norwegian adults discharged from inpatient mental health care (Iversen 2025), n = 267 Sinhala-speaking community members aged 16 to 75 (Perera 2020), n = 661 older adults aged 60 to 93 in Lima (Díaz Gamarra 2025) and n = 1414 Chinese university participants across two studies (n = 903 and n = 511) (Fung 2022); the Iversen, Krieger and Fung abstracts describe the correlations as strong, very high and good respectively without reporting numeric coefficients. Additional coefficients in the abstracts include r = 0.60 with the CHQ-CF87 mental health subscale, r = 0.43 with its self-esteem subscale and r = -0.34 with the Diabetes Family Conflict Scale (de Wit 2007), and r = -0.466 with the PAID-20 (Du 2023).
High precondition (rubric 1.6): the cited studies that meet the High precondition, each with a sample size and a statistic of this property; on this population-sensitive property at least one is working-adults or general.
| Study | Sample | Statistic | Population | DOI |
|---|---|---|---|---|
| Halliday 2017 | 3249 | r = -0.73 with PHQ-9 | other | 10.1016/j.diabres.2017.07.005 |
| Faruk 2021 | 269 | r = 0.542 with WEMWBS (convergent); r = -0.443 with PSS-10 (divergent) | general | 10.1017/gmh.2021.26 |
| Hajos 2013 | 384 Type 1 and 549 Type 2 (933 total) | r = 0.55 to 0.69 with PHQ-9, PAID and SF-12 mental component | other | 10.1111/dme.12040 |
| Du 2023 | 200 | r = -0.610 with HAM-D; r = -0.694 with PHQ-9; r = -0.466 with PAID-20 | other | 10.1186/s12888-023-05381-9 |
| Perera 2020 | 267 (219 paper, 48 online) | r = -0.45 with PHQ-9; r = -0.56 with K10 | general | 10.1186/s12955-020-01532-8 |
| de Wit 2007 | 91 | r = -0.67 with CES-D; r = 0.60 with CHQ-CF87 mental health; r = 0.43 with self-esteem; r = -0.34 with DFCS | other | 10.2337/dc07-0447 |
| Díaz Gamarra 2025 | 661 | r = 0.35 with Family APGAR | other | 10.3389/fpubh.2025.1670429 |
Confidence note (legacy, first pass; not the basis of the grade): High: numerous studies, consistent direction and plausible magnitudes; discriminant evidence against anxiety is thinner than convergent evidence against depression.
Criterion validity: reference standard Moderatewell-establishedevidence form: canonical
indirect diabetes, student, schizophrenia and adolescent samples; no working-adult or general-population sample (flag basis: Australian diabetes adults; students; schizophrenia patients; adolescents)
Criterion validity is well established against one target, depression case status. As a depression screener the WHO-5 performs well: against structured diagnostic interviews or established depression scales, reported areas under the ROC curve cluster around 0.81 to 0.88 (AUC 0.87 in Australian diabetes adults, Halliday 2017; 0.823 against clinical interview in Iranian students, Ghazisaeedi 2021; 0.882 in Chinese healthcare students, Yang 2023; 0.838 in schizophrenia patients, Fekih-Romdhane 2024). Sensitivity and specificity depend heavily on the cut-off: a cut-off below 13 gave 0.79/0.79 while below 8 gave 0.44/0.96 in the same Australian sample (Halliday 2017), a cut-off below 50 (on the 0 to 100 metric) gave 0.79 to 0.88 sensitivity in Dutch diabetes outpatients (Hajos 2013) and 0.89/0.86 in adolescents with type 1 diabetes (de Wit 2007); the anchoring review summarised the scale as a sensitive and specific depression screen across fields (Topp 2015). Importantly, this criterion evidence was earned in diagnostic and clinical-population settings, which differ from workplace deployment where no diagnostic gold standard is applied.
Confidence note (legacy, first pass; not the basis of the grade): Moderate: criterion validity against depression is strong and consistent but almost entirely from clinical and disease populations; criterion validity against organisational outcomes is Absent (no such study located).
Criterion validity: organisational Absent (searched; none found in the sweep to date)untestedevidence form: canonical
absence type: population-general searched; none found in the sweep to date
Against organisational and health outcomes proper (sickness absence, staff turnover, diagnosed conditions, productivity) no criterion-validity study was located in this pass. Workplace papers instead report cross-sectional or prospective associations, for example poor WHO-5 well-being being more common with low workplace social capital (Gao 2014), with adverse psychosocial work factors across 34 European countries (Schütte 2014) and prospectively in France (Bertrais 2021), lower well-being among self-employed than salaried workers in small enterprises (Park 2025), and higher odds of poor well-being with lower-quality supervisor behaviour and workplace social capital across 35 European countries (Kizuki 2020). These are construct-relevant associations, not criterion validation against an organisational outcome.
Internal consistency Highwell-establishedevidence form: canonical
direct nurses; also community sample, adults with diabetes, diabetes patients, bipolar patients, schizophrenia patients, students, healthcare students, adolescents and children (flag basis: three countries of nurses; Sinhala community sample; Chinese type 2 diabetes patients; Iranian students; Luxembourg adolescents)
Internal consistency is consistently good to excellent. Cronbach's alpha values cluster in the low-to-high 0.80s to low 0.90s across populations: alpha 0.90 in Australian adults with diabetes (Halliday 2017), 0.88 in Chinese type 2 diabetes patients (Du 2023), 0.82 in adolescents with type 1 diabetes (de Wit 2007), 0.83 in euthymic bipolar patients (Bonnín 2017), 0.85 in a Sinhala community sample (Perera 2020), 0.81 to 0.90 across three countries of nurses (Lara-Cabrera 2022), 0.80 in Portuguese adolescents (Carvalho 2025) and 0.80 in Arabic-speaking schizophrenia patients (Fekih-Romdhane 2024). Higher values around 0.91 to 0.94 appear in Iranian students (Ghazisaeedi 2021), a Norwegian post-discharge sample (Iversen 2025) and Chinese healthcare students (Yang 2023). Where reported, McDonald's omega agrees closely with alpha (omega 0.84 in Luxembourg adolescents, Brisson 2025; omega 0.86 to 0.91 in Japanese children, Adachi 2025; omega 0.908 to 0.935 in Chinese healthcare students, Yang 2023). A lower alpha of 0.75 was reported for the Bangla version (Faruk 2021). Coefficient omega was the reliability index in the updated German norms study (Kliem 2025). Sample sizes behind these coefficients were 3249 Australian adults with diabetes (Halliday 2017), 200 Chinese type 2 diabetes patients (Du 2023), 91 adolescents with type 1 diabetes (de Wit 2007), 104 euthymic bipolar patients and 40 healthy controls (Bonnín 2017), 267 Sinhala-speaking community members (Perera 2020), 678 nurses across Spain, Chile and Norway (Lara-Cabrera 2022), 1916 Portuguese school students (Carvalho 2025), 117 Arabic-speaking schizophrenia patients (Fekih-Romdhane 2024), 400 Iranian university students (Ghazisaeedi 2021), 2310 Norwegian post-discharge inpatients with alpha 0.910 (Iversen 2025), 343 Chinese healthcare students (Yang 2023), 9007 Luxembourg school attendees (Brisson 2025), 6983 Japanese students aged 10 to 15, where alpha was 0.84 to 0.88 alongside the omega values (Adachi 2025), 269 Bangladeshi participants (Faruk 2021) and 2515 German adults in the representative norms sample, whose abstract reports no omega value (Kliem 2025).
High precondition (rubric 1.6): the cited studies that meet the High precondition, each with a sample size and a statistic of this property.
| Study | Sample | Statistic | DOI |
|---|---|---|---|
| Iversen 2025 | 2310 | Cronbach's alpha 0.910 | 10.1007/s11136-025-04104-9 |
| Ghazisaeedi 2021 | 400 | internal consistency (WHO-5) .94 | 10.1007/s11469-021-00483-5 |
| Halliday 2017 | 3249 | alpha 0.90 | 10.1016/j.diabres.2017.07.005 |
| Bonnín 2017 | 104 patients with bipolar disorder and 40 healthy controls | Cronbach's alpha 0.83 | 10.1016/j.jad.2017.12.006 |
| Faruk 2021 | 269 | alpha 0.754 | 10.1017/gmh.2021.26 |
| Brisson 2025 | 9007 | McDonald's omega .84 | 10.1080/00223891.2025.2569138 |
| Carvalho 2025 | 1916 | Cronbach alpha 0.80 | 10.1159/000543728 |
| Du 2023 | 200 | Cronbach's alpha 0.88 | 10.1186/s12888-023-05381-9 |
| Fekih-Romdhane 2024 | 117 | alpha 0.80 | 10.1186/s12888-024-05814-z |
| Perera 2020 | 267 | Cronbach's alpha 0.85 | 10.1186/s12955-020-01532-8 |
| Yang 2023 | 343 | Cronbach's alpha 0.907 to 0.934; McDonald's omega 0.908 to 0.935 | 10.2147/PRBM.S437219 |
| de Wit 2007 | 91 | Cronbach's alpha 0.82 | 10.2337/dc07-0447 |
| Adachi 2025 | 6983 | WHO-5-J alpha 0.84 to 0.88; omega 0.86 to 0.91 (WHO-4-J alpha 0.82 to 0.88) | 10.3389/fpubh.2025.1662332 |
| Lara-Cabrera 2022 | 678 | Cronbach's alphas 0.81 to 0.90 | 10.3390/ijerph191610106 |
Confidence note (legacy, first pass; not the basis of the grade): High: many studies, large total N, alpha and omega both reported and consistently adequate to excellent across diverse populations.
Test-retest reliability Lowthinevidence form: canonical
indirect patients with type 1 diabetes, bipolar patients and healthcare students; no working-adult or general-population sample (flag basis: Danish patients with type 1 diabetes; euthymic bipolar patients; Chinese healthcare students; Sinhala sample)
| Coefficient | Type | Interval | Sample | Population | Evidence form |
|---|---|---|---|---|---|
| 0.87 (95% CI 0.82 to 0.90) | ICC | median five days | not stated in the summary | Danish patients with type 1 diabetes | canonical |
| 0.72 (r); 0.82 (ICC) | r and ICC | two weeks | not stated in the summary | Sinhala community sample | canonical |
| 0.83 | r | ten days | not stated in the summary | euthymic bipolar patients | canonical |
| 0.71 | retest coefficient (type not stated) | not stated | not stated in the summary | Bangla validation sample | canonical |
| 0.803 | ICC | about one week | not stated in the summary | Chinese healthcare students | canonical |
Test-retest reliability is the weakest-evidenced reliability property for the WHO-5 and deserves explicit flagging: dedicated studies are few, intervals are short, and the single study designed specifically to estimate it raised a measurement-error concern. The one purpose-designed test-retest and measurement-error study, in Danish patients with type 1 diabetes using a median five-day interval, found an intraclass correlation of 0.87 (95% CI 0.82 to 0.90) but a large minimal detectable change of 18.56 points on the 0 to 100 scale, which the authors judged a larger measurement error than desirable and flagged for further research (Schougaard 2022). Other estimates are incidental to validation studies and mostly over about two weeks: Pearson r = 0.72 with ICC 0.82 over two weeks in a Sinhala sample (Perera 2020), r = 0.83 over ten days in euthymic bipolar patients (Bonnín 2017), a retest coefficient of 0.71 in the Bangla validation (Faruk 2021) and ICC 0.803 over about one week in Chinese healthcare students (Yang 2023). No workplace-specific test-retest estimate was located in this pass, and intervals long enough to separate stability from genuine change in well-being are not represented.
Confidence note (legacy, first pass; not the basis of the grade): Low: only one purpose-designed study (which itself flagged a large minimal detectable change), remaining estimates incidental with short intervals, none from workplace samples.
Measurement invariance Moderatewell-establishedevidence form: canonical
general representative norms study sample and community sample; also post-discharge sample, older adults, schizophrenia patients, healthcare students, adolescents, children and university students (flag basis: representative German norms study; Sinhala community sample; older Peruvian adults; Chinese healthcare students; UAE university students, a non-working sample)
Measurement invariance across sex and age is generally supported in adults, with more mixed results across countries and in adolescents. In adults, configural, metric and scalar invariance across sex, age and education was reported in a nationwide Norwegian post-discharge sample (Iversen 2025); invariance across gender and age was supported in the representative German norms study (Kliem 2025); configural, metric and scalar invariance across sex held in a Sinhala community sample (Perera 2020); and equivalence across sex was shown in older Peruvian adults (Díaz Gamarra 2025) and cross-sex invariance in Arabic-speaking schizophrenia patients (Fekih-Romdhane 2024). In Chinese healthcare students the scale was invariant across eleven sociodemographic groupings and longitudinally over one week (Yang 2023). Strong measurement invariance across the presence or absence of a current major depressive episode was also demonstrated (Krieger 2013). Cross-national and adolescent evidence is more qualified. Across 15 European countries the five-item model did not achieve acceptable cross-country invariance in adolescents, and a four-item version was required for valid comparison (Cosma 2022); across 43 countries an IRT analysis found many item parameters non-invariant although overall differential test functioning was only modest (Sischka 2025). Full scalar invariance was reported across academic, birth-country, language, sex, socioeconomic and school subgroups in Luxembourg adolescents (Brisson 2025) and across age and gender in Japanese children (Adachi 2025). No study located in this pass tested invariance across occupational groups or between working and non-working populations. Preliminary configural, metric and scalar invariance across gender was also reported for an Arabic version in UAE university students, a non-working sample (Jairoun 2026).
Confidence note (legacy, first pass; not the basis of the grade): Moderate: scalar invariance across sex and age is repeatedly demonstrated in adults and large adolescent samples, but cross-country invariance of the five-item form is contested and occupation-based invariance is untested.
Responsiveness and MIC Lowthinevidence form: canonical
direct German hospital staff; also controlled clinical trials (flag basis: German hospital staff; controlled clinical trials)
Responsiveness is asserted more often than it is quantified with a defined minimal important change. The anchoring review concluded the WHO-5 is responsive and can serve as an outcome measure in controlled clinical trials, balancing wanted and unwanted treatment effects (Topp 2015), and a 2025 disease-area review across 552 studies reported that the scale detects treatment-related changes in well-being across many conditions (Domenech 2025). Direct workplace responsiveness evidence is limited to small intervention studies: a phase-II stress-preventive leadership intervention among German hospital staff reported improved WHO-5 well-being over three months (Stuber 2022). A formal minimal important change for the WHO-5 was not located in this pass; the closest quantitative anchor is the measurement-error study's minimal detectable change of 18.56 points on the 0 to 100 scale, which is a distribution-based detectable-change threshold rather than an anchor-based minimal important change and was itself flagged as large (Schougaard 2022).
Confidence note (legacy, first pass; not the basis of the grade): Low: responsiveness is supported qualitatively by two large reviews and small intervention studies, but a validated anchor-based minimal important change was not located, and workplace-specific responsiveness rests on small samples.
Populations, languages and norms
The WHO-5 has been translated into more than 30 languages and validated across an unusually broad range of populations (Topp 2015); a 2025 review catalogued its use across essentially all major disease areas in 552 studies (Domenech 2025). Validation samples located in this pass span general community adults (German, Kliem 2025; Sinhala, Perera 2020; Bangla, Faruk 2021), older adults (Bonsignore 2001; Díaz Gamarra 2025), adolescents and children (43-country and 15-country adolescent studies, Sischka 2025, Cosma 2022, in which the five-item form did not fit and a four-item form was recommended; Luxembourg, Brisson 2025; Portugal, Carvalho 2025; Japan, Adachi 2025), diabetes populations (de Wit 2007, Hajos 2013, Halliday 2017, Du 2023), severe mental illness (Nielsen 2023, Fekih-Romdhane 2024, Bonnín 2017, Iversen 2025) and occupational groups (nurses, Lara-Cabrera 2022; medical educators, Chan 2022; European employees, Schütte 2014). Updated general-population norms exist for Germany (Kliem 2025). No UK-specific validation, UK normative benchmark or UK workplace norm was located in this pass; the nearest representative reference distribution is the German general-population norm set, and no workplace norm set was located in any country. An Arabic adaptation for university students in the UAE reported alpha 0.833, composite reliability 0.848 and preliminary configural, metric and scalar invariance across gender, extending Arabic-language evidence to the Gulf student population (Jairoun 2026). Sample sizes across these validation and norm studies were 2515 adults in the German representative norm sample (Kliem 2025), 267 (Perera 2020), 269 (Faruk 2021), 367 subjects over 50 (Bonsignore 2001), 661 Peruvian older adults aged 60 to 93 with alpha 0.84 and omega 0.84 (Díaz Gamarra 2025), 9007 (Brisson 2025), 1916 (Carvalho 2025), 6983 (Adachi 2025), 74,071 adolescents from 15 European countries in HBSC 2018 (Cosma 2022), 91 (de Wit 2007), 384 type 1 and 549 type 2 Dutch outpatients (Hajos 2013), 3249 (Halliday 2017), 200 (Du 2023), 399 Danish schizophrenia spectrum patients (Nielsen 2023), 117 (Fekih-Romdhane 2024), 104 patients and 40 controls (Bonnín 2017), 2310 (Iversen 2025), 678 nurses in three countries (Lara-Cabrera 2022), 435 Hong Kong medical educators (Chan 2022), 33,443 employees from 34 European countries in the 2010 EWCS (Schütte 2014) and 568 UAE students (Jairoun 2026); the 2015 review covered 213 articles (Topp 2015) and the 43-country adolescent study drew on HBSC 2022 data without a total sample size in its abstract (Sischka 2025). Rater note (rubric 1.6, C-0009): of the sixteen counted entries on this cell, the one that carries a cut-off with its accuracy figures is Hajos 2013, validated in diabetes patients; the cell's population requirement is met by its working-adult and general-population entries and the cell holds under the populations rule as written. A working-adult or general-population cut-off or norm with its value is on the full-text check list.
Criticisms and controversies
Four themes recur in the critical literature. First, the identity of the construct: although marketed as a positive well-being index, the WHO-5's validation base is dominated by depression screening, and it correlates so strongly with depression measures (for example r = -0.73, Halliday 2017) that some authors treat it as effectively a measure of the severity of depression (Krieger 2013). This matters for a workplace registry: the scale's criterion evidence was earned in diagnostic and disease settings, so using it to screen a workforce is deploying a clinically validated instrument outside the diagnostic context in which that validity was established, without a diagnostic gold standard in play. Second, the first item ('cheerful and in good spirits') is repeatedly identified as problematic in cross-cultural and adolescent samples, to the point that a four-item WHO-4 has been proposed for valid cross-country adolescent comparison (Cosma 2022) and adopted on cultural grounds elsewhere (Adachi 2025). Third, cut-off scores for likely depression vary substantially between studies and populations (for example optimal cut-offs corresponding to different thresholds across Halliday 2017, Ghazisaeedi 2021, Du 2023, Fekih-Romdhane 2024), so no single cut-off transfers automatically to a new setting. Fourth, structural analyses frequently report an elevated RMSEA in the five-item model and occasional response-category disordering (Halliday 2017, Nielsen 2023, Iversen 2025), and floor effects on some items in low-well-being clinical samples (Iversen 2025). Measurement error over short retest intervals has been flagged as larger than desirable (Schougaard 2022).
References (37)
- Topp, Østergaard, Søndergaard, et al. (2015). The WHO-5 Well-Being Index: a systematic review of the literature. https://doi.org/10.1159/000376585
- Domenech, Kasujee, Koscielny, et al. (2025). Systematic Review of the Use of the WHO-5 Well-Being Index Across Different Disease Areas. https://doi.org/10.1007/s12325-025-03266-9
- Bech P, Olsen LR, Kjoller M, Rasmussen NK (2003). Measuring well-being rather than the absence of distress symptoms: a comparison of the SF-36 Mental Health subscale and the WHO-Five Well-Being Scale https://doi.org/10.1002/mpr.145
- Bonsignore, Barkow, Jessen, et al. (2001). Validity of the five-item WHO Well-Being Index (WHO-5) in an elderly population. https://doi.org/10.1007/BF03035123
- de Wit, Pouwer, Gemke, et al. (2007). Validation of the WHO-5 Well-Being Index in adolescents with type 1 diabetes. https://doi.org/10.2337/dc07-0447
- Hajos, Pouwer, Skovlund, et al. (2013). Psychometric and screening properties of the WHO-5 well-being index in adult outpatients with Type 1 or Type 2 diabetes mellitus. https://doi.org/10.1111/dme.12040
- Halliday, Hendrieckx, Busija, et al. (2017). Validation of the WHO-5 as a first-step screening instrument for depression in adults with diabetes: Results from Diabetes MILES - Australia. https://doi.org/10.1016/j.diabres.2017.07.005
- Ghazisaeedi, Mahmoodi, Arpaci, et al. (2021). Validity, Reliability, and Optimal Cut-off Scores of the WHO-5, PHQ-9, and PHQ-2 to Screen Depression Among University Students in Iran. https://doi.org/10.1007/s11469-021-00483-5
- Krieger, Zimmermann, Huffziger, et al. (2013). Measuring depression with a well-being index: further evidence for the validity of the WHO Well-Being Index (WHO-5) as a measure of the severity of depression. https://doi.org/10.1016/j.jad.2013.12.015
- Carrozzino, Christensen, Patierno, et al. (2022). Cross-cultural validity of the WHO-5 Well-Being Index and Euthymia Scale: A clinimetric analysis. https://doi.org/10.1016/j.jad.2022.05.111
- Nielsen, Lauridsen, Østergaard, et al. (2023). Structural validity of the 5-item World Health Organization Well-being Index (WHO-5) in patients with schizophrenia spectrum disorders. https://doi.org/10.1016/j.jpsychires.2023.12.028
- Kliem, Lohmann, Fischer, et al. (2025). Psychometric evaluation and updated community norms of the WHO-5 well-being index, based on a representative German sample. https://doi.org/10.3389/fpsyg.2025.1592614
- Cosma, Költő, Chzhen, et al. (2022). Measurement Invariance of the WHO-5 Well-Being Index: Evidence from 15 European Countries. https://doi.org/10.3390/ijerph19169798
- Sischka, Martin, Residori, et al. (2025). Cross-National Validation of the WHO-5 Well-Being Index Within Adolescent Populations: Findings From 43 Countries. https://doi.org/10.1177/10731911241309452
- Brisson (2025). Psychometric Evaluation and Sociodemographic Measurement Invariance of the WHO-5 Well-Being Index among Adolescents in Luxembourg. https://doi.org/10.1080/00223891.2025.2569138
- Schougaard, Laurberg, Lomborg, et al. (2022). Test-retest reliability and measurement error of the WHO-5 Well-being Index and the Problem Areas in Diabetes questionnaire (PAID) used in telehealth among patients with type 1 diabetes. https://doi.org/10.1186/s41687-022-00505-3
- Bonnín, Yatham, Michalak, et al. (2017). Psychometric properties of the well-being index (WHO-5) spanish version in a sample of euthymic patients with bipolar disorder. https://doi.org/10.1016/j.jad.2017.12.006
- Fung, Kong, Liu, et al. (2022). Validity and Psychometric Evaluation of the Chinese Version of the 5-Item WHO Well-Being Index. https://doi.org/10.3389/fpubh.2022.872436
- Perera, Jayasuriya, Caldera, et al. (2020). Assessing mental well-being in a Sinhala speaking Sri Lankan population: validation of the WHO-5 well-being index. https://doi.org/10.1186/s12955-020-01532-8
- Faruk, Alam, Chowdhury, et al. (2021). Validation of the Bangla WHO-5 Well-being Index. https://doi.org/10.1017/gmh.2021.26
- Lara-Cabrera, Betancort, Muñoz-Rubilar, et al. (2022). Psychometric Properties of the WHO-5 Well-Being Index among Nurses during the COVID-19 Pandemic: A Cross-Sectional Study in Three Countries. https://doi.org/10.3390/ijerph191610106
- Yang, Ma, Huang, et al. (2023). Measurement Properties and Optimal Cutoff Point of the WHO-5 Among Chinese Healthcare Students. https://doi.org/10.2147/PRBM.S437219
- Du, Jiang, Lloyd, et al. (2023). Validation of Chinese version of the 5-item WHO well-being index in type 2 diabetes mellitus patients. https://doi.org/10.1186/s12888-023-05381-9
- Fekih-Romdhane, Al Mouzakzak, Abilmona, et al. (2024). Validation and optimal cut-off score of the World Health Organization Well-being Index (WHO-5) as a screening tool for depression among patients with schizophrenia. https://doi.org/10.1186/s12888-024-05814-z
- Iversen, Kjøllesdal, Ellingsen-Dalskau, et al. (2025). Psychometric performance of the WHO-5 well-being index in a nationwide sample of inpatients discharged from specialised mental health care. https://doi.org/10.1007/s11136-025-04104-9
- Adachi, Takahashi, Mori, et al. (2025). Psychometric validation of the WHO-5 and WHO-4 well-being index scales for assessing psychological well-being and detecting depression in Japanese school-aged children: a community-based study. https://doi.org/10.3389/fpubh.2025.1662332
- Gao, Weaver, Dai, et al. (2014). Workplace social capital and mental health among Chinese employees: a multi-level, cross-sectional study. https://doi.org/10.1371/journal.pone.0085005
- Schütte, Chastang, Malard, et al. (2014). Psychosocial working conditions and psychological well-being among employees in 34 European countries. https://doi.org/10.1007/s00420-014-0930-0
- Kizuki, Fujiwara (2020). Quality of supervisor behaviour, workplace social capital and psychological well-being. https://doi.org/10.1093/occmed/kqaa070
- Bertrais, HÉRault, Chastang, et al. (2021). Multiple psychosocial work exposures and well-being among employees: prospective associations from the French national Working Conditions Survey. https://doi.org/10.1177/14034948211008385
- Stuber, Seifried-Dübon, Tsarouha, et al. (2022). Feasibility, psychological outcomes and practical use of a stress-preventive leadership intervention in the workplace hospital: the results of a mixed-method phase-II study. https://doi.org/10.1136/bmjopen-2021-049951
- Park, Kim, Sung (2025). Factors Affecting Subjective Well-Being in Workers at Small-Sized Enterprises: A Cross-Sectional Study from the 6th Korean Working Conditions Survey. https://doi.org/10.3349/ymj.2024.0441
- Chan, Liu, Lam, et al. (2022). Validation of the World Health Organization Well-Being Index (WHO-5) among medical educators in Hong Kong: a confirmatory factor analysis. https://doi.org/10.1080/10872981.2022.2044635
- Carvalho, Vieira Martins, Azevedo, et al. (2025). World Health Organization's Well-Being Index - WHO-5: Psychometric Performance of the Portuguese Version for Adolescents. https://doi.org/10.1159/000543728
- Díaz Gamarra M, et al. (2025). Psychometric evidence of the WHO-5 well-being index in a sample of participants from hospitals and older adults care centers in Peru https://doi.org/10.3389/fpubh.2025.1670429
- Lara-Cabrera ML, Bjorngaard JH, Salvesen O, et al. (2020). Psychometric properties of the Five-item World Health Organization Well-being Index used in mental health services: Protocol for a systematic review https://doi.org/10.1111/jan.14445
- Jairoun AA, et al. (2026). Translation, cultural adaptation, and psychometric validation of the WHO-5 well-being index for Arabic-speaking university students. https://doi.org/10.3389/feduc.2026.1899754
Record notes
[Upgraded from v0.1 to v0.2 structure in pass two; criterion field split, licence re-verified 2026-07-12.] Overall confidence in the WHO-5 as a well-being measure is High for structural validity, internal consistency, convergent validity and breadth of populations/languages; Moderate for measurement invariance (adults yes, cross-country five-item form contested); Low for test-retest reliability and for responsiveness with a defined minimal important change; and effectively Absent for criterion validity against organisational outcomes and for workplace norms. Schema stress-test notes: (1) The schema's single 'criterion_validity' field forced two very different evidence states into one cell, namely strong criterion evidence against depression versus no located evidence against organisational outcomes (absence, turnover, diagnosed conditions). The maintainers reported both explicitly and graded to the audience's actual need, but a registry might benefit from separating clinical-criterion from organisational-criterion validity. (2) The WHO-5 versus WHO-4 question sits awkwardly: the WHO-4 is a proposed reduced form, not a separate instrument, and evidence about item 1 belongs partly under structural validity, partly under invariance, and partly under criticisms; The maintainers have cross-referenced rather than duplicated. (3) The clinical-origin caveat required by the brief cuts across identity, criterion validity and criticisms; The maintainers have stated it in each place rather than confining it to one field. (4) Licensing 'free to use' is well supported in the secondary literature but the maintainers did not retrieve a formal WHO licence document in the latest review pass; the claim rests on peer-reviewed statements, so the maintainers have qualified it accordingly. (5) No UK data of any kind (validation, norms, workplace) surfaced; under rubric 1.1 that is a fact recorded in the populations cell, not a downgrade; the indirect flags on this record rest on population type (clinical and student samples where the working-adult evidence is missing).