item-set · dataset v0.9.0

ONS-4 personal wellbeing questions

Licence verified: 2026-07-12 · literature last reviewed: 2026-09-01 · grades last confirmed: 2026-09-03 · rubric v1.6

In plain English

ONS-4 personal wellbeing questions is an item set whose items are separate records (4 items). What it claims to measure: The four questions deliberately span three or four conceptually distinct facets of subjective well-being rather than a single latent trait: evaluative well-being (life satisfaction), eudaimonic well-being (the sense that things done in life are worthwhile), ... Licence for an employer or vendor: open licence (verified 2026-07-12).

The published evidence is weak for structural validity, convergent and discriminant validity and criterion validity against a reference standard. No published evidence was located for criterion validity against organisational outcomes (absence, turnover, performance), test-retest reliability, measurement invariance and responsiveness to change: the registry searched and found none in the sweep to date, so if you need those properties evidenced, this instrument does not yet carry them. 1 of the 3 graded properties rest on evidence from clinical, student or otherwise non-working samples; where the property is sensitive to population that evidence cannot carry a High grade, and the reason is stated on each cell below. 2 of the 3 graded properties rest on adult general-population samples rather than samples of working adults; the flag says so on each cell.

Structural validity: LowConvergent validity: LowCriterion (reference standard): Very lowCriterion (organisational): AbsentInternal consistency: n/aTest-retest: AbsentInvariance: AbsentResponsiveness: AbsentLicence: open

Grades summarise the published evidence for this instrument on its own terms; they are not comparable across instruments, and this page makes no recommendation. Whether an instrument fits your workforce is a judgement this registry informs but cannot make. Full evidence, with citations, below. This summary is generated from the record's data, not written by hand.

Who graded this. Every grade in this registry was assigned by one rater employed by the steward (1 rater, 0 independent), with AI assistance in literature retrieval and drafting. Grades and statuses are single-rater and frozen from first publication; they will not move until two named psychometric raters who are not employees of the steward have joined. Until then, automated sweeps add citations and flag cells for review; corrections of fact are made in public; no grade changes. Why, and how to volunteer.

Identity

Version: Four questions as introduced by the UK Office for National Statistics in the Annual Population Survey in 2011; each item uses an 11-point 0 to 10 response scale. The wording has been stable since introduction, with ONS relabelling the set as 'personal well-being' following public focus groups in 2013.

Structure: 4 (life satisfaction; worthwhile; happiness yesterday; anxiety yesterday), each administered and reported as a separate single item; ONS does not publish a summed total.

Original citation: Office for National Statistics (2011). Four personal well-being questions introduced in the Annual Population Survey. No primary journal article; the measure is defined and disseminated in ONS technical guidance and described in Dolan P, Metcalfe R (2012) Measuring Subjective Wellbeing: Recommendations on Measures for use by National Governments. Journal of Social Policy. DOI 10.1017/s0047279411000833

Steward / publisher: UK Office for National Statistics (ONS)

openopen licence for employer or vendor use. Open Government Licence v3.0; use encouraged by the steward.
Licence status (verified 2026-07-12): CONFIRMED. ONS states the four personal well-being questions are part of the government harmonised standards and 'they are freely available to others and their use is encouraged'. All ONS website content is published under the Open Government Licence v3.0 (Crown copyright). No fee, no registration; open status attaches to the original ONS wording and 0-to-10 scale, not to modified derivatives.
Source: https://www.ons.gov.uk/peoplepopulationandcommunity/wellbeing/methodologies/surveysusingthe4officefornationalstatisticspersonalwellbeingquestions (use encouraged; freely available) and OGL v3.0 site-wide footer; corroborated at https://analysisfunction.civilservice.gov.uk/policy-store/personal-well-being/
Steward page as read, Internet Archive: archived copy 1 · archived copy 2
The registry records the licence class and the steward's page, never a price: fees change without notice and the archived page is the record of what was read.

Constructs claimed

The four questions deliberately span three or four conceptually distinct facets of subjective well-being rather than a single latent trait: evaluative well-being (life satisfaction), eudaimonic well-being (the sense that things done in life are worthwhile), and hedonic affect split into positive (happiness yesterday) and negative (anxiety yesterday) components Benson 2019. This tripartite or four-part conception follows the Stiglitz-Sen-Fitoussi and national-accounts-of-well-being tradition that ONS adopted Dolan & Metcalfe 2012. The design intent is that each item is informative in its own right; the anxiety item is a negative-affect measure and is therefore worded and scored in the opposite direction to the other three.

Relations to other records

A relation is published fact about the literature, with the study it rests on. It never says which form an implementer should use.

Evidence

Structural validity Lowuntestedevidence form: canonical

general the instrument is fielded on general populations; the cited studies describe no further sample

state: assessed · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03 · under rubric v1.1: indirectness indirect → general (re-read before first publication, correction C-0005)

There is no confirmatory factor-analytic validation of the original ONS-4 as a unidimensional scale, and this is by design: ONS treats the four items as separate indicators of distinct constructs and publishes each separately rather than as a summed score Benson 2019. The one psychometric study that factor-analyses ONS-4-derived items examined a modified four-point short version (the Personal Wellbeing Score, PWS) embedded alongside other R-Outcomes measures; a scree plot suggested four or six factors and Kaiser's criterion four, with the four well-being items loading together and separately from co-administered health and experience measures, and inter-item correlations of r=0.51 to r=0.77 Benson 2019. That evidence supports the well-being items cohering as a set in that instrument, but it pertains to the reworded four-point PWS, not to the original 0 to 10 ONS-4 wording. Broader multi-country work confirms that evaluative, eudaimonic and affective well-being are empirically separable dimensions rather than one factor, which is consistent with ONS-4 being reported item by item Ruggeri 2020.

Confidence note (legacy, first pass; not the basis of the grade): Low. Direct factor-analytic evidence exists only for a modified derivative (single study, one English social-prescribing sample); the original ONS-4 is not validated as a scale and is not intended to be one.

Convergent and discriminant validity Lowthinevidence form: canonical

general the instrument is fielded on general populations; the cited studies describe no further sample

state: assessed · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03 · under rubric v1.1: indirectness direct → general (re-read before first publication, correction C-0005)

Convergent evidence for the ONS-4 items is largely indirect, resting on the wider single-item life-satisfaction literature and on derivative instruments rather than on the ONS wording itself. Single-item life-satisfaction measures correlate with the multi-item Satisfaction With Life Scale at zero-order r=0.62 to 0.64, rising to r=0.78 to 0.80 after disattenuation, and reproduce the SWLS pattern of associations with health, domain satisfaction and affect almost exactly (mean absolute difference in correlations 0.015 to 0.042) Cheung & Lucas 2014. In the modified PWS derivative, each well-being item correlated with its summary score at r=0.83 to 0.88 and the two evaluative items correlated at r=0.77 Benson 2019. UK multi-instrument comparison datasets now allow ONS-4 to be mapped against WEMWBS, EQ-5D and ICECAP-A, supporting moderate cross-measure convergence while confirming the instruments are not interchangeable Wickramasekera & Tsuchiya 2025. Discriminant behaviour of the anxiety item is notable: it is the negative-affect component and shows the weakest association with the evaluative items, consistent with affect and evaluation being separable Ruggeri 2020.

Confidence note (legacy, first pass; not the basis of the grade): Low. Convergent magnitudes are reasonable but are drawn mostly from single-item life-satisfaction studies and from a modified derivative, not from the original ONS-4 items as fielded.

Criterion validity: reference standard Very lowuntestedevidence form: canonical

indirect adult cohort and older adults; no working-adult or general-population sample described (flag basis: Finnish cohort; older adults)

state: assessed · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03

No study located in this pass tests the ONS-4 items against a diagnostic or health reference standard, and none is expected: the items measure subjective wellbeing, which has no diagnostic gold standard. Criterion evidence for the underlying constructs comes from the general subjective-well-being literature and predicts health and mortality: low self-reported life satisfaction predicted 20-year all-cause mortality in a Finnish cohort of 22,461 adults Koivumaa-Honkanen 2000, life satisfaction predicted all-cause mortality over 22 years in older adults Gana 2016, and a broad review links subjective well-being to health and longevity Diener & Chan 2011. Mendelian randomisation provides mixed causal support, finding little robust causal effect of subjective well-being on cardiometabolic disease Wootton 2018. This is construct-level evidence in general and older cohorts, not instrument-level criterion validation.

Confidence note (legacy, first pass; not the basis of the grade): Very low. Workplace and organisational criterion validity is absent; the available criterion evidence is for the general constructs (health, mortality) in non-workplace populations and does not use the ONS-4 items specifically.

Criterion validity: organisational Absent (searched; none found in the sweep to date)untestedevidence form: canonical

absence type: population-general searched; none found in the sweep to date

state: assessed_absent · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03

No study located in this pass tests ONS-4 against organisational or workplace outcomes such as sickness absence, staff turnover or productivity; criterion evidence for these items in an employment context is absent. A UK natural experiment used ONS-4-style items as outcomes when evaluating a change in the built environment, illustrating policy-outcome use rather than establishing criterion validity against an organisational endpoint Ram 2020.

Internal consistency Not applicable (category difference)untestedevidence form: canonical

absence type: category-error the property does not apply to this construct

state: not_applicable · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03

Internal consistency is not a meaningful property of ONS-4 as ONS uses it, because the four items are reported separately and Cronbach's alpha cannot be computed for a single item Cheung & Lucas 2014. The only alpha located is for the modified four-point PWS summary score, where Cronbach's alpha was 0.90 in one English social-prescribing sample; the authors themselves had expected 0.7 to 0.9 to justify an aggregate score, and note this value sits at the top of that range Benson 2019. That coefficient applies to the reworded PWS, not to the original ONS-4, and a high alpha on four items partly reflects deliberate content breadth rather than redundancy. No omega is reported.

Test-retest reliability Absent (searched; none found in the sweep to date)untestedevidence form: canonical

absence type: population-general searched; none found in the sweep to date

state: assessed_absent · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03

Test-retest reliability for the ONS-4 items as fielded is, on the evidence located in this pass, absent, and this is the clearest evidential gap in the record. The one instrument-specific validation study states explicitly that its anonymous, unlinked data did not permit test-retest reliability, inter-rater reliability or within-individual change to be estimated Benson 2019. The nearest available evidence is not classical test-retest but model-based reliability from the single-item life-satisfaction literature: using latent state-trait models on four national panels (combined N over 68,000), reliability estimates for single-item life satisfaction ranged from about 0.68 to 0.74 Lucas & Donnellan 2012. These estimates concern single-item life satisfaction generally, not the ONS worthwhile, happiness or anxiety items, and they are derived from longitudinal decomposition rather than a short-interval retest. No short-interval test-retest coefficient for any ONS-4 item was located.

Measurement invariance Absent (searched; none found in the sweep to date)untestedevidence form: canonical

absence type: population-general searched; none found in the sweep to date

state: assessed_absent · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03

No formal measurement-invariance analysis of ONS-4 (configural, metric or scalar, across sex, age, occupation, language or time) was located in this pass. ONS publishes personal well-being estimates broken down by age and sex, but publishing subgroup means is not a test of invariance and does not establish that the items function equivalently across groups. Evidence that mode of administration shifts subjective well-being scores is directly relevant to invariance across data-collection contexts: telephone respondents report systematically higher well-being than face-to-face or self-completion respondents Dolan & Kavetsos 2016, a concern the ONS-4 validation study also flags because some of its ratings were collected face-to-face and some by telephone Benson 2019. This implies non-trivial risk of non-invariance across survey modes that has not been formally quantified for ONS-4.

Responsiveness and MIC Absent (searched; none found in the sweep to date)untestedevidence form: canonical

absence type: population-general searched; none found in the sweep to date

state: assessed_absent · rubric v1.6 · literature as of 2026-09-01 · grade confirmed 2026-09-03

Responsiveness and a minimal important change (MIC) have not been established for ONS-4. Responsiveness was named as a design goal for the derivative PWS, but the validation study was a cross-sectional secondary analysis with no linkage between pre- and post-intervention responses, so it could not estimate within-individual change or responsiveness Benson 2019. ONS-4-style items have been used as outcomes in a longitudinal natural experiment on the built environment, demonstrating that the items are used to detect change over time in practice, but that study did not derive an anchor-based or distribution-based MIC Ram 2020. No minimal-important-change threshold for any ONS-4 item was located.

Populations, languages and norms

ONS-4 was developed for and is fielded on the UK general adult population through the Annual Population Survey, and UK population norms and benchmarks are published by ONS from that survey and reported against standard threshold bands: for life satisfaction, worthwhile and happiness, responses of 0 to 4 are classed low, 5 to 6 medium, 7 to 8 high and 9 to 10 very high, while for anxiety (reverse direction) 0 to 1 is low anxiety, 2 to 3 medium, 4 to 5 high and 6 to 10 very high Benson 2019. Beyond the UK general population, instrument-specific validation exists only in one English social-prescribing sample using a modified four-point derivative Benson 2019, and applied UK studies have used the items in specific subpopulations such as residents of a regenerated urban neighbourhood Ram 2020 and during COVID-19 among home-workers Hensher & Beck 2023. Multi-instrument UK comparison datasets now situate ONS-4 alongside other measures for benchmarking Wickramasekera & Tsuchiya 2025. No validated non-English translation of the ONS-4 wording was located in this pass; the closely related OECD core module, which adds an affect item, provides an internationally harmonised counterpart Dolan & Metcalfe 2012.

Criticisms and controversies

The most consequential issue for a registry is the scale-versus-four-items question. ONS-4 was built as four separate single items measuring distinct constructs and ONS provides no summary score, so summing the items into a composite (as some deployments do) is a departure from the instrument's design and psychometric basis Benson 2019. A second issue is that the strongest instrument-specific psychometrics (alpha 0.90, factor structure) belong to a reworded four-point derivative rather than the original 0 to 10 items, so those coefficients should not be read as evidence about ONS-4 as ONS fields it Benson 2019. Third, the anxiety item is negatively worded and reverse-scored, its distribution is strongly skewed towards low anxiety, and it behaves differently from the three positive items, which complicates any attempt to combine the four Benson 2019. Fourth, mode-of-administration effects are documented: telephone interviewing inflates reported well-being relative to other modes, a threat to comparability across surveys and over time that is particularly relevant when the items are used for benchmarking Dolan & Kavetsos 2016. Finally, single-item measurement remains debated; while single-item life-satisfaction measures perform comparably to multi-item scales on validity and reliability in large studies Cheung & Lucas 2014 Lucas & Donnellan 2012, that literature concerns life satisfaction specifically and does not directly certify the worthwhile, happiness or anxiety items, and the causal standing of subjective well-being for downstream health outcomes is itself contested Wootton 2018.

References (14)

  1. Benson T, Sladen J, Liles A, Potts HWW (2019). Personal Wellbeing Score (PWS): a short version of ONS4: development and validation in social prescribing https://doi.org/10.1136/bmjoq-2018-000394
  2. Benson T, Sladen J, Liles A, Potts HWW (2019). Correction: Personal Wellbeing Score (PWS): a short version of ONS4: development and validation in social prescribing https://doi.org/10.1136/bmjoq-2018-000394corr1
  3. Dolan P, Metcalfe R (2012). Measuring Subjective Wellbeing: Recommendations on Measures for use by National Governments https://doi.org/10.1017/s0047279411000833
  4. Lucas RE, Donnellan MB (2012). Estimating the Reliability of Single-Item Life Satisfaction Measures: Results from Four National Panel Studies https://doi.org/10.1007/s11205-011-9783-z
  5. Cheung F, Lucas RE (2014). Assessing the validity of single-item life satisfaction measures: results from three large samples https://doi.org/10.1007/s11136-014-0726-4
  6. Dolan P, Kavetsos G (2016). Happy Talk: Mode of Administration Effects on Subjective Well-Being https://doi.org/10.1007/s10902-015-9642-8
  7. Ruggeri K, Garcia-Garzon E, Maguire A, Matz S, Huppert FA (2020). Well-being is more than happiness and life satisfaction: a multidimensional analysis of 21 countries https://doi.org/10.1186/s12955-020-01423-y
  8. Wickramasekera N, Tsuchiya A (2025). A Large Scale Population Survey of Health and Wellbeing to Allow Comparisons Between Outcome Measures: the SIPHER-HWMIC Dataset https://doi.org/10.1007/s11205-025-03728-1
  9. Koivumaa-Honkanen H, Honkanen R, Viinamaki H, Heikkila K, Kaprio J, Koskenvuo M (2000). Self-reported life satisfaction and 20-year mortality in healthy Finnish adults https://doi.org/10.1093/aje/152.10.983
  10. Gana K, Broc G, Saada Y, Amieva H, Quintard B (2016). Subjective wellbeing and longevity: Findings from a 22-year cohort study https://doi.org/10.1016/j.jpsychores.2016.04.004
  11. Wootton RE, Lawn RB, Millard LAC, Davies NM, Taylor AE, et al. (2018). Evaluation of the causal effects between subjective wellbeing and cardiometabolic health: mendelian randomisation study https://doi.org/10.1136/bmj.k3788
  12. Diener E, Chan MY (2011). Happy People Live Longer: Subjective Well-Being Contributes to Health and Longevity https://doi.org/10.1111/j.1758-0854.2010.01045.x
  13. Ram B, Limb ES, Shankar A, Nightingale CM, Rudnicka AR, et al. (2020). Evaluating the effect of change in the built environment on mental health and subjective well-being: a natural experiment https://doi.org/10.1136/jech-2019-213591
  14. Hensher DA, Beck MJ (2023). Exploring how worthwhile the things that you do in life are during COVID-19 and links to well-being and working from home https://doi.org/10.1016/j.tra.2022.103579

Record notes

[Upgraded from v0.1 to v0.2 structure in pass two; criterion field split, licence re-verified 2026-07-12.] Overall confidence is Low, with two properties (test-retest, invariance, responsiveness/MIC) graded Absent. The dominant honesty problem this record surfaced is a substitution risk that the schema does not explicitly guard against: almost all instrument-specific psychometric coefficients (alpha 0.90, factor structure, item-total r=0.83 to 0.88) come from Benson 2019, but that study validated a modified four-point derivative (the Personal Wellbeing Score) with reworded items, a collapsed response scale and a reversed anxiety direction, not the original 0 to 10 ONS-4. The maintainers have flagged this inline everywhere a Benson coefficient appears, but a naive reader could still mistake these for original-ONS-4 properties; the schema would benefit from a field distinguishing 'evidence for this instrument as canonically fielded' from 'evidence for a named derivative'. Second, the scale-versus-single-items status is central and cuts across every property: ONS-4 is four separate single items with no ONS summary score, which makes internal consistency, structural validity and a single test-retest coefficient partly category errors when applied to the instrument as designed; The maintainers recorded these honestly as Low or Absent rather than forcing scale-level statistics. Third, for reliability and convergent validity the maintainers had to rely on the general single-item life-satisfaction literature (Lucas & Donnellan; Cheung & Lucas) because ONS-4-specific studies do not report these; this is indirect evidence (general-population samples, life satisfaction only, not the worthwhile/happiness/anxiety items as fielded) and the maintainers downgraded accordingly. Fourth, key steward facts (UK norms from the Annual Population Survey, Open Government Licence, National Statistics designation) are documented in ONS technical guidance that does not carry a DOI; The maintainers grounded the citable elements (threshold bands, National Statistics status, ONS encouragement of reuse) in Benson 2019 and did not invent a DOI for ONS web guidance. No workplace criterion validity was located, which is a material gap given the registry's audience. All fourteen DOIs were checked for resolution and retraction status in the latest review pass. [0.2.1] ONS-4 item split implemented per schema v0.2 rule 4: four first-class item records (see items list) with graded evidence retained at set level; items cross-linked to the question bank.