How the instrument registry grades evidence
The instrument registry is the open synthesis of the published evidence on instruments used to measure workplace health and wellbeing. For every instrument it records what it measures, how well, in which populations and languages, and on what licence terms, drawn from the literature and existing systematic reviews, every claim cited, every grade conservative. It is maintained, machine-readable and free to use, so that nobody choosing, building, licensing or reviewing a workplace measure has to reassemble the field's evidence themselves. This page explains how to read it, in plain language, and what the grades can and cannot tell you.
The eight properties, in one sentence each
| Property | The question it answers |
|---|---|
| Structural validity | Does the instrument have the internal shape its authors claim, for example one thing being measured rather than several tangled together? |
| Convergent validity | Does it agree with other measures it should agree with, and differ from ones it should not? |
| Criterion validity (reference standard) | Does it agree with a trusted benchmark, such as a structured clinical interview? |
| Criterion validity (organisational outcomes) | Does it predict things organisations care about, such as sickness absence or turnover? |
| Internal consistency | Do the items hang together as one scale? |
| Test-retest reliability | Does it give a stable answer when nothing has changed? |
| Measurement invariance | Does it mean the same thing across groups (sex, age, occupation, language) and over time? |
| Responsiveness | Does it detect genuine change, and how much change is meaningful? |
Each record also carries a ninth graded property, populations, languages and norms, which summarises the breadth and quality of the population evidence rather than a psychometric property in the strict sense. It sits on the record page, not in the matrix, and its grade should be read as a rough summary with the findings.
What the grades mean
High, Moderate, Low and Very low summarise the quality and consistency of the published evidence for that property, informed by the COSMIN approach to evaluating measurement studies. Absent means the registry searched for evidence on that property and found none in the sweep to date. It is a finding about the literature, not a grade of the instrument; from dataset v0.5.0 every Absent or Not-applicable cell also names its absence type (nothing found in any population, or a property that does not apply to the construct; evidence that exists only in other populations is graded and flagged indirect, never recorded as Absent), and from v0.3.0 every cell carries an evidence state that keeps "searched and found nothing" (assessed_absent) separate from "not yet searched" (not_assessed). The full scale, the downgrade rules and how a cell is derived are in the grading rubric, version 1.6. A tilde (~) marks a thin evidence base; an asterisk (*) marks evidence that is actively contested in the literature.
Grades are not comparable across instruments. A wellbeing index and a burnout inventory measure different things; their grades summarise different literatures answering different questions. The registry deliberately publishes no ranking and no "best instrument". The registry tells you what the evidence says about each instrument on its own terms, and it does not choose for you.
The honest current limits
- The grades are single-rater, and frozen. Every grade was assigned by one human rater, Zak Fenton, who created OWHS and is employed by the steward, applying the rules published as rubric v1.6 (version 1.0 was written after the grades, codifying the practice as applied; version 1.1, before first publication, redefined indirectness by population type rather than country, and every cell was re-read against it, logged as correction C-0004; version 1.2, also before first publication, followed a review pass of that re-read: the flag became three-valued with its basis stated on every cell, a High grade now needs two cited studies with a sample size and a statistic, the six C-0004 grade moves were reversed and six further High grades dropped to Moderate, logged as correction C-0005 with the earlier values kept on each cell; version 1.3, again before first publication, followed a second review pass: the evidence-form rule recodes rather than caps, the High precondition asks for a statistic of the property graded and, on population-sensitive properties, at least one working-adult or general-population study, criterion validity against a reference standard became population-sensitive, one Moderate returned to High and five High grades dropped to Moderate, logged as correction C-0006; version 1.4, after a third review pass that found no grade demonstrably wrong, names the statistics of each property that the High precondition counts and tests them by machine, keeps the evidence a cell dropped from High with in its history, states the internal-consistency rule for item sets, and maps every rule to its check: one High grade dropped to Moderate on a credible published disagreement, logged as correction C-0007; version 1.5, after a fourth review pass returned the first fit-to-publish verdict, requires that a High grade on populations, languages and norms rest on at least one study carrying a norm, cut-off, prevalence or reference value, splits criterion validity so that only an accuracy statistic counts against a reference standard, and reconciles the rubric's statistic lists with the machine check in both directions: one Moderate grade dropped to Low where the record had no reliability coefficient of its own, logged as correction C-0008; version 1.6, after a fifth review pass scoped to the changes of 1.5 returned fit to publish with one grade to move, reads the norm rule as written: one High grade on populations, languages and norms dropped to Moderate where the word that met the rule named a shared-variance figure and not a norm, one status went from well-established to thin, one judgement clause was corrected and one rater note added, logged as correction C-0009). Each of those passes was a review pass over the dataset by the same rater, assisted by AI in retrieval, screening and drafting. None was independent: no rater who is not an employee of the steward has yet seen the registry. Grades and statuses are frozen from first publication and stay frozen until two named psychometric raters who are not employees of the steward have joined. Until then, treat the citations as the registry's strongest content and the grades as its structured, disclosed, but not yet independently checked reading of them. If you could be one of those raters, the contribute page says what is involved.
- AI assistance and human accountability. AI assisted retrieval, screening, extraction, synthesis and checks against written rules. Earlier revisions also used AI-derived population flags and scripts to apply rule consequences, as disclosed in rubric section 10. The published grades were assigned or confirmed by the named rater and have not yet been independently reproduced by psychometric raters. Automated sweeps may propose citations and review flags; they may not change grades, statuses, evidence forms, population judgements or licence classifications. What the automation may and may not do is set out on the maintenance page.
- Synthesis, not systematic review. A synthesis here means reading across the published studies and reviews and stating what they show, including where they disagree. The registry does not run its own meta-analyses and its searches are not conducted to systematic-review standard. Where a COSMIN review exists for an instrument, the registry cites and summarises it; where none exists, the registry says so and grades what is there. The queries below describe a search design, not a reproducible log of every cell's search or proof of a running database workflow.
- Search coverage is incomplete. Searches may miss relevant work because of indexing, terminology, language, access restrictions or unavailable manuals and technical reports. The September 2026 run used web search under database-access restrictions; it did not query the major indexes directly.
- This is an initial set of 27 instruments. It includes measures represented in the OWHS question bank and is a bounded selection of the field, not a survey of UK employer usage or a census of measurement evidence. In this set, organisational criterion validity is recorded as Absent for 15 instruments and graded Low or Very low for a further 7. Grade and evidence status are distinct: 8 organisational cells carry the thin-status marker, including TIS-6, which is graded Moderate. Further instruments enter through the public admission route.
- Two organisational cells are listed for re-examination. The organisational criterion validity cells for the Work Ability Score and the Turnover Intention Scale (TIS-6) rest on recorded outcomes (register-based disability pension and absence days; actual leavers against stayers), so they stand under the stated outcome definition. Both were flagged in external review as cells to re-examine when the freeze lifts, and they will be, in public, under the rater conditions in force then.
- Licence terms change without notice. Each record states a licence class for employer or vendor use (open, free without permission, free for non-commercial use, free with registration, fee-bearing, no formal licence, or unverified), the date it was verified, and links to archived copies of the steward's page, so the basis of every licence claim is independently checkable. The registry records the class and the source, never a price: prices change and the registry would be wrong within the year.
Search design and its present limits
The September 2026 sweep used web search under direct-database access restrictions. The planned deterministic harvester will use instrument names and aliases, measurement-property terms and citation links; its exact queries, sources, search windows and retrieval logs will be published with the implementation. The search strings written for the earlier passes remain in the repository history; their historical hit counts are not a measure of retrieval completeness and not evidence of a live monthly workflow.
Earlier literature passes were non-systematic and were not logged per cell. The as-of dates on cells describe those passes. A new citation or a completed search is distinct from human confirmation of a grade.
Evidence reports and accepted corrections are recorded in the changelog and the sweep reports. Automated retrieval and the end-to-end update pipeline are being built; the maintenance page distinguishes what has run from what is planned. Grades remain frozen; a search result may prompt review but does not confirm or change a grade.