How to read this registry

Every convention in the registry exists to stop a reader over-trusting a number. This page is the method.

The instrument registry is the open synthesis of the published evidence on instruments used to measure workplace health and wellbeing. For every instrument it records what it measures, how well, in which populations and languages, and on what licence terms, drawn from the literature and existing systematic reviews, every claim cited, every grade conservative. It is maintained, machine-readable and free to use, so that nobody choosing, building, licensing or reviewing a workplace measure has to reassemble the field's evidence themselves.

Synthesis, not systematic review. A synthesis here means reading across the published studies and reviews and stating what they show, including where they disagree. The registry does not run its own meta-analyses and its searches are not conducted to systematic-review standard. Where a COSMIN review exists for an instrument, the registry cites and summarises it; where none exists, the registry says so and grades what is there. The search venues and strings the sweep runs are published, so the coverage can be checked and improved by anyone.

The grade scale

Each property carries one of High Moderate Low Very low, or one of two non-grades: Absent, meaning no published evidence was located, which is reported as a finding about the literature; and Not applicable, meaning the property is a category error for the instrument's type (a single item has no internal consistency to report). The two non-grades are deliberately distinct: absence is information, not-applicable is taxonomy, and a registry that lets them blur misleads exactly the reader it exists to protect. Every ungraded cell also names its absence type: population-general (searched, nothing found in any population) or category-error (the property does not apply to the construct, for example a criterion reference standard for a construct that has none, or measurement invariance of a single item). Evidence that exists only in other populations is never recorded as Absent: it is graded and flagged indirect. Ungraded cells carry no indirectness flag: there is no grade for the flag to qualify.

Evidence states: what the registry did, not what the literature says

Every cell also records the registry's own work on it. assessed: searched and graded. assessed_absent: searched, nothing found in the sweep to date, published as a finding with the search basis. not_assessed: not yet searched, which says nothing about the literature; new instruments enter the watchlist in this state until a rater grades them. not_applicable: category error. Each cell also carries the rubric version it was graded under, the date its literature was last searched, the date a human last confirmed the grade, and a review-due flag set whenever a sweep adds a citation after that confirmation. The full method is the grading rubric.

Who graded this, and the freeze

Grades and statuses are single-rater and are frozen from first publication until two named independent raters have joined. Corrections of fact still run; grade moves do not. The rubric 1.1 re-rate is logged as correction C-0004 and the rubric 1.2 reconciliation of the independent check, which reversed the C-0004 grade moves, as C-0005. C-0006 (2 September 2026) reconciles the second independent check of v0.5.0 under rubric 1.3: every change is a table in migrate_v0.6.py and is listed in the correction's note. C-0007 (2 September 2026) reconciles the third independent check of v0.6.0 under rubric 1.4: every change is a table in migrate_v0.7.py and is listed in the correction's note. C-0008 (3 September 2026) reconciles the fourth independent check of v0.7.0 under rubric 1.5: every change is a table in migrate_v0.8.py and is listed in the correction's note. C-0009 (3 September 2026) reconciles the fifth independent check of v0.8.0 under rubric 1.6: every change is a table in migrate_v0.9.py and is listed in the correction's note. Unfreeze condition: Two named psychometric raters who are not employees of the steward have joined the registry and the inter-rater procedure in the rubric is in force. The registry is looking for those raters: psychometricians or measurement researchers with no financial interest in the instruments they would grade (instrument authors are welcome as reviewers of their own records, never as raters of them). Write to hello@openworkplacehealth.org.

The four status tokens

The grade says how strong the evidence is; the status says what kind of literature produced it: well-established a mature, replicated evidence base; contested credible published disagreement (this badge is displayed prominently on record pages, never buried); thin few studies, small samples, or narrow settings; untested the specific claim has not been directly examined. Moderate-and-contested and Moderate-and-thin are different situations that previously shared a word; here they never do.

Indirectness: who the grade is for

Grades are for a working-adult audience, not bibliometric totals. Every grade carries a three-value indirectness flag with its reason: direct evidence earned on samples of working adults, in any country and any language; general evidence earned on adult general-population samples that include working adults without isolating them; indirect evidence earned in clinical, student, adolescent, older-adult or otherwise non-working samples. The flag and the grade are separate facts: a grade never moves because a flag moved. The flag bites only where the property is sensitive to who was sampled (convergent validity, both criterion validities, responsiveness, populations and languages): there a High grade needs direct or general evidence, and indirect evidence caps the grade at Moderate. Criterion validity against a reference standard is on that list because a cut-off validated in a clinic, where the condition is common, is not a cut-off validated in a workforce, where it is rare. For the properties that are about the instrument's own structure (internal consistency, structural validity, invariance, test-retest) the flag is recorded and shown but does not move the grade. Each flag names the samples it rests on, quoted from the cell's own findings, each naming a population (a sample size or a country on its own is not a reason), or says that it falls back to the population the instrument is fielded on because the cited studies describe no sample. Country and language are never a reason for a downgrade: where the evidence was earned is recorded on the cell as a fact, and whether an instrument behaves the same across languages is a graded property of its own, measurement invariance. A large clinical literature does not entitle an instrument to a High grade for workplace use; the flag is where that discipline lives. A High grade has a precondition: at least two cited studies with a sample size and a statistic of the property graded, listed on the cell as its High basis, and on a population-sensitive property at least one of them in a working-adult or general-population sample, and on populations, languages and norms at least one of them carrying a norm, cut-off, prevalence or reference value; a cell that cannot show them is Moderate at most. The numbers the grade rests on, with their intervals and sample sizes, are in the findings on each cell. A record-level deployment context caveat, where present, states once anything that cuts across every property (for example a clinical-origin instrument deployed in a workplace) and is shown at the top of the evidence section.

Relations between records

Where a published study links two records, the link is recorded as data and shown on both record pages: an item to its parent set, a single item or short form to the full instrument it was derived from, a single item to a multi-item instrument it has been shown to screen for (with reported sensitivity and specificity against the longer instrument's threshold), a single item to a multi-item instrument it corresponds with (a reported correlation only, which is convergence, not screening), and a composite survey to an instrument it fields verbatim. Every relation cites the study it rests on; a relation without one is not recorded. A relation is a fact about the literature, never a recommendation: the registry does not say which form an implementer should use, and it never states a threshold for action. The single-item measures are gathered on their own page.

Evidence-form provenance

Every grade names what the evidence was earned on: canonical (the fielded instrument, as versioned), derivative (a named reworded or short form), parent (a longer parent form), or mixed. Borrowed evidence can no longer be silently read as evidence for the fielded instrument.

Criterion validity is two properties, never one

Validity against a diagnostic or health reference standard and validity against organisational outcomes (absence, turnover, performance) diverge so sharply across this registry that a merged grade actively misleads, in exactly the direction that harms a workplace reader. They are graded separately everywhere, by schema rule.

Criterion validity against organisational outcomes

This property records whether an instrument's scores have been shown to relate to outcomes an organisation records: sickness absence, return to work, occupational health referral, enacted adjustments, benefit use, actual turnover, rated or objective performance, and safety incidents. In the current set of 27 instruments, organisational criterion validity is recorded as Absent for 15.

What counts as an outcome here. An organisational outcome is an event the organisation recorded, an outcome linked from a register, or a performance measure rated independently of the person completing the instrument. Self-reported constructs do not count in this property, however work-related they are: an association between an instrument and self-reported turnover intention, self-rated productivity loss or self-rated work capacity is convergent validity between two self-reports, and it is graded as convergent validity. Keeping the two apart is what stops this property from measuring itself.

What the evidence records. Each located finding names the outcome, its source, the study design, any follow-up interval, the effect and its direction, the sample and population, and the form of the instrument the evidence was earned on. In the current dataset version those findings are prose on the record page; from the next dataset version they are structured fields, so that a reader can filter the registry on which outcomes have located evidence for an instrument.

What this property does not say. It does not say the instrument will predict outcomes in your organisation: the evidence was earned in named populations and settings, which every finding states. It does not rank instruments, and the registry publishes no ordering of instruments on this or any other property. It says nothing about whether measuring something changes it: evidence about interventions is a different literature, and this registry does not cover it.

Test-retest is structured, not prose

Retest findings are recorded as structured entries (coefficient, coefficient type, interval, sample, population), so bundled ICCs and internal-consistency contamination cannot pass as stability evidence. In the current set of 27 instruments, test-retest reliability has 0 High grades and 11 Absent cells.

Licence currency

Licence status is verified against the steward's current distribution terms, with the verification date shown on every record. Founding papers and review literature are never acceptable licence sources; that rule was learned publicly (correction C-0001) and is now schema law. Where a steward's terms could not yet be verified, the record says so plainly and displays no settled status.

Licence classes

Each record carries a licence class for one audience, an employer or a vendor acting for one: open: An open licence that permits commercial and non-commercial use without permission or fee (Open Government Licence, CC BY, Apache and equivalents).; free-no-permission: The steward states the instrument is free to use and requires no permission, but publishes no formal open licence.; free-non-commercial: Free for research or non-commercial use; employer or vendor deployment requires the steward's permission and may be charged.; free-with-registration: Free for the registry's audience but only after registration or application with the steward.; fee-bearing: A paid licence is required for the registry's audience (employer or vendor workplace use), whatever the terms for other users.; no-formal-licence: No steward and no instrument-level licence exist (typically a generic single item published in a journal article).; unverified: The steward's current terms could not be read at the last verification; no class is asserted.. The registry records the class, the steward's source page and an archived copy of it, never a price.

The corrections policy, errata and right of reply

Disputed or wrong statements in published records are checked and fixed publicly: every correction carries the old value, the new value, and the source that settled it, on the corrections page. A registry that corrected itself silently would be indistinguishable from one that was never wrong; the log is the difference. An instrument's steward, developer or copyright holder has a right of reply: a written response is published beside the record, dated and unedited, with the registry's reply.

What the automation does and does not do

AI systems assist the rater with retrieval, screening, extraction and drafting, and run the monthly sweeps. They do not assign grades, statuses, evidence forms or indirectness flags, and nothing they produce reaches the dataset except through a change a human reviews and merges. What the automation may do, and where it has failed, is on the automation page.

What this registry is not

It is not part of the normative OWHS specification, it does not recommend instruments, it does not advise what to measure, and it never reproduces instrument item text. It grades evidence about instruments, never evidence about interventions. It describes the published evidence so that whoever chooses can choose with their eyes open. How an instrument comes to be here at all is a separate rule: how an instrument enters this registry.