How the instrument registry grades evidence

methods · updated 5 September 2026 · registry v0.9.0 · rubric v1.6

The instrument registry is the open synthesis of the published evidence on instruments used to measure workplace health and wellbeing. For every instrument it records what it measures, how well, in which populations and languages, and on what licence terms, drawn from the literature and existing systematic reviews, every claim cited, every grade conservative. It is maintained, machine-readable and free to use, so that nobody choosing, building, licensing or reviewing a workplace measure has to reassemble the field's evidence themselves. This page explains how to read it, in plain language, and what the grades can and cannot tell you.

The eight properties, in one sentence each

PropertyThe question it answers
Structural validityDoes the instrument have the internal shape its authors claim, for example one thing being measured rather than several tangled together?
Convergent validityDoes it agree with other measures it should agree with, and differ from ones it should not?
Criterion validity (reference standard)Does it agree with a trusted benchmark, such as a structured clinical interview?
Criterion validity (organisational outcomes)Does it predict things organisations care about, such as sickness absence or turnover?
Internal consistencyDo the items hang together as one scale?
Test-retest reliabilityDoes it give a stable answer when nothing has changed?
Measurement invarianceDoes it mean the same thing across groups (sex, age, occupation, language) and over time?
ResponsivenessDoes it detect genuine change, and how much change is meaningful?

Each record also carries a ninth graded property, populations, languages and norms, which summarises the breadth and quality of the population evidence rather than a psychometric property in the strict sense. It sits on the record page, not in the matrix, and its grade should be read as a rough summary with the findings.

What the grades mean

High, Moderate, Low and Very low summarise the quality and consistency of the published evidence for that property, informed by the COSMIN approach to evaluating measurement studies. Absent means the registry searched for evidence on that property and found none in the sweep to date. It is a finding about the literature, not a grade of the instrument; from dataset v0.5.0 every Absent or Not-applicable cell also names its absence type (nothing found in any population, or a property that does not apply to the construct; evidence that exists only in other populations is graded and flagged indirect, never recorded as Absent), and from v0.3.0 every cell carries an evidence state that keeps "searched and found nothing" (assessed_absent) separate from "not yet searched" (not_assessed). The full scale, the downgrade rules and how a cell is derived are in the grading rubric, version 1.6. A tilde (~) marks a thin evidence base; an asterisk (*) marks evidence that is actively contested in the literature.

Grades are not comparable across instruments. A wellbeing index and a burnout inventory measure different things; their grades summarise different literatures answering different questions. The registry deliberately publishes no ranking and no "best instrument". The registry tells you what the evidence says about each instrument on its own terms, and it does not choose for you.

The honest current limits

Search design and its present limits

The September 2026 sweep used web search under direct-database access restrictions. The planned deterministic harvester will use instrument names and aliases, measurement-property terms and citation links; its exact queries, sources, search windows and retrieval logs will be published with the implementation. The search strings written for the earlier passes remain in the repository history; their historical hit counts are not a measure of retrieval completeness and not evidence of a live monthly workflow.

Earlier literature passes were non-systematic and were not logged per cell. The as-of dates on cells describe those passes. A new citation or a completed search is distinct from human confirmation of a grade.

Evidence reports and accepted corrections are recorded in the changelog and the sweep reports. Automated retrieval and the end-to-end update pipeline are being built; the maintenance page distinguishes what has run from what is planned. Grades remain frozen; a search result may prompt review but does not confirm or change a grade.