The IRW Data Standard
Version 1.0 (beta) · 2026-09-19
Data standards are an essential mechanism of increasing reusability and interoperability of data amongst researchers. As such, they play a critical role in increasing the extent to which science is FAIR (findable, accessible, interoperable, reusable). The IRW data standard is meant to be a straightforward way of standardizing item response data such that it can be readily used by interested researchers in a wide variety of psychometric analyses.
You can use the standard without depositing anything with the IRW. To check a file against it, drop it on the browser validator (your file stays on your computer), or run pip install irw-validate and then irw-validate mydata.csv. Both report every problem against the numbered clauses below.
Cite as: The IRW Data Standard, version 1.0 (beta), 2026. https://itemresponsewarehouse.org/standard.html. Described in: Domingue, B. W., Braginsky, M., Caffrey-Maffei, L., Gilbert, J. B., Kanopka, K., Kapoor, R., Lee, H., Liu, Y., Nadela, S., Pan, G., Zhang, L., Zhang, S., & Frank, M. C. (2025). An introduction to the Item Response Warehouse (IRW): A resource for enhancing data usage in psychometrics. Behavior Research Methods, 57(10), 276. https://doi.org/10.3758/s13428-025-02796-y
How to read this page
The standard is a set of numbered clauses. MUST marks a requirement: a file that breaks one does not conform. SHOULD marks a strong expectation that a validator will warn about but that does not by itself make a file nonconforming. The clause numbers are stable; the validator names them in its reports ([C4]), and a later version will add clauses rather than renumber them.
Tables published before version 1.0 may not conform. The IRW predates this numbered standard, and some of its tables break a clause, most often C4 (a text NA in resp) or C5 (repeated responses with no column recording the occasion). They are being corrected, not exempted: released fixes are listed on the Corrections page, and known problems are tracked as issues. A table that does not conform is still usable, but check it before relying on it.
The standard defines the format of a response table. What the IRW itself chooses to host (licences, a minimum sample size, table names) is a separate matter, described under IRW intake policy at the end.
C1. Long format, with three required columns
A file MUST be in long format, one row per response, and MUST contain the columns id, item and resp, described in C2–C4.
C2. id
id is a persistent identifier for the focal unit, i.e., the thing being measured. This will typically be a person but in some cases may be, for example, a word (when interest is in establishing some property of the word). id MUST NOT be entirely missing, and SHOULD have no missing values.
C3. item
item is a persistent identifier for the probe being used to measure, and it is always authoritative: it MUST identify the probe, including in repeated-trial designs, where details of each trial go in trial_ columns rather than replacing item. The identifier SHOULD be kept in a form that makes subsequent matching to item text as straightforward as possible (e.g., identifiers that make it challenging to match back to original items, when available, should be avoided). item MUST NOT be entirely missing, and SHOULD have no missing values.
C4. resp
resp holds the item responses. Every value that is present MUST be numeric, so that it can be directly utilized in various psychometric models. Given its centrality, we make a few additional points:
A code for a missing or non-substantive answer (e.g.,
-9,99, “don’t know”) is not a response. Rows without a response SHOULD be omitted rather than carried with an emptyresp, and a text token such asNAis not a number.Exception: omitted versus not reached. Where the source documents separate codes for an item the respondent saw and skipped (omitted) and one they never reached (as IEA studies such as TIMSS and PIRLS do), an omitted item is a response: it SHOULD be scored
resp = 0, with the source’s code kept inresp_raw. Not-reached rows are not responses and SHOULD be omitted. Where the source does not distinguish the two, the rule above applies.Response values that were imputed in the original data SHOULD be removed.
Within an item, values are meant to be consistently coded so that higher numbers indicate a consistently meaningful change in a response (i.e., higher numbers indicate stronger agreement). However, values may load on the latent variable in opposite ways across items (i.e., higher numbers may indicate strong agreement for some items and weaker agreement for others).
While responses are most commonly discrete ordinal values (e.g., 1–5 Likert), bounded continuous responses are also permitted. Examples include visual-analogue or slider scales (e.g., 0–100 “does not describe me at all” to “describes me perfectly”). In such cases
respis stored as a float rather than an integer.
C5. One response per id and item per occasion
A given id–item pair MUST appear at most once, unless the rows are distinguished by a column that records the occasion of the response: wave, date, rater, a trial or session index (timepoint, trial, trialnum, order, session, occasion, period, block, subtest), or a trial_ column such as trial_number. A person rated by two raters, or answering the same item in two waves, is the design; the same response recorded twice is a defect.
C6. Other columns are named as defined here
Any column beyond id, item and resp SHOULD be one of the optional elements below, or one of the occasion columns in C5, or carry one of the prefixes cov_, itemcov_, qmatrix or trial_. A covariate recorded under a bare name (age rather than cov_age) cannot be told apart from a column with a defined meaning.
C7. Column order
The first three columns SHOULD be id, item, resp, in that order, followed by any optional columns.
Optional elements
These elements are optional features of IRW datasets. When IRW datasets have these columns, they have been consistently formatted as per the data standards described here.
resp_rawThe response as it was originally recorded, retained alongsiderespin cases where scoring discards information a user may want back. The typical example is a multiple-choice item:respis the scored 0/1, whileresp_rawholds the option the respondent actually selected (e.g.AthroughE), including whatever codes the original source used for an omitted or unreadable response.respremains the numeric column intended for psychometric modelling;resp_rawis not numeric and is not a substitute for it. Note the order of the two parts of the name: the column isresp_raw, notraw_resp.rtThe response time used to produce an item response is coded in seconds.dateThe calendar time at which a response was produced is included. Coding of this variable is done in two ways. In some cases, there were only relative dates (e.g., 30 days into data collection). In that case, we convert to the number of seconds since the first piece of data in the dataset was collected. In cases where more exact information was given (e.g., 1:30PM 03/04/2008), we convert to Unix time (i.e., seconds since JAN 01 1970. (UTC)).qmatrixWhen items are classified into a small number of skills (i.e., for the purposes of cognitive diagnostic modeling ), we have included these item-level classifications. Note that the column headers will beqmatrix1, …,qmatrixNwhen there areNskills.raterIn scenarios wherein the focal unit is being rated by other observers (perhaps along multiple dimensions), an identifier associated with the rater producing the rating.waveIn settings wherein data is collected longitudinally (but precise timing as indateis not available), we indicate the wave of data collection. Values are ordered such that larger values represent data collected at a later date relative to smaller values. Note that this is also used to indicate pre/post treatment in the case of data collected from randomized controlled trials (RCTs).treatAn indicator of whether a respondent was in a treatment group (1=Yes) in an experimental study (e.g., an RCT or quasiexperiment).cov_Covariates that are invariant for the foci of measurement (denoted byid; as an example, common covariates for persons may be gender and age) will be identified via thecov_prefix. NOTE: This was implemented as of V11.25; data added before that may not be reverse-compatible with this standardization element.itemcov_Covariates that are invariant for the measurement probe. NOTE: This was implemented as of V16.1; data added before that may not be reverse-compatible with this standardization element.trial_Details of individual trials in repeated-trial designs, such as the trial’s position (trial_number), its block (trial_block) or its list (trial_list). These sit alongsideitem, which still identifies the probe; atrial_index is what distinguishes repeated responses by oneidto oneitem(C5).item_familyAn identifier for groups of items that potentially have family resemblances such that assumptions related to local independence may be violated. Examples include items with a testlet structure, items that have common features (e.g., ‘how important is it that people wash their hands regularly?’ versus ‘how important is it that you wash your hands regularly’), or items that are clones of each other. Each element of a family of items will have a unique identifier (e.g., items in the first testlet have id1, items in the second testlet have id2, etc.) and items that are not a member of a family will beNA. NOTE: This was implemented as of V11.25; data added before that may not be reverse-compatible with this standardization element.
Additional considerations
In some cases, we deviate from the above rules for specific reasons. We describe those here. Some of the below may be considered more experimental features of the IRW; more information about these data will be forthcoming.
Process data: In general, a single row of data in the IRW corresponds to a unique response from an individual to an item. In the context of ‘process data’, more information about a response is available. In such cases, responses have identifiers that can be used to understand the sequence of actions that led to a given response.
Trials: Many constructs in cognitive psychology are measured via repeated trials of similar tasks wherein the probes either do not vary or vary in terms of some quantifiable feature of the stimulus. In such cases
itemstill identifies the probe (C3), and columns beginningtrial_record the trial itself, such as its position and block. Earlier versions of this page saiditemcould be uninformative in trial data; that reading is withdrawn.
The competition and nominal standards are experimental extensions and are not part of version 1.0.
IRW intake policy (not part of the standard)
A file can conform to the standard and still not be something the IRW hosts. These are the IRW’s own rules for what it accepts, and the validator reports them separately from conformance:
- Licence. The data must normally carry an explicit open licence (CC0, CC BY or CC BY-SA), or come with the owner’s written permission.
- Sample size. At least 100 unique
idvalues. - One file per scale. Responses to different instruments go in separate files.
- Table names.
author_year_construct, lowercase, at most 40 characters.
See Contributing data to the IRW for how to send a file.
Versions
| Version | Date | Change |
|---|---|---|
| 1.0 (beta) | 2026-09-19 | First numbered version, released as a beta. Clauses C1–C7 state the rules this page previously gave as prose, and the rules irw-validate already enforced: C5 (repeated responses) and C7 (column order) were enforced but not written here. Intake policy separated from the format; tables published earlier may not conform. item is always authoritative (C3); the earlier description of trial data, in which item could be uninformative, is withdrawn, and trial_ columns are defined as trial-level details. |