In a conjoint experiment, respondents see hypothetical profiles, such as candidates, immigrants, vaccines or policies, whose attributes are randomly assigned. They then choose between the profiles in a task, rate each one, or both, usually over several tasks. Every profile is a new random bundle of attribute levels, so no fixed probe can serve as item. Conjoint data therefore get their own layout and form their own data family in the IRW (source = "conj"), as the competition and nominal data do.
A table belongs here when respondents evaluate profiles whose attributes were randomized: at least two attributes, each with varying levels. That covers paired and single-profile conjoints, and factorial surveys or vignette experiments that randomize the attributes of a text vignette. A fixed set of vignettes shown identically to everyone, such as anchoring vignettes, is a set of ordinary items and belongs in the core IRW.
The data standard
One row per respondent × task × profile.
column
meaning
id
required
Respondent identifier. Never a platform worker ID or anything else that identifies a person.
task
required
Task index within respondent, in the order shown.
profile
required
Position within the task (1 = left or first). Single-profile designs use profile = 1 throughout.
choice
outcome
1 if this profile was chosen in the task, else 0.
rating
outcome
A numeric judgement of this profile, stored as in the source, without rescaling or reversal.
choice_<name>, rating_<name>
outcome
Further questions asked about the same tasks, for example which candidate a respondent would vote for and which would most reduce corruption.
attr_<name>
required, ≥ 2
The level of each attribute as respondents saw it, stored as text rather than a numeric code. An attribute the design left off a profile is the text (not shown), never a blank cell.
attrpos_<name>
optional
Row position of the attribute in the profile, when attribute order was randomized.
cov_<name>
optional
Respondent covariates, as in the core standard. A survey weight is cov_survey_weight.
trial_<name>
optional
Other task-level details, such as an experiment arm or a framing condition. trial_repeat_of marks a task that repeats an earlier one (often the first task shown again at the end, to measure reliability) with the number of the task it repeats.
Choice or rating:
choice is a pick among the profiles of a task. A single-profile accept/reject question is also a choice, with an opt-out, since rejecting is the outside option.
Everything else asked about a profile is a rating, including yes/no judgements that do not pick among profiles, such as “is this candidate a Democrat?” asked of each profile.
Respondent covariates keep the names and codings of the source, except for these, which mean the same in every table:
column
values
cov_gender
female, male or other; missing when not answered
cov_age
age in years at the survey
cov_birth_year
year of birth
cov_age_group
an age band, as text (“18-29”)
cov_education
the main education question, as the answer text in the source’s own categories and language
cov_party_id
party identification, as answer text; a US 7-point scale is cov_party_id7
cov_attention_pass
1 = passed an attention check, 0 = failed (cov_attention_pass_1, _2, … when there are several)
cov_duration_sec
survey or module duration in seconds
cov_survey_weight
the per-respondent survey weight
Codes are turned into text only from the study’s own codebook. A covariate whose codes no source explains keeps them, with a _code suffix (cov_gender_code).
Further rules:
A forced choice has exactly one chosen profile per task. In a design that offered “neither”, an opt-out task has choice = 0 on every profile.
Rows with no outcome are left out. A row with a choice but a missing rating is kept.
(not shown) is the one reserved attribute value. It is used when a design hides some attributes, shows a subset in some arms, or leaves a clause out, and it reads the same in every table and language. Levels respondents actually saw keep their own text, even when they describe an absence, such as “None”. A blank attribute cell never appears.
One table holds one experiment: one attribute set and one design. Several questions about the same tasks stay in the same table. Country samples are kept together when the original authors analysed them together and the attribute text is shared, with the country in cov_country. Otherwise each sample is its own table.
Derived variables, such as dummy codings of attributes or “co-partisan” flags, are dropped, since they can be rebuilt from the attributes.
What each table means
The tables share one layout but not one meaning. choice is “vote for” in one table and “admit” in another, and ratings run 1–7, 0–10 or 0–100. For each table, the IRW records:
the question wording and scale anchors of every outcome;
whether an opt-out was offered;
any randomization restrictions: combinations of levels that were not allowed, kept separate from unequal level probabilities, since only the first requires estimating within the allowed combinations;
whether attribute order was fixed or randomized (per respondent or per task);
whether the study had a survey weight, and whether the table keeps it;
whether profiles were shown as an attribute table, a text vignette or an image;
the country, the language respondents saw, and the language of the stored attribute text;
whether task and profile were recorded in the source or inferred from row order.
These records are kept in data/conjoint/ in the IRW repository, alongside each table’s processing script, so that tables can be compared or pooled with their differences in view.
A few things are left to the user. Attribute levels are text even when they are numbers, such as a price or an age. No baseline level is recorded: estimates of the average marginal component effect (AMCE) need one, and the analyst chooses it, while marginal means do not.
Getting back to id, item, resp
Any conjoint table converts mechanically to the core layout. Each outcome becomes an item named after its question (choice, rating, …), its value becomes resp, and task, profile and attributes become trial_ columns. irw_conj_long() in R and irw.conj_long() in Python do this, so tools written for the core layout can be used, for example an explanatory item response model of the outcome on the attributes. The long view does not keep which profiles shared a task, so analyses of choice between profiles work from the conjoint layout. The reverse is not possible, which is why the conjoint layout is the stored form.
Accessing the conjoint data
The conjoint data are in the Redivis dataset irw_conjoint and can be accessed from both the R and Python irw packages. Tables can be listed, fetched, described, filtered and cited in both. The table metadata (irw_metadata(source = "conj")) carries the design of each experiment: the numbers of respondents, tasks, profiles and attributes, the outcomes, the country and language, and the randomization restrictions.
library(irw)# View available conjoint datairw_list_tables(source ="conj")# Fetch a conjoint tabledf <-irw_fetch("kreps_2020_covid_vaccine", source ="conj")head(df)# The core id/item/resp viewlong <-irw_conj_long(df)table(long$item)# Design metadata, and tables with a rating outcome fielded in the US or UKmeta <-irw_metadata(source ="conj")irw_filter(source ="conj", outcome ="rating", country =c("US", "GB"))# Citationsirw_save_bibtex("kreps_2020_covid_vaccine", source ="conj", output_file ="conjoint.bib")
Code
import irw# View available conjoint datairw.list_tables(source="conj")# Fetch a conjoint tabledf = irw.fetch("kreps_2020_covid_vaccine", source="conj")df.head()# The core id/item/resp viewlong= irw.conj_long(df)long["item"].value_counts()# Design metadata, and tables with a rating outcome fielded in the US or UKmeta = irw.metadata(source="conj")irw.filter(source="conj", outcome="rating", country=["US", "GB"])# Citationsirw.save_bibtex(["kreps_2020_covid_vaccine"], source="conj", output_file="conjoint.bib")
An example
Bansak, Hainmueller and Hangartner (2016) asked 18,021 voters in 15 European countries to compare pairs of hypothetical asylum seekers. The average marginal component effect (AMCE) of an attribute level is the change in the probability that a profile is chosen when the attribute takes that level rather than a baseline, averaged over the other attributes. With the paper’s baselines and its survey weights (top-coded at 6), a linear probability model on the IRW table gives the same AMCEs the paper reports: +0.13 for a doctor and +0.09 for a teacher (against unemployed), −0.11 for a Muslim (against Christian), and −0.15 for an economic migrant (against political persecution).
Code
d <-irw_fetch("bansak_2016_asylum", source ="conj")d$attr_occupation <-relevel(factor(d$attr_occupation), "Unemployed")d$attr_religion <-relevel(factor(d$attr_religion), "Christian")d$attr_reason <-relevel(factor(d$attr_reason), "Persecution for political views")fit <-lm(choice ~ attr_occupation + attr_religion + attr_reason, data = d,weights =pmin(cov_survey_weight, 6))coef(fit)
The standard errors should be clustered by respondent, for example with sandwich::vcovCL(fit, cluster = ~id), because each respondent made several choices.