Skip to content
GitHub

Codebook

Every column in the Peekbank schema. See Data Schema for how the tables fit together.

Fields are required unless marked optional, and a → points at the table a foreign key refers to.

dataset_id key
integer
row identifier for the datasets table indexing from zero
lab_dataset_id
string · optional
lab-internal short name of the dataset
dataset_name
string
a descriptive label, could be a short citation. Needs to be unique in the database
shortcite
string · optional
shortened citation (author (year) APA-style in-text citation)
cite
string · optional
full citation (APA format)
dataset_aux_data
json · optional
optional additional information about the dataset in the format of a JSON (see aux_data tab for information about existing columns)
subject_id key
integer
database unique participant identifier (indexing from zero)
sex
string · optional
sex of the participants (male/ female/ other/ unspecified) One of female , male , other , unspecified .
native_language
string · optional
participants' native language (in ISO 6390-2); potentially a string of multiple language codes separated by a ", " (in the case of bilinguals; note that there *must* be a space after the comma)
lab_subject_id
string · optional
lab unique participant identifer
subject_aux_data
json · optional
optional additional information about the subject in the format of a JSON (see aux_data tab for information about existing columns)
administration_id key
integer
row identifier for the administrations table indexing from zero
dataset_id → datasets
integer
row identifier for datasets indexing from zero (linked to datasets)
age
float · optional
age of the participant in months (conversion from days to months: days / (365.25/12)); conversion from decimal years to months: years * 12; conversion from whole number years to months: years * 12 + 6)
lab_age
float · optional
lab-internal age recorded for subject (before converting to standardized age in months; conversion from days to months: days / (365.25/12))
lab_age_units
string · optional
unit used for lab-internal age recorded for subject (before converting to standardized age in months); should typically be days, months, or years One of days , months , years , epochs .
subject_id → subjects
integer
row identifier for subjects indexing from zero (linked to subjects)
monitor_size_x
integer · optional
width of the monitor (in pixels)
monitor_size_y
integer · optional
height of the monitor (in pixels)
sample_rate
float · optional
sampling rate of the eyetracker (in Hz) / frame rate for manual gaze coding
tracker
string · optional
name of the eye-tracker type/ data source (SMI/ iCoder/ Tobii/ etc.)
coding_method
string · optional
method used in the experiment for coding gaze data (eyetracking, manual gaze coding, automated gaze coding, preprocessed eyetracking - meaning eyetracking with only aoi data available) One of eyetracking , manual gaze coding , automated gaze coding , preprocessed eyetracking .
administration_aux_data
json · optional
optional additional information about the administration in the format of a JSON (see aux_data tab for information about existing columns)
trial_id key
integer
row identifier for the trials table indexing from zero
trial_order
integer · optional
index of the trial in order of presentation for this participant during the experiment
trial_type_id → trial_types
integer
row identifier for the trial_types table indexing from zero
excluded
boolean · optional
whether or not the trial was excluded in the data analysis conducted in the original experiment for which the data was collected
exclusion_reason
string · optional
the reason the trial was excluded in the original experiment that is the source of the data
trial_aux_data
json · optional
optional additional information about the trial in the format of a JSON (see aux_data tab for information about existing columns)
trial_type_id key
integer
row identifier for the trial_types table indexing from zero
aoi_region_set_id → aoi_region_sets
integer · optional
unique aoi_region identifier (indexing from zero)
dataset_id → datasets
integer
row identifier for the datasets table indexing from zero
distractor_id → stimuli
integer
row identifier for the distractor stimulus indexing from zero
target_id → stimuli
integer
row identifier for the target stimulus indexing from zero
full_phrase
string · optional
Full phrase prompting the target
full_phrase_language
string · optional
language of the full phrase (options: ISO 639-2/B language 3 digits code, multiple, artificial); see here for table and information on ISO 639-2/B language 3 digits code (use code from 639-2/B column): https://en.wikipedia.org/wiki/List_of_ISO_639-1_codes ; other link: https://gist.github.com/gantian127/8345007938faf611fa6d )
489 accepted values

aar, abk, ace, ach, ada, ady, afa, afh, afr, ain, aka, akk, alb, ale, alg, alt, amh, ang, anp, apa, ara, arc, arg, arm, arn, arp, art, arw, asm, ast, ath, aus, ava, ave, awa, aym, aze, bad, bai, bak, bal, bam, ban, baq, bas, bat, bej, bel, bem, ben, ber, bho, bih, bik, bin, bis, bla, bnt, tib, bos, bra, bre, btk, bua, bug, bul, bur, byn, cad, cai, car, cat, cau, ceb, cel, cze, cha, chb, che, chg, chi, chk, chm, chn, cho, chp, chr, chu, chv, chy, cmc, cop, cor, cos, cpe, cpf, cpp, cre, crh, crp, csb, cus, wel, dak, dan, dar, day, del, den, ger, dgr, din, div, doi, dra, dsb, dua, dum, dut, dyu, dzo, efi, egy, eka, gre, elx, eng, enm, epo, est, ewe, ewo, fan, fao, per, fat, fij, fil, fin, fiu, fon, fre, frm, fro, frr, frs, fry, ful, fur, gaa, gay, gba, gem, geo, gez, gil, gla, gle, glg, glv, gmh, goh, gon, gor, got, grb, grc, grn, gsw, guj, gwi, hai, hat, hau, haw, heb, her, hil, him, hin, hit, hmn, hmo, hrv, hsb, hun, hup, iba, ibo, ice, ido, iii, ijo, iku, ile, ilo, ina, inc, ind, ine, inh, ipk, ira, iro, ita, jav, jbo, jpn, jpr, jrb, kaa, kab, kac, kal, kam, kan, kar, kas, kau, kaw, kaz, kbd, kha, khi, khm, kho, kik, kin, kir, kmb, kok, kom, kon, kor, kos, kpe, krc, krl, kro, kru, kua, kum, kur, kut, lad, lah, lam, lao, lat, lav, lez, lim, lin, lit, lol, loz, ltz, lua, lub, lug, lui, lun, luo, lus, mac, mad, mag, mah, mai, mak, mal, man, mao, map, mar, mas, may, mdf, mdr, men, mga, mic, min, mis, mkh, mlg, mlt, mnc, mni, mno, moh, mon, mos, mul, mun, mus, mwl, mwr, myn, myv, nah, nai, nap, nau, nav, nbl, nde, ndo, nds, nep, new, nia, nic, niu, nno, nob, nog, non, nor, nqo, nso, nub, nwc, nya, nym, nyn, nyo, nzi, oci, oji, ori, orm, osa, oss, ota, oto, paa, pag, pal, pam, pan, pap, pau, peo, phi, phn, pli, pol, pon, por, pra, pro, pus, qaa-qtz, que, raj, rap, rar, roa, roh, rom, rum, run, rup, rus, sad, sag, sah, sai, sal, sam, san, sas, sat, scn, sco, sel, sem, sga, sgn, shn, sid, sin, sio, sit, sla, slo, slv, sma, sme, smi, smj, smn, smo, sms, sna, snd, snk, sog, som, son, sot, spa, srd, srn, srp, srr, ssa, ssw, suk, sun, sus, sux, swa, swe, syc, syr, tah, tai, tam, tat, tel, tem, ter, tet, tgk, tgl, tha, tig, tir, tiv, tkl, tlh, tli, tmh, tog, ton, tpi, tsi, tsn, tso, tuk, tum, tup, tur, tut, tvl, twi, tyv, udm, uga, uig, ukr, umb, und, urd, uzb, vai, ven, vie, vol, vot, wak, wal, war, was, wen, wln, wol, xal, xho, yao, yap, yid, yor, ypk, zap, zbl, zen, zgh, zha, znd, zul, zun, zxx, zza, multiple, artificial, other

point_of_disambiguation
integer · optional
target onset time (relative to trial onset) in ms
target_side
string · optional
part of the screen in which the target stimulus appears (left or right in 2-choice design) One of left , right .
condition
string · optional
information on the condition manipulation in the given trial (as close as possible as language from the original study)
lab_trial_id
string · optional
trial label assigned within lab (or descriptive label of the trial)
vanilla_trial
boolean · optional
whether or not the trial is a "standard" ("vanilla") looking-while-listening trial or includes additional manipulations. Some requirements for being considered a standard/ vanilla trial include: Familiar words (also no part-words), both target and distractor are familiar objects (i.e., each could be valid familiar targets; non-prototypical objects are OK, but unnatural/ artificially modified exemplars are not), target word is the first point of disambiguation, well-formed/grammatical carrier phrase, no relevant “stuff”/information prenominally (e.g. semantically informative verbs, adjectives), no nonsense words anywhere (including carrier phrase), no language/ speaker/ accent/ etc. switching within trial, no duplicated target words, no intentional background noise or audio filtering. For a full list of current criteria, see the tab "current vanilla criteria". Reach out to the Peekbank team for guidance on handling borderline cases.
trial_type_aux_data
json · optional
optional additional information about the trial type in the format of a JSON (see aux_data tab for information about existing columns)
stimulus_id key
integer
row identifier for the stimuli table indexing from zero
dataset_id → datasets
integer
row identifier for datasets indexing from zero (linked to datasets)
stimulus_novelty
string · optional
whether the stimulus word (not image) is novel or familiar (options: novel, familiar) One of novel , familiar .
original_stimulus_label
string · optional
label (i.e. label for the image when it appears in target position or the label for the image as described in the experiment design) assigned to the stimulus (in the original language of the study)
english_stimulus_label
string · optional
label assigned to the stimulus translated into English
stimulus_image_path
string · optional
relative file path to the visual image stimulus (relative to raw_data/) - has to point to an existing file, should be NA if no such file exists in the raw data
lab_stimulus_id
string · optional
an id or name assigned to the stimulus by the original authors
image_description
string · optional
a short natural language description of the image (can be multiple words)
image_description_source
string · optional
an explanation of the source of the natural language description (Options: taken from the image path; based on experiment documentation; generated by the importer/ Peekbank team) One of image path , experiment documentation , Peekbank discretion .
stimulus_aux_data
json · optional
optional additional information about the stimulus in the format of a JSON (see aux_data tab for information about existing columns)
aoi_timepoint_id key
integer
row identifier for the aoi_timepoints table indexing from zero
aoi
string
name of the looking location (target, distractor, other, missing). other = on-screen looking, but not in the AOI regions for target or distractor. missing = missing data (could be off-screen look, blink, tracker error, etc.) One of target , distractor , other , missing .
administration_id → administrations
integer
row identifier for the administrations table indexing from zero (linked to administrations)
t_norm
integer
timestamp in ms, centered on the point of disambiguation of the target (= 0 ms)
trial_id → trials
integer
row identifier for the trials table indexing from zero (linked to trials)
xy_timepoint_id key
integer
row identifier for the xy_timepoints table indexing from zero
administration_id → administrations
integer
row identifier for the administrations table indexing from zero (linked to administrations)
trial_id → trials
integer
row identifier for the trials table indexing from zero (linked to trials)
x
integer · optional
x-coordinate of gaze position at timepoint t (origin (0,0) is bottom left)
y
integer · optional
y-coordinate of gaze position at timepoint t (origin (0,0) is bottom left)
t_norm
integer
timestamp in ms, centered on the point of disambiguation of the target (= 0 ms)
aoi_region_set_id key
integer
row identifier for the aoi_region_sets table indexing from zero
l_x_max
integer · optional
maximum x-coordinate of the left AOI region (with the bottom left corner of the screen as origin (0,0))
l_x_min
integer · optional
minimum x-coordinate of the left AOI region (with the bottom left corner of the screen as origin (0,0))
l_y_max
integer · optional
maximum y-coordinate of the left AOI region (with the bottom left corner of the screen as origin (0,0))
l_y_min
integer · optional
minimum y-coordinate of the left AOI region (with the bottom left corner of the screen as origin (0,0))
r_x_max
integer · optional
maximum x-coordinate of the right AOI region (with the bottom left corner of the screen as origin (0,0))
r_x_min
integer · optional
minimum x-coordinate of the right AOI region (with the bottom left corner of the screen as origin (0,0))
r_y_max
integer · optional
maximum y-coordinate of the right AOI region (with the bottom left corner of the screen as origin (0,0))
r_y_min
integer · optional
minimum y-coordinate of the right AOI region (with the bottom left corner of the screen as origin (0,0))

The *_aux_data columns hold JSON. What a contributing lab may record in them is documented below; the top-level keys are optional, and the indented fields sit inside them.

subject_aux_data

cdi_responses
JSON · optional
JSON object for CDI responses
instrument_type
string · optional
CDI version/instrument used (e.g., wg / ws / wsshort)
measure
string · optional
CDI measure (comprehension or production)
rawscore
numeric · optional
CDI raw vocabulary score
percentile
numeric · optional
CDI percentile (as reported in the raw data)
age
numeric · optional
age of the participant in months at the time the instrument was administered (conversion from days to months: days / (365.25/12)); conversion from decimal years to months: years * 12; conversion from whole number years to months: years * 12 + 6)
language
string · optional
CDI language (in Wordbank format: Wordbank to ISOcode Mapping)
lui_responses
JSON · optional
JSON object for Language Use Inventory responses
rawscore
numeric · optional
LUI raw score
age
numeric · optional
age of the participant in months at the time the instrument was administered (conversion from days to months: days / (365.25/12)); conversion from decimal years to months: years * 12; conversion from whole number years to months: years * 12 + 6)
language
string · optional
LUI language (in Wordbank format)
lds_rawscore
JSON · optional
JSON object for Language Development Survey responses
rawscore
numeric · optional
LDS raw score
age
numeric · optional
age of the participant in months at the time the instrument was administered (conversion from days to months: days / (365.25/12)); conversion from decimal years to months: years * 12; conversion from whole number years to months: years * 12 + 6)
language
string · optional
LDS language (in Wordbank format)
lab_visit_num
numeric · optional
number of times the participant has been involved in previous experiments
lang_exposures
JSON · optional
JSON object for language exposures
language
string · optional
language that a child is exposed to (in Wordbank format: Wordbank to ISOcode Mapping)
exposure
numeric · optional
percentage exposure in a given language
lang_measures
JSON · optional
JSON object for language measures other than the CDI
instrument_type
string · optional
The specific language development measure used (e.g., LUI, LDS)
rawscore
numeric · optional
raw score in the language measure instrument
age
numeric · optional
age of the participant in months at the time the instrument was administered (conversion from days to months: days / (365.25/12)); conversion from decimal years to months: years * 12; conversion from whole number years to months: years * 12 + 6)
language
string · optional
language of the measure (in Wordbank format)
native_language_non_iso
string · optional
name of participants' native language if the language is not represented in ISO 6390-2

trial_type_aux_data

full_phrase_language_non_iso
string · optional
language of the full phrase if the language is not represented in ISO 6390-2