Skip to content
GitHub

PeekbankR Functions

Every function here takes the connection created by connect_to_peekbank(), as described in Data Access - PeekbankR.

One function per table in the data schema, which describes what each table holds. Every one takes a connection; the rest of the arguments narrow what comes back.

FunctionTableFilter argumentsSize in 2026.1
get_datasets()datasets-44 rows
get_subjects()subjects-3,915 rows
get_administrations()administrationsdataset_name, dataset_id, age5,596 rows
get_trials()trialsdataset_name, dataset_id134,929 rows
get_trial_types()trial_typesdataset_name, dataset_id6,918 rows
get_stimuli()stimulidataset_name, dataset_id1,834 rows
get_aoi_timepoints()aoi_timepointsdataset_name, dataset_id, age36.1M rows, 3.4 GB
get_xy_timepoints()xy_timepointsdataset_name, dataset_id, age10.0M rows, 994 MB
get_aoi_region_sets()aoi_region_sets-44 rows

dataset_name and dataset_id are separate arguments that do the same job, and each takes a vector, so dataset_name = c("pomper_saffran_2016", "reflook_v4") works. age takes a single age in months or a minimum and maximum. Called without filters, a function returns its whole table, so filter the timepoints unless you really want all of it.

datasets, subjects, administrations, trials, trial_types and stimuli each carry a *_aux_data column holding JSON: CDI scores on subjects, and whatever else a contributing lab recorded that the schema has no column for. The get_ functions return it as a raw string, so unpack it before use:

subjects <- get_subjects(connection = con) %>%
  unpack_aux_data()

An escape hatch for anything the get_ functions do not cover. Takes BigQuery Standard SQL against the tables above; string comparison is case-sensitive.

get_sql_query("SELECT dataset_name, shortcite FROM datasets", connection = con)

Downloads the per-dataset READMEs, which record import decisions and dataset quirks, into dataset_readmes/. These always come from the latest released files, not from the release your connection is pinned to.

get_readmes(datasets = "pomper_saffran_2016")

Downloads the stimulus images into stimulus_data/, skipping any already on disk, and returns the stimulus table with a column of local paths.

stimuli <- download_stimuli(con, datasets = "reflook_v4")

CDI scores arrive as JSON in subject_aux_data. Once unpacked and cleaned, this adds the reference age and year used along with gender-specific and general percentiles.

cdi <- get_subjects(connection = con) %>%
  unpack_aux_data() %>%
  tidyr::unnest(subject_aux_data) %>%
  dplyr::filter(!is.na(cdi_responses)) %>%
  tidyr::unnest(cdi_responses) %>%
  cleanup_cdi_data() %>%
  populate_cdi_percentiles()