ImmunData: A data structure for storing adaptive immune receptor repertoire data¶
Description¶
ImmunData stores adaptive immune receptor repertoire (AIRR)
data and the rules used to turn observed sequences or cells into data
units for analysis. Think AnnData or SeuratObject, but for immune
repertoires.
You work with an ImmunData object after importing bulk or
single-cell AIRR-seq data. The major idea behind ImmunData
is that because sequencing provides only information about sequences
and, for single-cell data, cell barcodes, the your responsibility is to
determine, which sequences you want to treat as the same receptor,
repertoire, or stratum (group of repertoires). You define these analysis
units with schemas. A schema is a stored set of column names and
chain-selection rules that tells ImmunData how to group
observations. Those definitions are kept inside ImmunData
to ensure that downstream functions count, filter, and compare the same
units consistently. Repertoire and strata schemas can be changed later
to re-aggregate repertoires differently, e.g., merge receptors from
different clusters into per-patient clusters. Receptor schema is fixed
once and for all, so if you want to work with a different receptor
definition, e.g., use "CDR3aa + V gene" instead of just "CDR3aa" as a
definiton for a unique receptor, you will need to create a separate
ImmunData object.
ImmunData is immutable, meaning that functions that
transform an ImmunData object return a new object, and the
original object is not changed. Due to multiple optimisations on the
backend, it does not mean that you re-create the whole dataset each time
you run a, let’s stay, a filter. However, it does affect analysis
workflow significantly. You can read about it more on the website and in
tutorials.
From observed data to analysis units¶
ImmunData connects observed records to user-defined
analysis units:
- A chain observation is an observed receptor-chain sequence, such as a TRA, TRB, or IGH sequence. Chain observations form the main table.
- A barcode is an observed identifier for a cell in single-cell data. It links chains found in the same cell.
- A receptor is a virtual analysis unit that you define. For example, you may define it by CDR3 sequence alone, by CDR3 and V gene, or as a paired TRA-TRB receptor. The receptor schema records which chain features and loci must match for observations to receive the same receptor identifier.
- A repertoire is a virtual collection of receptors that you define from annotation columns. For example, one repertoire may contain all receptors from one sample, or from one donor at one time point.
- A stratum is a virtual collection of repertoires for a comparison. For example, one stratum may contain all repertoires from one treatment arm.
These definitions do not change the observed sequences. They determine
how observations are grouped and counted during analysis. The resulting
hierarchy is chain observations and
barcodes -> receptors -> repertoires -> strata.
Inspect and transform an object¶
Print an object for a compact overview. Use
$receptors for the receptor
table, $repertoires for one
summary row per repertoire, and
$strata for one row per stratum.
Most analysis functions accept the complete ImmunData
object directly.
Common transformations include:
-
filter_immundata()to keep selected chains, cells, or receptors; -
mutate_immundata()to calculate annotation columns; -
annotate()to add external biological information; -
agg_repertoires()to define repertoires; and -
agg_strata()to group repertoires into strata.
Create an object¶
Create an ImmunData object with
read_repertoires(), or reopen a saved object with
read_immundata(). Do not call the
$new() constructor in analysis
code. Direct construction is reserved for package developers.
Lazy data and storage¶
The chain-level table uses duckplyr and can remain on disk. Filtering,
mutation, and aggregation stay lazy when possible, so large datasets do
not need to be loaded fully into R memory. Downstream analysis functions
in the immunarch package are designed to accept lazy
ImmunData objects. Pass the object directly; you usually do
not need to call dplyr::collect(). Collect data only when
another function explicitly requires an in-memory data frame or when you
want to inspect a small table in R.
Objects created by read_repertoires() are backed by files
in their output folder. Keep that folder while you use the object. Use
write_immundata() to save a transformed object and
read_immundata() to reopen it.
Public fields¶
-
schema_receptor -
A named list defining the virtual receptor unit. The
featureselement names the chain columns used to group observations, such as CDR3 sequence and V gene. Thechainselement selects one chain or a paired set of chains. -
schema_repertoire -
A character vector naming annotation columns whose unique combinations
define one repertoire, such as
sample_idorc(“donor_id”, “timepoint”). It isNULLwhen repertoires have not been defined. -
schema_strata -
A character vector naming repertoire-level columns whose unique
combinations define one stratum, such as
treatment. It isNULLwhen strata have not been defined.
Active bindings¶
-
receptors - A derived duckplyr table of distinct receptors. For a paired receptor, the selected chain features are shown side by side.
-
annotations -
The lazy duckplyr table of retained chain observations and their
biological annotations. For most tasks, pass the complete
ImmunDataobject to a transformation function or usecollect(idata)to inspect this table in memory. -
repertoires -
A small table with one row per repertoire, the columns that define it,
and summary statistics such as
n_barcodesandn_receptors. It isNULLwhen repertoires have not been defined. -
strata -
A small table with one row per stratum, its label, and the
repertoire-level columns that define it. It is
NULLwhen strata have not been defined. -
provenance -
Read-only named list describing the snapshot origin and storage context
carried by this object. Retrieve the complete list with
idata$provenance, or one field with, for example,idata$provenance$current_path. The fields are:-
home_path: project home used for managed snapshots and artifacts. The original ingestion snapshot is stored directly in this folder; it isNULLfor an object with no persisted home. -
current_path: exact folder of the most recently loaded or written snapshot. Transformations preserve this source path until the transformed object is written as another snapshot; it isNULLfor an object that has never been loaded from or written to disk. -
snapshot_root: derived managed-snapshot root,home_path/snapshots, orNULLwhenhome_pathisNULL. -
artifacts_root: derived project-level root for optional external tool outputs,home_path/artifacts, orNULLwhenhome_pathisNULL. -
artifacts_path: derived namespace for artifacts associated with the most recently loaded or written snapshot. It isartifacts_root/rootfor the original ingestion,artifacts_root/\for a managed snapshot, and/vNNN artifacts_root/by-id/\for a detached explicit snapshot. External tools can append\and create that directory; artifact contents are not part of/\ ImmunData. Write a transformed object as a new snapshot before storing artifacts that should be associated with the transformed data. -
snapshot_id: unique identifier generated when the snapshot is written;NULLfor an in-memory object that has never been written. -
lineage: ordered list of ingestion and snapshot events leading to the current snapshot.
idata$provenanceis an error. -
Methods¶
Public methods
ImmunData$new()
Low-level constructor for package developers. Analysis code must create
an ImmunData object with read_repertoires() or
reopen one with read_immundata().
Usage
ImmunData\$new( schema, annotations, repertoires = NULL, provenance = NULL, strata = NULL )
Arguments
-
schema -
A character vector or named list. A character vector names the features
used to define a chain-agnostic receptor. A named list is created by
make_receptor_schema()and can also select receptor chains. -
annotations - A duckplyr table. It contains retained chain observations, receptor identifiers, and biological annotations.
-
repertoires -
A data frame or
NULL. It contains one row per repertoire and its summary statistics and is usually created byagg_repertoires(). -
provenance -
A list or
NULL. It contains internal storage and snapshot history. -
strata -
A data frame or
NULL. It contains one row per stratum, its label, and the repertoire-level columns that define it.
ImmunData$clone()
The objects of this class are cloneable with this method.
Usage
ImmunData\$clone(deep = FALSE)
Arguments
-
deep - Whether to make a deep clone.
See Also¶
read_repertoires(), read_immundata(),
write_immundata(), agg_repertoires(),
agg_strata(), filter_immundata(),
mutate_immundata(), annotate()
Examples¶
library("immundata")
library(immundata)
library(dplyr)
options(immundata.verbose = FALSE)
# Load the small dataset included with immundata, then define one repertoire
# for each treatment-response group.
idata <- get_test_idata() |>
agg_repertoires(schema = "Response")
idata$repertoires |>
select(Response, n_barcodes, n_receptors) |>
arrange(Response)
#> # A tibble: 2 × 3
#> Response n_barcodes n_receptors
#> * <chr> <dbl> <int>
#> 1 FR 955 871
#> 2 PR 947 867
# Expected result:
# Response n_barcodes n_receptors
# FR 955 871
# PR 947 867
# Under the current receptor definition, the full-response (FR) repertoire
# contains 955 chain observations grouped into 871 receptor units.
# Keep only the full-response repertoire. filter() returns a new object.
fr_only <- idata |>
filter(Response == "FR")
tibble(
original_repertoires = nrow(idata$repertoires),
filtered_repertoires = nrow(fr_only$repertoires)
)
#> # A tibble: 1 × 2
#> original_repertoires filtered_repertoires
#> <int> <int>
#> 1 2 1