Retrieve the peptide metadata table into DuckDB, forcing atomic types
Source:R/peptide-library.R
get_peptide_library.RdThis function uses the phiperio logging utilities for
consistent, ASCII-only progress messages and timing. Long-running steps are
bracketed with .ph_with_timing(), and informational/warning/error
messages are emitted via .ph_log_info(), .ph_log_ok(), .ph_warn(),
and .ph_abort().
Downloads each requested library RDS once, sanitizes types (logical, character, numeric), and writes it into a DuckDB cache on disk.
Subsequent calls return a lazy
tbl_dbiwithout loading into R memory.
Value
A dplyr::tbl_dbi pointing to the requested library: the
peptide_meta_<name> table for a single library, or a view stacking the
tables of several. The returned object carries an attribute "duckdb_con"
with the open DBI connection.
Details
Caching: A persistent DuckDB database is created under the user cache
directory (via tools::R_user_dir("phiperio", "cache")). You can override
this location with options(phiperio.cache_dir = \"...\"). Each library is
stored in its own peptide_meta_<name> table. The force_refresh argument
bypasses the fast path and rebuilds the cache.
Several libraries: The libraries are stacked by column name in a view
named after them (e.g. peptide_meta_combined_icam). Columns that only some
libraries have are NA for the peptides of the others. Peptide IDs carry a
library-specific prefix, so they do not collide.
Sanitization: Columns are stripped of attributes, list-columns are
flattened, textual "NaN" and numeric NaN are coerced to NA. Binary 0/1
fields are converted to logical, "TRUE"/"FALSE" (case-insensitive) are
converted to logical, and numeric-looking character columns (beyond trivial
0/1) are converted to numeric. All other atomic types are preserved.
Integrity check: If a SHA-256 checksum is provided, a warning is logged when the downloaded file’s checksum does not match the expected value.
Examples
lib <- get_peptide_library()
#> [13:09:20] INFO Retrieving peptide metadata into DuckDB cache
#> -> get_peptide_library(library = combined, force_refresh =
#> FALSE)
#> duckdb keeps downloaded extensions and secrets in a temporary directory:
#> ℹ /tmp/Rtmpx3O1yL/duckdb
#> This is removed when the R session ends.
#> • Extensions are re-downloaded each session.
#> • Secrets are lost.
#> ℹ Run duckdb(shared_home = TRUE) (or create ~/.duckdb) to keep them (suitable for most users).
#> ℹ Run duckdb(shared_home = FALSE) to accept the temporary directory (and silence this message).
#> ℹ See ?duckdb_storage for details and alternatives.
#> [13:09:20] INFO Opened DuckDB connection
#> - cache dir:
#> /home/runner/.cache/R/phiperio/peptide_meta/phip_cache.duckdb
#> - tables: peptide_meta_combined
#> [13:09:20] OK Using cached peptide_meta_combined (fast path)
#> [13:09:20] OK Retrieving peptide metadata into DuckDB cache - done
#> -> elapsed: 0.024s