Generated API

This page is generated from public Python docstrings with mkdocstrings. The parameter-level migration map and R correspondence remain in the repository file PYTHON_API_REFERENCE.md; this page follows the actual installed Python signatures.

Core

pycardinal.core

Core dataset types and processing infrastructure.

MSImagingArrays

Bases: SpectralImagingArrays

Mass-spectrometry dataset stored as ragged m/z/intensity spectra.

mz property writable

Per-pixel m/z arrays.

intensity property writable

Per-pixel intensity arrays.

MSImagingExperiment

Bases: SpectralImagingExperiment

Shared-m/z mass-spectrometry imaging experiment.

The intensity matrix is always oriented (features, pixels); each row corresponds to the matching sorted m/z value in feature_data.

mz property writable

Shared m/z values, one per matrix row.

intensity property writable

Intensity matrix oriented as (features, pixels).

is_centroided()

Return whether the data is known to be centroided.

__getitem__(key)

Subset by feature rows and pixel columns.

SpectralImagingArrays

Bases: SpectralImagingData

Dataset of per-pixel spectra whose m/z axes may differ.

SpectralImagingData

Shared base for spectral-imaging datasets.

spectra_data property

Named spectral arrays stored by the dataset.

pixel_data property writable

Metadata with one row per pixel/spectrum.

coord property

Pixel coordinate columns.

run property

Run labels for each pixel.

processing property writable

Queued processing steps, in execution order.

__len__()

Number of pixels/spectra in the dataset.

copy(deep=True)

Copy the dataset, optionally copying underlying array values.

SpectralImagingExperiment

Bases: SpectralImagingData

Shared-domain data matrix with features in rows and pixels in columns.

feature_data property writable

Metadata with one row per shared-domain feature.

shape property

Matrix shape as (features, pixels).

spectra(name='intensity')

Return a named shared-domain matrix (features by pixels).

__getitem__(key)

Subset by dataset[feature_selector, pixel_selector].

MassDataFrame

Bases: XDFrame

Feature metadata with a finite, nondecreasing mz column.

mz property

Sorted m/z values as a NumPy array.

PositionDataFrame

Bases: XDFrame

Pixel metadata with numeric x/y coordinates and a run.

Coordinates can be supplied as a DataFrame, a mapping, or an N x 2 (or N x 3) array. If omitted, run defaults to a single run.

coord property

Coordinate columns in their original x/y/z order.

run property

Categorical run labels, one per pixel.

XDFrame

Bases: DataFrame

DataFrame base that records columns identifying domain keys.

key_columns is used instead of keys to preserve pandas' DataFrame.keys() method.

ProcessingStep dataclass

A queued spectrum function and the arguments needed to call it.

SpectraArrays

Bases: MutableMapping[str, Any]

Store named arrays with identical shapes.

Ragged spectra may be stored as sequences: their shared outer dimension is checked here, while each spectrum's internal length may differ.

names property

Names of the contained arrays in insertion order.

copy()

Return a new container sharing the stored array values.

add_processing(obj, fn, label, metadata=None, **fn_kwargs)

Return a shallow copy with a processing function appended to its queue.

The function is called as fn(intensity, mz, **fn_kwargs). It may return a new intensity vector or a pair (new_mz, new_intensity).

process(obj, *, n_jobs=None, chunk_size=None, verbose=False)

Apply queued processing steps and return a new dataset.

Each step is applied to every spectrum in queue order. Shared-domain data must retain the same m/z axis across all pixels. File output is deferred to the I/O phase and is not accepted here.

reset(obj)

Return a shallow copy with all pending processing steps removed.

Processing

pycardinal.processing

Spectral preprocessing operations.

reduce_baseline(obj, method='locmin', *, window=31, iterations=40)

Queue baseline estimation and subtraction for every spectrum.

Methods are locmin, hull, snip, and median. The returned dataset is a copy with a deferred step; call process to apply it. Output intensities are clipped at zero after baseline subtraction.

bin_spectra(obj, ref=None, *, method='sum', resolution=None, tolerance=None, units='ppm', mass_range=None)

Eagerly bin spectra onto shared m/z values.

Aggregation methods sum, mean, max, and min are supported. Interpolation names from the R API are recognized but not yet implemented.

estimate_domain(xlist, width='median', units='relative')

Estimate a regular shared domain from per-spectrum m/z vectors.

estimate_reference_mz(obj, width='median', units='ppm')

Return the shared m/z axis or estimate one from ragged spectra.

estimate_reference_peaks(obj, method='diff', snr=2.0)

Detect reference peaks from the mean spectrum of a dataset.

peak_align(obj, ref=None, *, method='diff', snr=2.0, tolerance=None, units='ppm', binratio=2.0, n_jobs=None)

Apply queued steps and align detected peaks to a shared sparse matrix.

When ref is omitted, local maxima are detected, merged by tolerance, and used as the reference axis. This is an eager operation and returns centroided feature-by-pixel data.

peak_pick(obj, ref=None, method='diff', snr=2.0, type_='height', tolerance=None, units='ppm', **options)

Queue local-maximum peak picking, optionally extracting a reference.

Methods diff, sd, mad, quantile, filter, and cwt select noise estimates/detection backends. type_ is height or area. Without ref, ragged spectra become per-spectrum peak lists; shared-axis experiments keep the shared axis and zero non-peak samples.

peak_process(obj, ref=None, *, method='diff', snr=2.0, type_='height', tolerance=None, units='ppm', sample_size=None, binratio=2.0, filter_freq=True, n_jobs=None, **peak_options)

Convenience pipeline for peak picking followed by peak alignment.

If sample_size is set and ref is omitted, evenly spaced spectra are peak-picked to estimate a reference before extracting peaks from the full dataset. A fraction in (0, 1) is treated as a proportion; values at least 1 are interpreted as a count.

Features and spatial utilities

pycardinal.features

Feature and pixel selection utilities.

features(obj, query=None, **conditions)

Return zero-based feature indices matching metadata predicates.

query is a pandas expression, such as "mz >= 500 and quality > 0.8". Keyword conditions use exact column names for equality, or column__op with eq, ne, lt, le, gt, ge, in, or notin. A callable condition receives the metadata Series and returns a boolean mask. Multiple predicates are combined with logical AND.

pixels(obj, query=None, **conditions)

Return zero-based pixel indices matching pixel metadata predicates.

Query strings and keyword conditions follow :func:features semantics.

subset_features(obj, query=None, **conditions)

Return a shared-domain experiment with matching feature rows.

subset_pixels(obj, query=None, **conditions)

Return a dataset containing pixels matching metadata predicates.

subset(obj, select=None, subset=None)

Subset shared-domain data by feature and pixel selectors.

select indexes feature rows and subset indexes pixel columns. For ragged MSImagingArrays, only pixel selection is defined.

colocalized(obj, i=None, mz=None, ref=None, threshold='median', n=np.inf, sort_by='cor')

Rank features by colocalization against a reference image or feature.

i selects a feature index, mz selects matching feature m/z values, and ref may be a numeric pixel vector or a logical mask. The returned DataFrame contains the ranking columns plus feature metadata.

slice_image(obj, i=None, run=None, simplify=True, drop=True, **conditions)

Extract selected feature intensities as 2D ion-image rasters.

Coordinates are rasterized with rows ordered by ascending y and columns by ascending x. Missing grid positions are NaN. With simplify=True, results are stacked as (feature, run, y, x) and singleton dimensions are removed when drop=True. Otherwise a feature-major list of per-run image lists is returned. Three-dimensional coordinates are not yet rasterized.

pycardinal.spatial

Spatial analysis utilities.

find_neighbors(coord_or_obj, r=1, groups=None, metric='maximum', p=2, matrix=False)

Find each point's neighbors within radius r.

The point itself is included. metric="maximum" uses Chebyshev distance; euclidean and minkowski use the supplied Minkowski p. Group labels prevent neighbors from crossing run or other group boundaries.

spatial_dists(x, y, coord=None, r=1, neighbors=None, neighbors_weights=None, weights=None, byrow=True, metric='euclidean', p=2)

Return neighborhood-weighted distances from observations in x to y.

The result has one row per center coordinate and one column per observation in y. Set byrow=False for a feature-by-pixel matrix.

spatial_weights(x, coord=None, r=1, neighbors=None, weights='gaussian', sd=None, matrix=False)

Calculate Gaussian spatial weights, optionally modulated by data similarity.

For an imaging experiment, pixel coordinates and feature vectors are inferred. Otherwise x supplies coordinates unless coord is explicitly provided. Adaptive weights multiply the spatial Gaussian by a per-neighborhood Gaussian based on Euclidean distances between data vectors.

Summaries and models

pycardinal.summarize

Summary statistics for shared-domain spectral-imaging data.

row_stats(x, stat, *, na_rm=False)

Reduce each row of a two-dimensional dense or sparse matrix.

col_stats(x, stat, *, na_rm=False)

Reduce each column of a two-dimensional dense or sparse matrix.

summarize_features(obj, stat='mean', groups=None, *, na_rm=False)

Add per-feature summaries across pixels to feature metadata.

Group labels refer to pixels. Grouped columns are named group.stat.

summarize_pixels(obj, stat={'tic': 'sum'}, groups=None, *, na_rm=False)

Add per-pixel summaries across features to pixel metadata.

Group labels refer to features. Grouped columns are named group.stat.

pycardinal.stats

Statistical and machine-learning methods.

ContrastTest dataclass

Estimated contrasts derived from a means test.

MeansTest dataclass

Regression coefficients and fitted models from a means test.

SpatialCV dataclass

Fold-level scores and optional fitted models from cross-validation.

SpatialDGMM dataclass

Per-feature Gaussian-mixture segmentation result.

predict(newdata)

Predict one class label per selected feature and observation.

log_lik()

Return the summed fitted-data log likelihood.

SpatialFastmap dataclass

Spatially smoothed low-dimensional projection result.

x property

Return projected observation scores.

predict(newdata, **_)

Project new observations with the fitted model.

SpatialKMeans dataclass

K-means cluster labels, centers, and feature correlations.

predict(newdata)

Assign new observations to fitted clusters.

top_features(n=np.inf, sort_by='correlation')

Rank features by cluster correlation.

SpatialOPLS dataclass

Bases: SpatialPLS

Orthogonal partial-least-squares result.

coef()

Return feature-by-response regression coefficients.

residuals()

Return training response residuals.

SpatialPCA dataclass

Principal-component scores, loadings, and fitted PCA state.

x property

Return observation scores.

rotation property

Return feature loadings.

predict(newdata)

Project new observations into the fitted PCA space.

SpatialPLS dataclass

Partial-least-squares regression or classification result.

predict(newdata=None, ncomp=None, type='response', simplify=True)

Predict responses or classes using one or more components.

fitted(type='response')

Return fitted responses or classes for training data.

top_features(n=np.inf, sort_by='vip')

Rank features using fitted model loadings.

SpatialShrunkenCentroids dataclass

Supervised or unsupervised shrunken-centroid result.

predict(newdata)

Predict class labels for new observations.

fitted(type='class')

Return fitted classes or probabilities.

predict_from_observations(matrix)

Predict labels from an observation-by-feature matrix.

top_features(n=np.inf, sort_by='statistic')

Rank features by between-class center differences.

NMF(x, ncomp=3, method='als', random_state=0, **kwargs)

Fit nonnegative matrix factorization with the selected solver.

OPLS(x, y, ncomp=3, retx=True, center=True, scale=False, **kwargs)

Fit the available OPLS-compatible supervised baseline.

PCA(x, ncomp=3, center=True, scale=False, **kwargs)

Fit principal components to observations by features.

PLS(x, y, ncomp=3, method='nipals', center=True, scale=False, **kwargs)

Fit a supervised PLS baseline for regression or classification.

contrast_test(fit, specs, method='pairwise', emm_adjust='none')

Return simple pairwise contrasts from a fitted means test.

cross_validate(fit_fn, x, y=None, folds=None, predict_fn=None, keep_models=False, **fit_kwargs)

Fit and score a model independently on each validation fold.

means_test(x, fixed, random=None, samples=None, response='intensity', reduced='~1', use_lmer=False, **kwargs)

Fit a formula-based means model to per-pixel experiment summaries.

segmentation_test(x, fixed, random=None, samples=None, class_=1, response='intensity', reduced='~1')

Run a means test using the first spatial segmentation class.

spatial_dgmm(x, i=None, coord=None, r=1, k=2, groups=None, weights='gaussian', neighbors=None, annealing=True, compress=True, random_state=0, **kwargs)

Fit independent Gaussian mixtures for selected feature columns.

spatial_fastmap(x, coord=None, r=1, ncomp=3, weights='gaussian', neighbors=None, transpose=True, niter=10, **kwargs)

Fit a neighborhood-smoothed low-dimensional spatial projection.

spatial_kmeans(x, coord=None, r=1, k=2, ncomp=None, weights='gaussian', neighbors=None, transpose=True, niter=10, centers=True, correlation=True, random_state=0, **kwargs)

Cluster observations with K-means for one or more cluster counts.

spatial_shrunken_centroids(x, y=None, coord=None, r=1, k=2, s=0, weights='gaussian', neighbors=None, bags=None, priors=None, init=None, threshold=0.01, niter=10, random_state=0, **kwargs)

Fit a supervised or unsupervised shrunken-centroid baseline.

top_features(fit, n=np.inf, sort_by='vip')

Rank features for a supported Phase 6 result object.

Plotting and ROI

pycardinal.plotting

Plotting helpers for spectra and ion images.

image_model(fit, type='x', ax=None, **kwargs)

Plot a model's spatial values when pixel coordinates are attached.

plot_image(obj, feature=None, i=None, superpose=False, scale=False, ax=None, cmap='viridis', **kwargs)

Plot one or more ion images and return the Matplotlib axes.

plot_model(fit, type='scores', ax=None, **kwargs)

Plot common result-object fields such as scores, centers, or clusters.

plot_spectra(obj, i=None, superpose=False, xlim=None, ylim=None, ax=None, **kwargs)

Plot selected spectra and return the Matplotlib axes.

pycardinal.roi

ROI selection and categorical-mask helpers.

make_factor(*, ordered=False, **named_masks)

Combine named boolean masks into a first-match categorical factor.

select_roi(obj, mode='region', polygon=None, points=None, tolerance=0.5, **_)

Return a pixel mask from a polygon or selected coordinate points.

polygon and points provide a deterministic, noninteractive path for scripts and tests. When neither is provided, an interactive Matplotlib polygon selector is opened for mode="region".

I/O and simulation

pycardinal.io

Mass spectrometry imaging file input and output.

convert_arrays_to_experiment(obj, *, mz=None, mass_range=None, resolution=None, units='ppm', guess_max=1000, tolerance=None)

Bin ragged spectra onto a shared m/z axis and return an experiment.

If all spectra already share exactly the same axis, it is reused. For varying axes, supply mz or both mass_range and resolution. Nearest-axis bins outside tolerance are omitted; default tolerance is half the requested resolution (or half a local axis interval).

convert_experiment_to_arrays(obj)

Return nonzero per-pixel spectra from a shared-domain experiment.

read_imzml(file, *, memory=True, check=False, mass_range=None, resolution=None, units='ppm', guess_max=1000, as_='auto', parse_only=False, verbose=False)

Read an imzML file into a mass-imaging dataset.

This first implementation reads spectra eagerly. memory=False is reserved for future file-backed loading and currently raises NotImplementedError rather than silently loading the full file.

Parameters:
  • file (str | Path) –

    Path to the imzML XML file. Its matching .ibd sidecar must be present next to it.

  • memory (bool, default: True ) –

    Must be True; lazy/file-backed loading is not implemented yet.

  • check (bool, default: False ) –

    If true, ask pyimzml to validate the imzML/ibd checksum where supported.

  • mass_range (tuple[float, float] | None, default: None ) –

    Optional (minimum_mz, maximum_mz) used when converting processed spectra to a shared m/z axis.

  • resolution (float | None, default: None ) –

    Positive shared-axis spacing. Required with mass_range when converting processed spectra to an experiment unless mz is otherwise available through :func:convert_arrays_to_experiment.

  • units (str, default: 'ppm' ) –

    "mz" for absolute spacing or "ppm" for parts-per-million.

  • guess_max (int, default: 1000 ) –

    Number of evenly spaced spectra sampled to infer a mass range when conversion to a shared axis is requested.

  • as_ (str, default: 'auto' ) –

    "auto" preserves processed spectra as ragged arrays and reads continuous data as an experiment. "arrays" and "experiment" force a representation.

  • parse_only (bool, default: False ) –

    Return parsed metadata and coordinates without reading spectra.

  • verbose (bool, default: False ) –

    Emit a short progress message through the standard Python logger.

Returns:
Raises:
  • FileNotFoundError –

    If the XML or binary sidecar is absent.

  • ValueError –

    If as_ or conversion parameters are invalid.

  • NotImplementedError –

    If memory=False is requested before file-backed loading exists.

write_imzml(obj, file, *, bundle=True, verbose=False)

Write a dataset and its binary sidecar as an imzML pair.

With bundle=True, file names a directory and output XML uses that directory's name. With bundle=False, file is the output XML path.

read_msi_data(file, **kwargs)

Read a supported mass-spectrometry imaging file by extension.

write_msi_data(obj, file, **kwargs)

Write a supported mass-spectrometry imaging file by extension.

read_analyze(file, **kwargs)

Read Analyze 7.5 data (not implemented yet).

write_analyze(obj, file, **kwargs)

Write Analyze 7.5 data (not implemented yet).

pycardinal.simulate

Synthetic mass spectra and imaging experiments for examples and tests.

simulate_spectra(n=1, npeaks=50, mz=None, intensity=None, from_=None, to=None, by=400, sdpeaks=None, sdpeakmult=0.2, sdnoise=0.1, sdmz=10, resolution=1000, fmax=0.5, baseline=0, decay=10, units='ppm', centroided=False, random_state=None)

Generate one or more reproducible profile or centroided spectra.

Returns a MassDataFrame with one m/z value per row. A single spectrum uses an intensity column; multiple spectra use intensity_1 through intensity_n. Profile output is sampled on a shared domain. Centroided output uses the supplied/theoretical peak positions as its shared domain.

add_shape(pixel_data, center, size, shape='circle', name=None)

Return pixel metadata with a circle or square ROI mask column added.

preset_image_def(preset=1, nrun=1, npeaks=30, dim=(20, 20), peakheight=np.e, peakdiff=np.e, sdsample=0.2, jitter=True, random_state=None)

Build pixel/feature designs for one of nine example image presets.

Presets provide compact, deterministic-in-design examples rather than reproductions of Cardinal's R random-number stream. Presets 1-8 are 2D; preset 9 creates a 3D lattice when dim has three entries.

simulate_image(pixel_data=None, feature_data=None, preset=None, *, from_=None, to=None, by=400, sdrun=1.0, sdpixel=1.0, spcorr=0.3, sar=False, resolution=1000, fmax=0.5, units='ppm', centroided=False, continuous=True, random_state=None, **preset_kwargs)

Simulate a complete imaging experiment from a design or preset.

Matching boolean columns in pixel_data and numeric columns in feature_data define where each feature is present and its mean signal. If preset is supplied, preset_image_def creates that design. random_state accepts an integer seed or a NumPy Generator.