Transform Commands¶
Transform and process data within the QPX ecosystem.
Overview¶
The transform command group provides tools for processing and transforming QPX data into various downstream formats. These commands enable gene annotation and protein-level quantification from feature data.
Available Commands¶
- gene-map - Map genes from FASTA
- protein-properties - Fill protein coverage, molecular weight and peptide positions from an optional FASTA
- normalize-accessions - Normalize protein accession formats (full ↔ bare)
- update-metadata - Update sample/run metadata from a revised SDRF
- quantify - Protein quantification via mokume (DirectLFQ, MaxLFQ, iBAQ, TopN, etc.)
gene-map¶
Map gene information from FASTA to parquet format.
Description¶
Enriches protein identifications with gene names read from the ``GN=`` field of the FASTA headers. The FASTA is the only source: nothing is fetched over the network. Given ``--dataset``, every quantification view that carries protein accessions (pg and feature) is annotated, and the dataset's MuData view is rebuilt when one is present, so the ``.h5mu`` never describes stale genes. Dataset mode requires flat Parquet views sharing one file prefix; partitioned pg/feature views are rejected before any files are changed. An output folder must be empty and outside the source dataset.
Parameters¶
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
--parquet-path |
FILE | No | - | QPX PSM, feature or pg parquet file path (single-file mode) |
--dataset |
DIRECTORY | No | - | Path to a QPX dataset directory (annotates pg + feature, refreshes the h5mu) |
--in-place |
FLAG | No | - | Dataset mode: overwrite the dataset's files instead of writing to --output-folder |
--fasta |
FILE | Yes | - | FASTA database file path |
--output-folder |
DIRECTORY | No | - | Output directory for generated files |
--verbose |
FLAG | No | - | Enable verbose logging |
Usage Examples¶
Basic Example¶
Map gene information to parquet file:
# A single parquet file
qpxc transform gene-map \
--parquet-path ./output/psm.parquet \
--fasta proteins.fasta \
--output-folder ./output
# A whole QPX dataset, refreshing its h5mu in place
qpxc transform gene-map \
--dataset ./qpx_output \
--fasta proteins.fasta \
--in-place
Annotating a Feature File¶
qpxc transform gene-map \
--parquet-path ./output/feature.parquet \
--fasta tests/examples/fasta/Homo-sapiens.fasta \
--output-folder ./output
Annotating a Whole Dataset¶
qpxc transform gene-map \
--dataset ./qpx_output \
--fasta tests/examples/fasta/Homo-sapiens.fasta \
--in-place
Every quantification view that carries protein accessions (pg and feature) is
annotated, and the dataset's MuData view is rebuilt when one is present, so the
.h5mu never describes stale genes. Pass --output-folder instead of --in-place
to write an annotated copy and leave the source dataset untouched.
Output Files¶
- Output: Enhanced parquet file(s) with gene information
- Format: Parquet file in output folder
- Added Fields: Gene names and metadata from FASTA headers
- Dataset mode: annotated
pg+featureviews, plus a rebuilt.h5muwhen the dataset has one
Best Practices¶
- Use the same FASTA the search used, so every identified protein can be mapped
- Gene names come from the
GN=field of the FASTA headers; entries without it stay unmapped - Enable verbose mode for debugging
protein-properties¶
Description¶
Fills protein properties a producer did not record, from the FASTA used for the search. DIA-NN never reports protein sequences, and some consensusXML files carry protein hits without them. For target rows only, and only where the value is null:
| Field | Computed as |
|---|---|
pg.molecular_weight |
Average mass of the anchor protein's sequence, in kDa (same definition as the OpenMS consensusXML converter) |
pg.sequence_coverage |
Percent of the anchor covered by the dataset's target peptides whose evidence names it: PSM protein_accessions, feature group membership, recorded positions |
feature.pg_positions |
Every one-based occurrence of the peptide in each member of its protein group |
A value the producer recorded is never overwritten. The FASTA is optional: proteins absent from it — DIA-NN's internal decoys, contaminants from another database, isoforms not in the file — keep a null value and are counted in the report. FASTA decoy entries (DECOY_, REV_, ...) are skipped so they cannot collide with a target's accession. The step is appended to the provenance view with the FASTA's SHA-256.
The same step runs after conversion when --fasta is passed to qpxc convert diann or qpxc convert openms-consensus.
Parameters¶
| Parameter | Description |
|---|---|
--dataset |
QPX dataset directory (flat Parquet views) |
--fasta |
Protein FASTA used for the search (plain or .gz) |
--in-place |
Overwrite the dataset's files |
--output-folder |
Write an annotated copy instead (empty, outside the dataset) |
Usage Examples¶
# Annotate an existing dataset in place
qpxc transform protein-properties \
--dataset ./PXD017199 \
--fasta Homo-sapiens-uniprot-reviewed-contaminants.fasta \
--in-place
# Or during conversion
qpxc convert diann --report-path report.parquet ... --fasta search.fasta
Best Practices¶
- Use the exact database the search used. The command warns when fewer than half of the target protein groups are found.
- Accessions match both as the full identifier (
sp|P12345|NAME) and as the bare accession (P12345). - Coverage reflects the peptides exported in the dataset (after FDR filtering). On PXD000612 it agrees with OpenMS's recorded coverage to a median difference of 0.0 points (93% of groups within 1 point, correlation 0.9995).
normalize-accessions¶
Normalize protein accession formats between full UniProt form (sp|ACC|NAME) and bare form (ACC).
Description¶
Forward (default): converts full UniProt identifiers to bare accessions. sp|P04114|APOB_HUMAN → P04114 CONTAM_sp|CONTAM_P02768|... → CONTAM_P02768 Reverse: converts bare accessions back to full UniProt format. Requires a FASTA database to look up the full identifiers. Normalizes anchor_protein and pg_accessions in both feature.parquet and pg.parquet files.
Parameters¶
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
--dataset |
DIRECTORY | Yes | - | Path to a QPX dataset directory |
--direction |
TEXT | No | forward |
'forward' (sp|ACC|NAME → ACC) or 'reverse' (ACC → sp|ACC|NAME) |
--fasta |
FILE | No | - | FASTA database file (required for --direction reverse) |
--in-place |
FLAG | No | - | Overwrite original files instead of writing to --output |
--output |
DIRECTORY | No | overwrites in place | Output directory (default: overwrites in place) |
--verbose |
FLAG | No | - | Enable verbose logging |
Usage Examples¶
Normalize protein accession formats:
# Forward: strip sp|...|... to bare accessions
qpxc transform normalize-accessions \
--dataset ./my_dataset --direction forward --in-place
# Reverse: restore full UniProt identifiers from FASTA
qpxc transform normalize-accessions \
--dataset ./my_dataset --direction reverse \
--fasta proteins.fasta --in-place
# Forward to a new directory (non-destructive)
qpxc transform normalize-accessions \
--dataset ./my_dataset --direction forward \
--output ./my_dataset_normalized
update-metadata¶
Update sample.parquet and run.parquet metadata from a revised SDRF file, with safety checks on protected fields.
Description¶
Re-generates sample.parquet and run.parquet from the new SDRF. Original files are backed up as *.parquet.bak before overwriting. Safe to update (metadata-only, no impact on data): disease, organism_part, cell_type, cell_line, sex, age, treatment, individual, ancestry, developmental_stage, and any additional characteristics[*] columns. Protected fields (will BLOCK unless --force is used): instrument, enzymes, modification_parameters, fraction, dissociation_method, label/channel mapping, data file references. If --old-sdrf is provided, the tool compares old vs new SDRF and blocks if any protected fields changed. Without --old-sdrf, the safety check is skipped (useful for first-time metadata enrichment).
Parameters¶
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
--dataset |
DIRECTORY | Yes | - | Path to a QPX dataset directory |
--sdrf |
FILE | Yes | - | Path to the updated SDRF TSV file |
--old-sdrf |
FILE | No | - | Path to the original SDRF (for safety checks). If omitted, protected-field validation is skipped. |
--force |
FLAG | No | - | Apply changes even if protected fields (instrument, enzymes, modifications, fractions, labels) have changed. Use with caution. |
--verbose |
FLAG | No | - | Enable verbose logging |
Usage Examples¶
Update dataset metadata from SDRF:
# Update with safety check
qpxc transform update-metadata \
--dataset ./my_project \
--sdrf ./updated_sdrf.tsv \
--old-sdrf ./original_sdrf.tsv
# Update without safety check (first-time enrichment)
qpxc transform update-metadata \
--dataset ./my_project \
--sdrf ./enriched_sdrf.tsv
# Force update even if protected fields changed
qpxc transform update-metadata \
--dataset ./my_project \
--sdrf ./new_sdrf.tsv \
--old-sdrf ./old_sdrf.tsv --force
quantify¶
Compute protein-level quantification from QPX feature data using mokume.
Description¶
Reads a QPX feature.parquet file, extracts peptide-level intensities, and computes protein-level quantification using the selected method. Supported methods: directlfq — DirectLFQ intensity traces (default) maxlfq — MaxLFQ delayed normalization topn — Average of N most intense peptides top3 — Average of 3 most intense peptides ibaq — Intensity-Based Absolute Quantification (requires --fasta) sum — Sum of all peptide intensities
Parameters¶
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
--feature-path |
FILE | Yes | - | QPX feature.parquet file path |
--method |
TEXT | No | directlfq |
Quantification method (directlfq, maxlfq, topn, top3, ibaq, sum) |
--fasta |
FILE | No | - | FASTA database (required for ibaq method) |
--sdrf |
FILE | No | - | SDRF used to map QPX run/channel intensities to samples and conditions |
--enzyme |
TEXT | No | Trypsin | Enzyme for iBAQ digestion (default: Trypsin) |
--topn-n |
INTEGER | No | 3 | N for TopN method (default: 3) |
--threads |
INTEGER | No | -1 | Parallel threads for MaxLFQ (-1 = all cores) |
--output |
PATH | Yes | - | Output file path (.parquet, .tsv, or .csv) |
--normalize |
FLAG | No | - | Normalize quantification values |
--organism |
TEXT | No | human | Organism for iBAQ (default: human) |
--ploidy |
INTEGER | No | 2 | Ploidy for iBAQ ruler (default: 2) |
--cpc |
FLOAT | No | 200 | Cell copies per cell for iBAQ ruler (default: 200) |
--min-aa |
INTEGER | No | 7 | Min peptide length for iBAQ (default: 7) |
--max-aa |
INTEGER | No | 30 | Max peptide length for iBAQ (default: 30) |
--verbose |
FLAG | No | - | Enable verbose logging |
Supported Methods¶
| Method | Description | Extra Requirements |
|---|---|---|
directlfq |
DirectLFQ intensity traces (default) | pip install "qpx[quantify]" |
maxlfq |
MaxLFQ delayed normalization | pip install "qpx[quantify]" |
topn |
Average of N most intense peptides | pip install "qpx[quantify]" |
top3 |
Average of 3 most intense peptides | pip install "qpx[quantify]" |
ibaq |
Intensity-Based Absolute Quantification | pip install "qpx[quantify]" + --fasta |
sum |
Sum of all peptide intensities | pip install "qpx[quantify]" |
Usage Examples¶
DirectLFQ (default)¶
qpxc transform quantify \
--feature-path ./qpx_output/feature.parquet \
--method directlfq \
-o proteins_directlfq.parquet
iBAQ (requires FASTA)¶
qpxc transform quantify \
--feature-path ./qpx_output/feature.parquet \
--method ibaq --fasta proteome.fasta \
-o proteins_ibaq.tsv
MaxLFQ with 8 threads¶
qpxc transform quantify \
--feature-path ./qpx_output/feature.parquet \
--method maxlfq --threads 8 \
-o proteins_maxlfq.parquet
TopN with normalization¶
qpxc transform quantify \
--feature-path ./qpx_output/feature.parquet \
--method topn --topn-n 5 --normalize \
-o proteins_top5.parquet
Output Files¶
- Parquet:
.parquetfiles with protein-level quantification - TSV:
.tsvfiles (tab-separated) — determined by output file extension - Content: Protein accessions, sample IDs, and quantified intensities
Common Issues¶
Issue: mokume is not installed / DirectLFQ is not installed
- Solution: Install with
pip install "qpx[quantify]"(pullsmokume[directlfq]>=0.1.0)
Issue: --fasta option is required for the ibaq method
- Solution: Provide a FASTA database file with
--fasta
Best Practices¶
- Ensure QPX feature.parquet contains valid
anchor_protein,intensities, andrun_file_namefields - If using older QPX datasets, ensure fields have been migrated from legacy names (
precursor_charge→charge,id_scan→scan) - Decoy entries (
is_decoy=true) and zero-intensity rows are automatically filtered - Use
--normalizefor cross-sample normalization - Use
--threadsto control parallelism for MaxLFQ
Related Commands¶
- Convert Commands - Convert raw data to QPX format
- Visualization Commands - Visualize transformed data
- Statistics Commands - Analyze transformed data