Peptide Identification Criteria in LC-MS Data Packages
This guidance is intended for qualified laboratory researchers preparing, reviewing, or archiving LC-MS peptide identification data packages. It focuses exclusively on analytical documentation, data integrity, quality control, and research workflow considerations necessary to support transparent and reproducible peptide identifications.
Key identification criteria to document
A defensible peptide identification must include clear, reproducible evidence. At minimum, document: mass accuracy relative to instrument specification; retention time alignment with standards or replicates; presence and quality of fragment ions supporting sequence assignment; and any database search scores or statistical metrics (including how false discovery rate was estimated). Record the search engine(s) and versions, database(s) and versions, enzyme specificity, fixed/variable modifications, precursor and fragment tolerance settings, and post-processing filters used to generate the reported identifications. When applicable, include orthogonal evidence such as synthetic peptide spectra or stable-isotope standards.
Documentation and data integrity practices
Maintain raw data and derived files together with full metadata and instrument logs. Capture software parameter files, spectral peaklists, and the exact commands or GUI exports used to perform searches or spectral matching. Ensure traceability by using consistent file naming, timestamps, and a version-controlled directory structure. Apply ALCOA+ principles (Attributable, Legible, Contemporaneous, Original, Accurate; plus Complete, Consistent, Enduring, Available) to entries in electronic notebooks, LIMS, and audit trails. Record calibration actions, mass calibration files, and any reprocessing steps with rationale and approver signatures.
Quality control and review workflow
Implement QC at multiple stages: instrument performance checks (mass accuracy, resolution, intensity response), chromatographic performance (peak shape, retention-time stability), blank and carryover assessments, and system suitability runs. For identification-specific QC, include decoy database results or target-decoy analysis to justify FDR thresholds, inspect representative MS/MS spectra visually, and confirm that identifications are consistent across biological or technical replicates. Document acceptance criteria for signal-to-noise and peak integration boundaries used during review. Maintain records of failed QC runs and subsequent corrective actions.
Packaging, reporting, and reviewer checklist
A well-constructed data package for peptide identification should include: raw LC-MS files (vendor or open format), centroided/mzML exports if used, search engine output and summary statistics, annotated spectra for key peptides, list of acceptance criteria and the rationale for thresholds, and a metainformation file (instrument, operator, project, sample processing notes). Provide a short reviewer checklist that confirms presence and adequacy of (1) raw data and metadata, (2) parameter and database versions, (3) QC evidence, (4) annotated exemplar spectra, and (5) documented re-analysis or confirmation where applicable. Keep reviewer notes and final sign-off in the package.
Maintaining reproducibility and facilitating re-use
Encourage reproducibility by sharing parameter files and, when possible, open-format exports (mzML, mzIdentML). Annotated spectra and clear links between raw scans and final identifications reduce ambiguity during secondary review. Use persistent identifiers for datasets and keep a changelog for any reprocessing to support transparent provenance tracking. Adhering to community reporting guidelines and including sufficient metadata will streamline downstream interpretation and re-analysis [1,2].
Sources
- https://pmc.ncbi.nlm.nih.gov/articles/PMC3434779/
- https://www.mcponline.org/article/S1535-9476(20)33869-X/fulltext
Not for human consumption. For laboratory research use only.
