Cloud-Based DIA Proteomics Pipeline
Designed and validated a reproducible AWS cloud pipeline for automated Data-Independent Acquisition (DIA) proteomics quantification. The pipeline moves raw mass-spectrometry data through a secure S3-to-EC2/Batch workflow, orchestrates multi-stage processing with Nextflow, and runs DIA-NN in containerized environments - transforming instrument RAW files into analysis-ready protein/precursor matrices, QC visualizations, and reproducible HTML reports, with built-in resumability, provenance tracking, and checksummed archival.
Cloud EngineeringAWSNextflowDockerDIA-NNProteomics View project on GitHub