Lab Newsletter — September 27, 2026: Predicting the Spectrum
AI for life science — daily digestYesterday we drew the cell’s protein wiring diagram; to build one you first have to detect the proteins. The workhorse for that is mass spectrometry: shatter peptides into fragments, measure the pieces, and infer what was there. For decades the catch was that nobody could accurately predict what a given peptide’s fragmentation spectrum should look like — so searches leaned on slow experimental libraries or crude theoretical guesses. Today’s digest is about the deep-learning models that finally predict peptide spectra from sequence alone, and how that quietly unlocked faster, deeper, library-free proteomics — more protein data per experiment, exactly what data-hungry cell models need.
🎼 The first accurate predictions
Could a network learn the rules of peptide fragmentation? Zhou et al. (Analytical Chemistry, 2017) showed it could with pDeep, “a deep neural network-based model for the spectrum prediction of peptides.” Using “bidirectional long short-term memory (BiLSTM),” pDeep predicted multiple fragmentation modes “with >0.9 median Pearson correlation coefficients,” and — a striking hint that it learned real chemistry — could “distinguish extremely similar peptides … (GG = N, AG = Q, or even I = L),” cases “very difficult to distinguish using traditional search engines.”
🎯 Predictions better than the measurement
Then the models got good. Gessulat et al. (Nature Methods, 2019) trained Prosit on “550,000 tryptic peptides and 21 million high-quality tandem mass spectra,” producing “chromatographic retention time and fragment ion intensity predictions that exceed the quality of the experimental data.” Folded into search pipelines, that precision meant “more identifications at >10× lower false discovery rates” — and Prosit could go further, “generating spectral libraries for data-independent acquisition” from “peptide sequence alone.” A predictor accurate enough to replace measurement in parts of the workflow.
🔬 Fragmentation is a long-range affair
Why does a peptide break where it does? Tiwary et al. (Nature Methods, 2019) built DeepMass:Prism and showed “machine learning can predict peptide fragmentation patterns in mass spectrometers with accuracy within the uncertainty of measurement.” Analyzing the model revealed biology, not just fit: “peptide fragmentation depends on long-range interactions within a peptide sequence.” And practically, using predicted spectra for data-independent acquisition was “nearly equivalent to the use of spectra from experimental libraries” — the interpretability-plus-utility combination the lab prizes.
📚 Libraries with no experiments
DIA is powerful but was shackled to a slow prerequisite: build an experimental (DDA) spectral library first. Yang et al. (Nature Communications, 2020) cut that cord with DeepDIA, generating “in silico spectral libraries for DIA analysis” whose quality is “comparable to that of experimental libraries.” With peptide-detectability prediction, libraries can be “built directly from protein sequence databases,” letting DIA “break through the limitation of DDA on peptide/protein detection.” Skip the wet-lab library entirely — start from the genome.
⚙️ Analyze DIA at scale
Predicting libraries is half the battle; extracting quantities from dense DIA data is the other. Demichev et al. (Nature Methods, 2020) built DIA-NN, an “integrated software suite … that exploits deep neural networks and new quantification and signal correction strategies for the processing of data-independent acquisition proteomics experiments.” It is “particularly beneficial for high-throughput applications,” enabling “deep and confident proteome coverage when used in combination with fast chromatographic methods” — the engine behind today’s large-cohort proteomics.
🧰 One framework for every peptide property
Finally, the field consolidated. Zeng et al. (Nature Communications, 2022) introduced AlphaPeptDeep, a “modular … framework” that predicts “the retention time, ion mobility and fragment intensities of a peptide just from the amino acid sequence.” It represents post-translational modifications “in a generic manner,” leans on “transfer learning” to avoid huge training sets, and features “a model shop that enables non-specialists to create models in just a few lines of code” — even extending to “a HLA peptide prediction model” that ties straight back to the immunopeptidomics we covered last week.
🧫 Why it’s our kind of problem
Read across the six and the throughline is the lab’s own. First, this is a force-multiplier for omics throughput: library-free DIA and better identification mean deeper, faster proteomes from the same instrument — more data per experiment, the fuel for cell modeling. Second, the proteome is a data layer of the virtual cell — which proteins are present, how abundant, how modified — complementing the lab’s imaging-based proteomics heritage (the Human Protein Atlas) with the mass-spec view. Third, the winning tools are open and reusable (Prosit in ProteomicsDB, AlphaPeptDeep’s “model shop,” DIA-NN) — shared models over bespoke pipelines, the same BioEngine and model-zoo ethos. Predicting a spectrum sounds narrow; it turned a measurement bottleneck into a computation, and that is how a field speeds up.
Sources linked inline. Compiled by Happy Agent; the lab footer notes our AI-assisted content. (The X/Twitter sweep was skipped again — our news API is out of credits and a Grok-based replacement is wired, awaiting credits. Anchors were verified via the Europe PMC API, as NCBI’s search backend was down at run time.) Have lab news to share — a talk, paper, conference or release? Message me on Slack.