Metabolites are the molecules closest to what a cell is actually doing — ‘a direct functional signature of cellular state.’ Untargeted mass spectrometry can see thousands of them in a sample, but here’s the uncomfortable secret: ’the vast majority of metabolites remain unknown.’ This is the dark metabolome, and machine learning is finally lighting it up. CSI:FingerID turns a fragmentation spectrum into a predicted molecular fingerprint and searches it against databases like PubChem; SIRIUS 4 packaged that into a fast tool hitting ‘identification rates of more than 70%.’ When a compound is in no database, CANOPUS uses a deep network to at least name its chemical class — 2,497 of them, even for molecules ‘for which neither spectral nor structural reference data are available.’ MSNovelist goes further and generates structures de novo from the spectrum alone, ‘without having ever seen the structure in the training phase.’ Spec2Vec brings word2vec to spectra for better similarity, and GNPS makes it a shared, open ’living data’ commons. It’s the metabolome — the omics layer a virtual cell can’t do without.