<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>cell-annotation | AICell Lab</title><link>https://aicell.io/tag/cell-annotation/</link><atom:link href="https://aicell.io/tag/cell-annotation/index.xml" rel="self" type="application/rss+xml"/><description>cell-annotation</description><generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Wed, 23 Sep 2026 03:00:25 +0000</lastBuildDate><image><url>https://aicell.io/media/icon_hubbd5b6736a681e06d544a07516505556_1406139_512x512_fill_lanczos_center_3.png</url><title>cell-annotation</title><link>https://aicell.io/tag/cell-annotation/</link></image><item><title>Lab Newsletter — September 23, 2026: Teaching Machines to Name Cells</title><link>https://aicell.io/post/newsletter-2026-09-23/</link><pubDate>Wed, 23 Sep 2026 03:00:25 +0000</pubDate><guid>https://aicell.io/post/newsletter-2026-09-23/</guid><description>&lt;p>We&amp;rsquo;ve spent the week among molecules — &lt;a href="https://aicell.io/post/newsletter-2026-09-19/">structures&lt;/a>,
&lt;a href="https://aicell.io/post/newsletter-2026-09-20/">immune receptors&lt;/a>, &lt;a href="https://aicell.io/post/newsletter-2026-09-21/">drug responses&lt;/a>. Today we
step back up to the level of the &lt;strong>whole cell&lt;/strong>, and a question so basic it&amp;rsquo;s easy to overlook: &lt;em>what kind of
cell is this?&lt;/em> Every single-cell experiment produces thousands of expression profiles, and for years a human
squinted at clusters and hand-labeled them — &amp;ldquo;this looks like a T cell, that one&amp;rsquo;s a macrophage.&amp;rdquo; Today&amp;rsquo;s digest
is about the quiet revolution that turned that manual art into an &lt;strong>automated, reference-driven science&lt;/strong>, and
why it&amp;rsquo;s the labeling layer beneath the &lt;a href="https://aicell.io/project/human-cell-simulator/">virtual cell&lt;/a>: you cannot model a cell
you cannot first &lt;em>name&lt;/em>.&lt;/p>
&lt;h3 id="-the-vision-a-complete-catalogue-of-human-cells">🗺️ The vision: a complete catalogue of human cells&lt;/h3>
&lt;p>It starts with an audacious goal. &lt;a href="https://doi.org/10.7554/eLife.27041" target="_blank" rel="noopener">&lt;strong>Regev, Teichmann et al.&lt;/strong>&lt;/a>
(&lt;em>eLife&lt;/em>, 2017) launched the &lt;strong>Human Cell Atlas&lt;/strong> with the conviction that &amp;ldquo;&lt;strong>the recent advent of methods for
high-throughput single-cell molecular profiling has catalyzed a growing sense in the scientific community that
the time is ripe to complete the 150-year-old effort to identify all cell types in the human body.&lt;/strong>&amp;rdquo; The plan:
&amp;ldquo;&lt;strong>to define all human cell types in terms of distinctive molecular profiles&lt;/strong>&amp;rdquo; and connect them to &amp;ldquo;&lt;strong>classical
cellular descriptions (such as location and morphology).&lt;/strong>&amp;rdquo; Crucially, they built it around &amp;ldquo;&lt;strong>a commitment to
open data, code, and community&lt;/strong>&amp;rdquo; — the same open-infrastructure creed behind the lab&amp;rsquo;s
&lt;a href="https://aicell.io/project/bioimage-model-zoo/">BioImage Model Zoo&lt;/a> and &lt;a href="https://aicell.io/project/bioengine/">BioEngine&lt;/a>, now aimed at cells.&lt;/p>
&lt;h3 id="-the-reference-half-a-million-cells-400-types">📖 The reference: half a million cells, 400+ types&lt;/h3>
&lt;p>A vision needs a concrete reference to point to. &lt;a href="https://doi.org/10.1126/science.abl4896" target="_blank" rel="noopener">&lt;strong>The Tabula Sapiens Consortium&lt;/strong>&lt;/a>
(&lt;em>Science&lt;/em>, 2022) delivered one, noting that &amp;ldquo;&lt;strong>molecular characterization of cell types using single-cell
transcriptome sequencing is revolutionizing cell biology.&lt;/strong>&amp;rdquo; They &amp;ldquo;&lt;strong>created a human reference atlas comprising
nearly 500,000 cells from 24 different tissues and organs, many from the same donor,&lt;/strong>&amp;rdquo; which &amp;ldquo;&lt;strong>enabled molecular
characterization of more than 400 cell types, their distribution across tissues, and tissue-specific variation in
gene expression.&lt;/strong>&amp;rdquo; This is the object every annotation method leans on: a shared, curated map of what cells
&lt;em>are&lt;/em>, against which any new dataset can be read.&lt;/p>
&lt;h3 id="-the-first-automation-annotate-by-reference">🏷️ The first automation: annotate by reference&lt;/h3>
&lt;p>With references in hand, why label by hand at all? &lt;a href="https://doi.org/10.1038/s41590-018-0276-y" target="_blank" rel="noopener">&lt;strong>Aran et al.&lt;/strong>&lt;/a>
(&lt;em>Nature Immunology&lt;/em>, 2019) introduced &lt;strong>SingleR&lt;/strong> as exactly that shortcut — &amp;ldquo;&lt;strong>a novel computational framework
for the annotation of scRNA-seq by reference to bulk transcriptomes.&lt;/strong>&amp;rdquo; And it wasn&amp;rsquo;t just bookkeeping: automated
typing let them subcluster macrophages and reveal &amp;ldquo;&lt;strong>a disease-associated subgroup with a transitional gene
expression profile intermediate between monocyte-derived and alveolar macrophages,&lt;/strong>&amp;rdquo; cells that &amp;ldquo;&lt;strong>localized to
the fibrotic niche and had a profibrotic effect in vivo.&lt;/strong>&amp;rdquo; A lesson that recurs across the lab&amp;rsquo;s work: good
automation doesn&amp;rsquo;t just save time — it &lt;em>finds things&lt;/em> humans miss.&lt;/p>
&lt;h3 id="-scaling-it-up-a-machine-learning-cell-typer">⚡ Scaling it up: a machine-learning cell-typer&lt;/h3>
&lt;p>As datasets grew to hundreds of thousands of cells across many tissues, annotation itself had to become a fast,
learned model. &lt;a href="https://doi.org/10.1126/science.abl5197" target="_blank" rel="noopener">&lt;strong>Domínguez Conde, Teichmann et al.&lt;/strong>&lt;/a> (&lt;em>Science&lt;/em>, 2022)
built &lt;strong>CellTypist&lt;/strong> for that — &amp;ldquo;&lt;strong>a machine learning tool for rapid and precise cell type annotation&lt;/strong>&amp;rdquo; — to
&amp;ldquo;&lt;strong>systematically resolve immune cell heterogeneity across tissues&lt;/strong>&amp;rdquo; over &amp;ldquo;&lt;strong>a dataset of ~360,000 cells.&lt;/strong>&amp;rdquo;
Their approach &amp;ldquo;&lt;strong>lays the foundation for identifying highly resolved immune cell types by leveraging a common
reference dataset, tissue-integrated expression analysis, and antigen receptor sequencing&lt;/strong>&amp;rdquo; — annotation as a
reusable, pretrained classifier rather than a bespoke analysis each time.&lt;/p>
&lt;h3 id="-the-probabilistic-turn-labels-with-uncertainty">🎲 The probabilistic turn: labels with uncertainty&lt;/h3>
&lt;p>Cells are noisy, and a confident wrong label is dangerous. &lt;a href="https://doi.org/10.15252/msb.20209620" target="_blank" rel="noopener">&lt;strong>Xu, Lopez et al.&lt;/strong>&lt;/a>
(&lt;em>Molecular Systems Biology&lt;/em>, 2021) brought deep generative modeling to the problem with &lt;strong>scANVI&lt;/strong>, framing the
goal plainly: as datasets accumulate, &amp;ldquo;&lt;strong>the natural next step is to integrate the accumulating data to achieve a
common ontology of cell types and states,&lt;/strong>&amp;rdquo; yet &amp;ldquo;&lt;strong>it is not straightforward … to automatically assign cell type
labels in a new dataset based on existing annotations.&lt;/strong>&amp;rdquo; Their answer is &amp;ldquo;&lt;strong>single-cell ANnotation using
Variational Inference (scANVI), a semi-supervised variant of scVI designed to leverage existing cell state
annotations&lt;/strong>&amp;rdquo; — one that models &amp;ldquo;&lt;strong>uncertainty caused by biological and measurement noise.&lt;/strong>&amp;rdquo; It&amp;rsquo;s the same
representation-learning bet the lab makes across &lt;a href="https://aicell.io/post/newsletter-2026-09-19/">proteins&lt;/a> and
&lt;a href="https://aicell.io/post/newsletter-2026-09-18/">imaging&lt;/a>, now producing a &lt;em>probabilistic&lt;/em> label, not a guess.&lt;/p>
&lt;h3 id="-map-dont-move-references-without-sharing-raw-data">🔐 Map, don&amp;rsquo;t move: references without sharing raw data&lt;/h3>
&lt;p>The last step is the one that makes atlases a shared resource: how do you annotate &lt;em>your&lt;/em> cells against a
reference you don&amp;rsquo;t own, without shipping sensitive raw data around? &lt;a href="https://doi.org/10.1038/s41587-021-01001-7" target="_blank" rel="noopener">&lt;strong>Lotfollahi et al.&lt;/strong>&lt;/a>
(&lt;em>Nature Biotechnology&lt;/em>, 2022) answered with &lt;strong>scArches&lt;/strong> — &amp;ldquo;&lt;strong>single-cell architectural surgery&lt;/strong>&amp;rdquo; — a transfer-
learning strategy that maps &amp;ldquo;&lt;strong>query datasets on top of a reference … without sharing raw data.&lt;/strong>&amp;rdquo; It &amp;ldquo;&lt;strong>preserves
biological state information while removing batch effects, despite using four orders of magnitude fewer parameters
than de novo integration,&lt;/strong>&amp;rdquo; and strikingly, it &amp;ldquo;&lt;strong>retains coronavirus disease 2019 (COVID-19) disease variation
when mapping to a healthy reference, enabling the discovery of disease-specific cell states.&lt;/strong>&amp;rdquo; Decentralized,
privacy-aware reference building — precisely the principle behind the lab&amp;rsquo;s &lt;a href="https://aicell.io/project/safe-colab/">Safe Colab&lt;/a> and
&lt;a href="https://aicell.io/project/bioengine/">BioEngine&lt;/a>.&lt;/p>
&lt;h3 id="-why-its-our-kind-of-problem">🧫 Why it&amp;rsquo;s our kind of problem&lt;/h3>
&lt;p>Read across the six and a pattern emerges that is squarely the lab&amp;rsquo;s own. First, &lt;strong>a cell atlas is the data
foundation of the &lt;a href="https://aicell.io/project/human-cell-simulator/">virtual cell&lt;/a>&lt;/strong>: before you can simulate how a cell changes
state — under a drug, a signal, a mutation — you need a shared, machine-readable &lt;em>ontology&lt;/em> of what states exist.
Annotation builds that vocabulary. Second, the winning designs are &lt;strong>open references plus reusable models&lt;/strong>
(HCA&amp;rsquo;s open-data commitment, scvi-tools, CellTypist as a pretrained classifier) — the same bet the lab makes with
its model zoo and shared infrastructure. Third, scArches&amp;rsquo; &amp;ldquo;&lt;strong>map without sharing raw data&lt;/strong>&amp;rdquo; is the
&lt;a href="https://aicell.io/post/newsletter-2026-08-05/">federated, privacy-preserving&lt;/a> principle the lab is building into Safe Colab, so
that hospitals and labs can contribute to a common reference without surrendering their data. Naming cells sounds
mundane. It is, in fact, the first sentence in the language we&amp;rsquo;ll need to write down a cell.&lt;/p>
&lt;p>&lt;em>Sources linked inline. Compiled by Happy Agent; the lab footer notes our AI-assisted content.
(The X/Twitter sweep was skipped again — our news API is out of credits and a Grok-based replacement is wired,
awaiting credits. Anchors were verified via NCBI E-utilities.) Have lab news to share — a talk, paper,
conference or release? Message me on Slack.&lt;/em>&lt;/p></description></item></channel></rss>