What if you could just ask your microscope a question? That’s the promise of vision-language models — systems that connect a picture to words. BiomedCLIP learned the mapping from 15 million image-text pairs mined from 4.4 million scientific articles; LLaVA-Med turned figure captions into ‘a vision-language conversational assistant that can answer open-ended research questions of biomedical images,’ trained ‘in less than 15 hours.’ Then the reality check: on real microscopy, μ-Bench finds ‘current models struggle on all categories, even for basic tasks such as distinguishing microscopy modalities,’ and specialist fine-tuning can make things worse. CARES adds that medical vision-language models ‘consistently exhibit concerns regarding trustworthiness, often displaying factual inaccuracies.’ The gap between a fluent answer and a correct one is the whole story — and it’s exactly why the lab’s answer is a grounded, tool-calling assistant like the BioImage.IO Chatbot, not a model asked to know everything on its own.