MedGemma is a collection of open vision-language foundation models tuned for medical images and text, built on the Gemma 3 architecture. The suite includes multimodal MedGemma 4B and 27B models and a MedGemma 27B Text variant, plus a specialized 400M-parameter medical image encoder, MedSigLIP, trained on millions of medical image-text pairs. Models accept images, text, or both and were optimized to retain Gemma 3’s general-language capabilities while gaining medical domain understanding. The release includes examples of fine-tuning for chest X-rays, histopathology, dermatology and electronic health record question answering, emphasizing practical downstream adaptation.
MedGemma delivers consistent gains over the Gemma 3 base and outperforms or matches larger proprietary models on multiple benchmarks. Out-of-distribution improvements include 2.6-10% in medical image question answering, 15.5-18.1% in chest X‑ray finding classification and a 10.8% increase in agentic evaluations. MedSigLIP provides data-efficient and zero-shot image classification and retrieval with performance comparable to many specialized medical encoders. Fine-tuning MedGemma proves especially effective with limited labeled data, producing strong results across MedQA, MedMCQA, PubMedQA, MMLU Med, AfriMed‑QA and AgentClinic. The open models are presented as a foundation for accelerating medical AI research and downstream application development.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.