
Google has released MyoCardBench, a comprehensive AI benchmark that evaluates large language models in clinically authentic cardiovascular care scenarios. The benchmark features 2,263 items from 13 task-specific datasets and has been tested by 16 cardiology physicians. Initial results show promising performance from top models, but significant gaps remain in certain areas.

MedLoCoMo is a new medical dialogue benchmark for large language models, testing their ability to reason over longitudinal patient histories. The benchmark contains 100 patient timelines with an average of 1,669.8 turns, 29.7 sessions, and 74,512.2 tokens per conversation. This release has significant implications for the development of more accurate and reliable medical AI systems.

Google's latest research reveals a breakthrough in using large language models (LLMs) to automatically detect critical inconsistencies in Electronic Health Records (EHRs), impacting nearly 70% of patient admissions. This formative study, leveraging Gemini 2.5 Pro and Flash, lays the groundwork for enhancing patient safety and clinical reasoning by identifying errors across diverse medical domains. While promising, the research also highlights key challenges in temporal reasoning and domain-specific knowledge that future AI solutions must overcome.

Meta has released a new AI model that combines reinforcement learning with large language models to create a more transparent and reliable insulin pump controller for Type 1 Diabetes patients. The model, called LLM-T1D, has shown promising results in blood sugar control and safety verification. This breakthrough has the potential to revolutionize the treatment of Type 1 Diabetes and improve the lives of millions of people worldwide.

Researchers from Anthropic have introduced a new diagnostic to evaluate the physics literacy of large language models (LLMs) in parallel physical worlds. The study tested three LLMs, including Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro, and found significant gaps in their ability to reason about unfamiliar physics frameworks. The results have important implications for the development and application of LLMs in scientific and technical domains.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.