
Google has released MyoCardBench, a comprehensive AI benchmark that evaluates large language models in clinically authentic cardiovascular care scenarios. The benchmark features 2,263 items from 13 task-specific datasets and has been tested by 16 cardiology physicians. Initial results show promising performance from top models, but significant gaps remain in certain areas.

Google's latest research reveals a breakthrough in using large language models (LLMs) to automatically detect critical inconsistencies in Electronic Health Records (EHRs), impacting nearly 70% of patient admissions. This formative study, leveraging Gemini 2.5 Pro and Flash, lays the groundwork for enhancing patient safety and clinical reasoning by identifying errors across diverse medical domains. While promising, the research also highlights key challenges in temporal reasoning and domain-specific knowledge that future AI solutions must overcome.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.