
DeepSeek's new AI model uses a unified multimodal learner for clinical prediction, simplifying the process and achieving state-of-the-art results. This approach converts all patient data into a single natural language sequence and fine-tunes a pretrained language model. The model outperforms task-specific multimodal baselines and a clinically deployed gradient boosting system.

A new study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries, identifying Unsupported Claims as the dominant failure mode. The study evaluates three language models, including GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash, and proposes a knowledge-graph-augmented retrieval system to improve faithfulness scores. This research has significant implications for the use of LLMs in high-stakes contexts such as healthcare.

Google researchers introduce HCC-STAR, a clinical-reasoning LLM for risk stratification and treatment guidance in hepatocellular carcinoma. This model achieves state-of-the-art performance in treatment recommendation and risk stratification. The study demonstrates the potential of AI in precision therapy for HCC patients.

MedEvoEval introduces a groundbreaking framework for evaluating AI doctor agents in simulated clinical settings. By tracking cross-episode learning and decision-making, it addresses critical gaps in medical AI evaluation. This tool enables developers to measure knowledge retention, resource allocation, and behavioral adaptation over time.

Meta researchers achieved 87.69% accuracy in predicting primary ICD-10 diagnosis categories by combining frozen medical LLM representations with multimodal EHR data. Their approach outperformed existing models and demonstrated strong cross-dataset adaptability.

A new arXiv study shows OpenEvidence's specialized clinical tool beats top general‑purpose models (Claude Opus 4.8, Gemini 3.1 Pro, GPT‑5.5) on 620 real‑world point‑of‑care questions. Physicians across 30 specialties rated the specialized tool higher on accuracy, utility, source quality, verifiability and completeness.
OpenAI introduces ATHENA-R1, an AI agent for treatment reasoning that outperforms language models and tool-use systems. Trained on 212 biomedical tools, ATHENA-R1 achieves 94.7% accuracy on open-ended drug reasoning and 82.9% on treatment reasoning. This breakthrough has significant implications for the healthcare industry and AI research.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.