
Anthropic has released version 0.115.0 of its SDK for Python, introducing new features such as support for Managed Agents event delta streaming and agent overrides. This update aims to enhance the functionality and usability of the Anthropics SDK, providing developers with more tools to work with AI models. The release is part of Anthropic's ongoing efforts to improve its offerings and stay competitive in the AI market.

DeepSeek's new study explores the effectiveness of learned stopping in reasoning models, finding that it can improve performance in certain tasks. The study introduces LearnStop, a hidden-state-free checkpoint stopper, and evaluates its performance across 18 task-model settings. The results show that learned stopping can be useful in tasks where many questions become correct before full budget but do not exhibit a single reliable scalar stopping signal.

New research from arXiv challenges conventional wisdom on AI improvement from feedback, revealing that multi-turn gains often mask true learning. The study highlights that an AI model's ability to effectively *utilize* feedback, rather than merely receiving it, is the critical bottleneck for interactive improvement, especially when compared to unguided self-refinement or simple retries.

A revolutionary method called Poller leverages large language models to evaluate poetry understanding with near-human accuracy, reducing errors by up to 94.55% in specific dimensions. This AI advancement bridges automation and human expertise in literary analysis.

OpenAI's latest research demonstrates how large language models can automate training data labeling for entity matching, reducing manual effort by 99% and slashing costs. This breakthrough enables faster, cheaper AI deployment for businesses.

Cohere's study reveals how transformer models develop situation modeling and mentalizing capabilities through training stages. Key findings show FBT performance depends on model size, training volume, and post-training methods, but remains fragile in complex scenarios.

MedEvoEval introduces a groundbreaking framework for evaluating AI doctor agents in simulated clinical settings. By tracking cross-episode learning and decision-making, it addresses critical gaps in medical AI evaluation. This tool enables developers to measure knowledge retention, resource allocation, and behavioral adaptation over time.

Meta researchers achieved 87.69% accuracy in predicting primary ICD-10 diagnosis categories by combining frozen medical LLM representations with multimodal EHR data. Their approach outperformed existing models and demonstrated strong cross-dataset adaptability.
Microsoft secretly embedded 'AARD code' in Windows 3.1 betas to sabotage DR DOS, triggering fake errors. This led to a $280M settlement with Caldera, Inc., revealing anti-competitive tactics in tech's past.

Emily Bender addresses misconceptions about her 2021 paper on 'stochastic parrots' and large language models, explaining their limitations and impact on AI discourse.

Anthropic unveils Claude Sonnet 5, a groundbreaking AI model offering enhanced agentic capabilities at lower costs. Targeting developers and businesses, the model aims to outperform competitors like GPT-5.5 and Gemini Pro while reducing operational expenses.

For two years Europe has been fixated on catching up in the model race, but the real edge may come from how companies integrate AI into their workflows. This article explores why architecture, not sheer size, will define Europe’s AI future.

The 246th LWiAI podcast breaks down Google’s Gemini 3.5 flash model, the multimodal Gemini Omni video engine, Elon Musk’s lost lawsuit, and OpenAI’s breakthrough on an 80‑year‑old Erdős geometry problem. We unpack the technical specs, market ripples, and what developers should watch next.

A new arXiv study shows OpenEvidence's specialized clinical tool beats top general‑purpose models (Claude Opus 4.8, Gemini 3.1 Pro, GPT‑5.5) on 620 real‑world point‑of‑care questions. Physicians across 30 specialties rated the specialized tool higher on accuracy, utility, source quality, verifiability and completeness.
OpenAI introduces ATHENA-R1, an AI agent for treatment reasoning that outperforms language models and tool-use systems. Trained on 212 biomedical tools, ATHENA-R1 achieves 94.7% accuracy on open-ended drug reasoning and 82.9% on treatment reasoning. This breakthrough has significant implications for the healthcare industry and AI research.
OpenAI introduces IMCBench, a benchmark for multimodal large language models in image-grounded medical conversations, evaluating safety, accuracy, and uncertainty in diagnosis. The benchmark tests eight models, with Claude Opus 4.6 achieving the highest overall score. The results highlight the need for multi-dimensional evaluation frameworks in medical AI.
Meta has introduced AnTenA, a novel AI system that leverages large language models to explain hidden patterns in human narratives. This system uses task-agnostic and task-specific prompts to analyze co-clustered latent patterns from tensor decomposition. AnTenA has the potential to revolutionize the field of explainable AI.
Researchers introduce a gravitational interpretation of fine-tuning reversion, explaining how AI models can revert to earlier behaviors. This phenomenon is caused by dominant behavioral manifolds created during early training phases. The study provides insights into the safety and stability of AI models, with significant implications for the AI industry.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.