
DeepSeek's new study explores the effectiveness of learned stopping in reasoning models, finding that it can improve performance in certain tasks. The study introduces LearnStop, a hidden-state-free checkpoint stopper, and evaluates its performance across 18 task-model settings. The results show that learned stopping can be useful in tasks where many questions become correct before full budget but do not exhibit a single reliable scalar stopping signal.

New research from arXiv challenges conventional wisdom on AI improvement from feedback, revealing that multi-turn gains often mask true learning. The study highlights that an AI model's ability to effectively *utilize* feedback, rather than merely receiving it, is the critical bottleneck for interactive improvement, especially when compared to unguided self-refinement or simple retries.

NVIDIA's full-stack inference software, optimized for the Blackwell platform, slashes token costs by 5x on DeepSeek V4. Companies like Baseten and Cognition leverage NVIDIA's tools to scale AI workloads efficiently, marking a shift from hardware specs to cost-per-token economics.

NVIDIA's BioNeMo Agent Toolkit now powers Anthropic's Claude Science, enabling researchers to accelerate drug discovery and genomic analysis with natural language workflows. This collaboration marries NVIDIA's GPU computing with Claude Science's AI agents for faster scientific innovation.

Podcasting platform Riverside has expanded into newsletter publishing with an AI-driven feature that transforms audio content into written newsletters. This move aims to streamline content repurposing for creators, leveraging AI to summarize podcasts, generate summaries, and publish cross-platform newsletters.

A revolutionary method called Poller leverages large language models to evaluate poetry understanding with near-human accuracy, reducing errors by up to 94.55% in specific dimensions. This AI advancement bridges automation and human expertise in literary analysis.

Saltroad, a clinician‑led speech and language therapy provider, secured £1.5 million in funding and bought AI documentation platform Ogma to streamline therapy for children. The move aims to cut admin, standardise notes, and expand access across the UK.

OpenAI's latest research demonstrates how large language models can automate training data labeling for entity matching, reducing manual effort by 99% and slashing costs. This breakthrough enables faster, cheaper AI deployment for businesses.

Cohere's study reveals how transformer models develop situation modeling and mentalizing capabilities through training stages. Key findings show FBT performance depends on model size, training volume, and post-training methods, but remains fragile in complex scenarios.

A groundbreaking AI system combines time-series forecasting, anomaly detection, and LLM-driven analysis to deliver actionable energy insights. This end-to-end solution reduces alert noise for facility managers while maintaining high accuracy across 16 real-world scenarios.

MedEvoEval introduces a groundbreaking framework for evaluating AI doctor agents in simulated clinical settings. By tracking cross-episode learning and decision-making, it addresses critical gaps in medical AI evaluation. This tool enables developers to measure knowledge retention, resource allocation, and behavioral adaptation over time.

Meta researchers achieved 87.69% accuracy in predicting primary ICD-10 diagnosis categories by combining frozen medical LLM representations with multimodal EHR data. Their approach outperformed existing models and demonstrated strong cross-dataset adaptability.

A groundbreaking study reveals that traditional safety methods for AI agents are fundamentally flawed. Instead of relying on refusal-based content safety, the paper advocates for action alignment and least privilege enforcement to ensure secure, user-intent-driven AI systems.
Microsoft secretly embedded 'AARD code' in Windows 3.1 betas to sabotage DR DOS, triggering fake errors. This led to a $280M settlement with Caldera, Inc., revealing anti-competitive tactics in tech's past.

Emily Bender addresses misconceptions about her 2021 paper on 'stochastic parrots' and large language models, explaining their limitations and impact on AI discourse.

A groundbreaking MIT study reveals how labeling AI agents as 'employees' leads to worse human oversight, as companies like OpenAI push agentic AI tools. New research shows a 18% drop in error detection when AI work is framed as coming from 'digital coworkers.'

OpenAI and Sceye lead groundbreaking advancements in AI collaboration and stratospheric internet. Discover how these innovations are reshaping workplaces and global connectivity.

Billions flood longevity research as scientists explore AI-driven cellular reprogramming to reverse aging. MIT’s Roundtable unpacks the science, funding, and ethical dilemmas behind this cutting-edge field.

Anthropic unveils Claude Sonnet 5, a groundbreaking AI model offering enhanced agentic capabilities at lower costs. Targeting developers and businesses, the model aims to outperform competitors like GPT-5.5 and Gemini Pro while reducing operational expenses.

Qualcomm's acquisition of Modular and SambaNova's $10B valuation reveal a seismic shift in AI. Dave Munichiello of GV explains how software is becoming as valuable as silicon in an era of hardware scarcity.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.