
Anthropic, the world's most valuable AI company, has made a groundbreaking discovery in mechanistic interpretability, shedding light on the inner workings of its AI models. This breakthrough has significant implications for the AI industry, developers, and businesses. The company's research has the potential to revolutionize the way we understand and interact with AI systems.

LongMedBench is a new benchmark for evaluating medical agents in long-horizon clinical decision-making. It provides a realistic assessment of AI models in medical care, emphasizing longitudinal interactions and multi-session decision-making. This benchmark has significant implications for the development of more accurate and reliable medical AI systems.

DeepSeek's new Director system accelerates distributed MoE serving via online proactive expert placement, reducing end-to-end latency by 11-55%. This breakthrough has significant implications for the AI industry, enabling faster and more efficient model serving. The Director system uses prediction-driven expert placement and online migration to minimize downtime and optimize performance.

Google has released a new model called Graph-Regularized Agentic Context Evolution (GRACE) to improve the reliability of long-horizon agentic context evolution under distribution shift. This model maintains the persistent instruction component as a typed semantic graph and validates proposed updates within the local typed neighborhoods of modified nodes. The results show a significant improvement in strict reliability compared to the baseline models.

HuggingFace has released a new AI model, Qwonkeau-v0.2-0.9B, which is a fully linearized version of Qwen3.5-0.8B with RWKV-7 and MesaNet layers. The model has 0.9B parameters and is available for use on the HuggingFace platform. This release is a significant development in the field of natural language processing and has the potential to improve the performance of various AI applications. The Qwonkeau-v0.2-0.9B model is part of the Qwonkeau collection, which includes multiple models with different architectures and parameters.

Anthropic has unveiled groundbreaking research detailing its ability to 'read' the internal states, or 'thoughts,' of its Claude AI models. This pivotal study reveals the existence of a 'global workspace' within LLMs, offering unprecedented insights into their complex decision-making processes and significantly advancing the field of AI interpretability.

OpenAI's GPT-5.6 Sol Ultra has solved the 50-year-old Cycle Double Cover Conjecture, a fundamental problem in graph theory. The proof was generated in under an hour using 64 subagents working in parallel.

OpenAI has officially launched its highly anticipated new family of models, spearheaded by GPT-5.6, marking a significant leap forward in generative AI capabilities. This release promises substantial improvements across diverse areas, including a critical focus on bolstering cybersecurity applications and overall model safety.

OpenAI has launched GPT-5.6, a new family of models that promises to deliver more intelligence from every token, stronger performance per dollar, and more capability on demand. The models have been trained to get more useful work from every token and have achieved state-of-the-art results across various fields. GPT-5.6 sets a new standard for both intelligence and efficiency, outperforming previous and competing frontier models with fewer tokens and at lower estimated cost.

OpenAI introduces GPT-5.6, a groundbreaking AI model that sets new standards for intelligence and efficiency. This model achieves state-of-the-art results in various fields, outperforming previous models at lower costs. GPT-5.6 is available in three variants: Sol, Terra, and Luna, catering to different needs and budgets.

Google researchers introduce HCC-STAR, a clinical-reasoning LLM for risk stratification and treatment guidance in hepatocellular carcinoma. This model achieves state-of-the-art performance in treatment recommendation and risk stratification. The study demonstrates the potential of AI in precision therapy for HCC patients.
US lawmakers are investigating the growing use of Chinese AI models by American companies, citing concerns over censorship, security risks, and the impact on domestic alternatives. The probe is specifically looking at companies such as Cursor and Airbnb, and the use of models like DeepSeek. This investigation highlights the complexities of the AI landscape and the need for careful consideration of the origins and implications of AI models.

A new arXiv paper by Alibaba researchers details a ReAct-style agentic setup integrating Large Language Models with SageMath, a powerful Computer Algebra System. This novel approach demonstrates substantial performance gains across frontier LLMs in solving research-level mathematical problems, significantly narrowing the capability gap between open-weight and closed models and paving the way for automated conjecture discovery.

OpenAI CEO suggests that video games can be a superior training data source than the internet for achieving AGI, with significant implications for the AI industry.

New research reveals a critical vulnerability in advanced reasoning AI models, where logically inconsistent prompts can force them into 'overthinking,' leading to denial-of-service attacks. This 'Evolutionary Prompt Attack' significantly increases resource consumption and poses a serious threat to commercial LLM providers like OpenAI, Google, and DeepSeek.

A groundbreaking Arxiv paper introduces LLMForge, a multi-model text-to-CAD framework that enables automatic generation of parametric 3D mechanical designs from natural language. This framework, featuring innovative iterative refinement and VLM-based critique, demonstrates remarkable success, with top models like DeepSeek-V3.2 achieving near-perfect mesh generation and showing compact models can rival larger systems.

Alibaba's Qwen team has unveiled a novel Reinforcement Learning (RL) approach, RLVR, designed to significantly enhance data-efficient code-switched Automatic Speech Recognition (ASR). This method uses verifiable rewards and a two-pass refinement process to adapt audio-language models, achieving state-of-the-art performance with just 10% of the data typically required.

OpenAI's GPT-5.6 model is expected to launch publicly in a matter of days, marking a significant milestone in AI development. Despite the impending release, OpenAI is still holding back on certain features. The launch is anticipated to have far-reaching implications for the AI industry and beyond.

Anthropic releases its most powerful public AI model, Claude Fable 5, marking a significant milestone in AI development. This model is designed to provide more accurate and informative responses. The release is expected to have a substantial impact on the AI industry and its applications.

Meituan has released LongCat-2.0, a 1.6T-parameter open MoE model with native 1M context and LongCat Sparse Attention. The model is designed for agentic coding and has been trained on over 35 trillion tokens. It aims to provide reliable and efficient code understanding, generation, and execution inside agent workflows.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.