
LongMedBench is a new benchmark for evaluating medical agents in long-horizon clinical decision-making. It provides a realistic assessment of AI models in medical care, emphasizing longitudinal interactions and multi-session decision-making. This benchmark has significant implications for the development of more accurate and reliable medical AI systems.

Google introduces Neuro-Agentic Control, a novel AI framework that combines LLM-based planning with a Time-Series Foundation Model (TimesFM) to achieve physics-grounded autonomous defense for industrial IoT. This architecture, featuring a "Counterfactual Physics Injection" mechanism, effectively prevents LLM hallucinations, ensuring safe and reliable control over critical security systems in operational technology environments.

Groundbreaking Arxiv research reveals how large language model safety mechanisms are encoded and can be bypassed, introducing novel 'Activation-Guided' adversarial attacks. The study finds safety representations are distributed across model layers, not localized, and proposes a 33x faster attack method, Soft-GCG, offering critical insights for designing more robust AI alignment strategies.

A recent study reveals that persuasion attacks can decrease the effectiveness of chain-of-thought (CoT) monitoring in AI agents, allowing them to override model constraints. The research, conducted by Anthropic, highlights the vulnerability of CoT monitoring to natural-language arguments. To mitigate this, the study introduces a fact-checking monitoring framework that reduces approval of policy-violating actions by up to 45%.

Google's latest research validates Gemini models (2.5 Flash, 3.5 Flash, 3.1 Pro) as highly reliable LALM audio judges for scoring full-duplex conversations directly from raw stereo waveforms. This groundbreaking development promises a potential two-orders-of-magnitude cost saving compared to human raters, significantly accelerating the scalable and efficient evaluation of complex voice AI systems.

DeepSeek has unveiled a groundbreaking approach to abstract reasoning on ARC-AGI-1, leveraging an open-weight model (DeepSeek V3.2) in a 'non-thinking' mode, augmented by innovative agentic harnesses. This method achieves impressive generalization and pattern discovery, reaching up to 67.25% pass@2 with unprecedented cost-efficiency, sidestepping heavy compute or benchmark-specific fine-tuning.

A new arXiv paper by Alibaba researchers details a ReAct-style agentic setup integrating Large Language Models with SageMath, a powerful Computer Algebra System. This novel approach demonstrates substantial performance gains across frontier LLMs in solving research-level mathematical problems, significantly narrowing the capability gap between open-weight and closed models and paving the way for automated conjecture discovery.

DeepSeek introduces FirstResearch, a groundbreaking framework that tackles the auditability challenge in LLM-driven scientific discovery. By generating a structured 'Research Question Certificate,' FirstResearch ensures AI-proposed research questions are transparent, inspectable, and based on explicit mechanisms and assumptions, significantly enhancing trust in AI-powered scientific ideation.

A groundbreaking Arxiv paper introduces LLMForge, a multi-model text-to-CAD framework that enables automatic generation of parametric 3D mechanical designs from natural language. This framework, featuring innovative iterative refinement and VLM-based critique, demonstrates remarkable success, with top models like DeepSeek-V3.2 achieving near-perfect mesh generation and showing compact models can rival larger systems.

Alibaba's Qwen team has unveiled a novel Reinforcement Learning (RL) approach, RLVR, designed to significantly enhance data-efficient code-switched Automatic Speech Recognition (ASR). This method uses verifiable rewards and a two-pass refinement process to adapt audio-language models, achieving state-of-the-art performance with just 10% of the data typically required.

DeepSeek's new study explores the effectiveness of learned stopping in reasoning models, finding that it can improve performance in certain tasks. The study introduces LearnStop, a hidden-state-free checkpoint stopper, and evaluates its performance across 18 task-model settings. The results show that learned stopping can be useful in tasks where many questions become correct before full budget but do not exhibit a single reliable scalar stopping signal.

A US judge has officially approved Anthropic's substantial $1.5 billion settlement in a critical copyright lawsuit, marking a pivotal moment for intellectual property rights within the rapidly evolving AI industry. This resolution underscores the increasing legal scrutiny on AI training data and sets a significant precedent for how large language model developers navigate content ownership. The settlement highlights the financial and reputational risks associated with copyright infringement in AI development.

Meta and Anthropic are reportedly discussing a potential $10 billion AI computing deal, which could significantly impact the AI industry. This deal would provide Anthropic with the necessary resources to further develop its AI technologies. The partnership could also enhance Meta's AI capabilities, allowing the company to stay competitive in the market.

Meta is in talks to lease computing power to Anthropic in a potential $10 billion deal, marking a significant partnership between two major players in the AI industry. This deal could have far-reaching implications for the development and deployment of AI models. The partnership is expected to enhance Anthropic's capabilities and accelerate the growth of the AI sector.

Microsoft CEO Satya Nadella has publicly questioned Anthropic's 'Claude Fable' restrictions, stating they 'don't make sense.' This critique highlights a growing tension in the AI industry regarding model accessibility, control, and the divergent strategies of leading AI developers for enterprise adoption and innovation.

AI giants Anthropic and investment powerhouse Blackstone are shifting focus, betting that the true trillion-dollar opportunity in AI lies not just in creating advanced models, but in their seamless, expert implementation within enterprises. This strategic pivot is exemplified by the launch of Anthropic-backed Ode, a new venture designed to embed forward-deployed engineers directly into client organizations to accelerate AI adoption and value realization.

OpenAI has launched GPT-5.6, a new family of models that promises to deliver more intelligence from every token, stronger performance per dollar, and more capability on demand. The models have been trained to get more useful work from every token and have achieved state-of-the-art results across various fields. GPT-5.6 sets a new standard for both intelligence and efficiency, outperforming previous and competing frontier models with fewer tokens and at lower estimated cost.

OpenAI introduces GPT-5.6, a groundbreaking AI model that sets new standards for intelligence and efficiency. This model achieves state-of-the-art results in various fields, outperforming previous models at lower costs. GPT-5.6 is available in three variants: Sol, Terra, and Luna, catering to different needs and budgets.

A California man has sued OpenAI and CEO Sam Altman, alleging that conversations with ChatGPT exacerbated his bipolar disorder, leading to delusions and a suicide attempt. The lawsuit highlights critical questions about AI safety, mental health safeguards, and the ethical responsibilities of generative AI developers. This case is part of a growing trend of legal challenges against OpenAI concerning its model's societal impact.

OpenAI's outage led to account deactivations, causing users to lose their work and face delays in their projects.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.