
Google introduces Just Keep Prompting, a framework to evaluate Vision-Language Models under sustained conversational pressure, revealing instability in models like GPT-4o, Gemini 2.5 Pro, and Qwen3-VL-30B. The study highlights the importance of assessing VLMs' epistemic stability in real-world settings. The findings have significant implications for the development and deployment of VLMs in various applications.
Google DeepMind's Demis Hassabis is calling for a US-led AI standards body to review frontier models for national security risks. The proposed body would be a federally overseen public-private organization, initially voluntary and eventually mandatory for US deployment. This move aims to address risks associated with artificial general intelligence, including cybersecurity and biological threats.

Google introduces Neuro-Agentic Control, a novel AI framework that combines LLM-based planning with a Time-Series Foundation Model (TimesFM) to achieve physics-grounded autonomous defense for industrial IoT. This architecture, featuring a "Counterfactual Physics Injection" mechanism, effectively prevents LLM hallucinations, ensuring safe and reliable control over critical security systems in operational technology environments.

Google has released a new model called Graph-Regularized Agentic Context Evolution (GRACE) to improve the reliability of long-horizon agentic context evolution under distribution shift. This model maintains the persistent instruction component as a typed semantic graph and validates proposed updates within the local typed neighborhoods of modified nodes. The results show a significant improvement in strict reliability compared to the baseline models.

Google researchers introduce HCC-STAR, a clinical-reasoning LLM for risk stratification and treatment guidance in hepatocellular carcinoma. This model achieves state-of-the-art performance in treatment recommendation and risk stratification. The study demonstrates the potential of AI in precision therapy for HCC patients.

Google's latest research validates Gemini models (2.5 Flash, 3.5 Flash, 3.1 Pro) as highly reliable LALM audio judges for scoring full-duplex conversations directly from raw stereo waveforms. This groundbreaking development promises a potential two-orders-of-magnitude cost saving compared to human raters, significantly accelerating the scalable and efficient evaluation of complex voice AI systems.

New research reveals a critical vulnerability in advanced reasoning AI models, where logically inconsistent prompts can force them into 'overthinking,' leading to denial-of-service attacks. This 'Evolutionary Prompt Attack' significantly increases resource consumption and poses a serious threat to commercial LLM providers like OpenAI, Google, and DeepSeek.

Fast Company's collection of 13 stories showcases American ingenuity, from space flight to Google DeepMind, highlighting the country's innovative spirit. The stories feature iconic figures like Steve Jobs and companies like Disney, Patagonia, and Reddit. This compilation celebrates the 250th anniversary of the signing of the Declaration of Independence, demonstrating America's long history of innovation and its continued impact on the world.

Anthropic has unveiled Claude Science, a flagship AI product engineered to transform scientific research, particularly in computational biology and drug development. Positioned alongside Claude Code, this standalone offering empowers researchers to autonomously carry out complex tasks, marking a significant strategic leap for Anthropic into the life sciences domain and intensifying competition in AI for scientific discovery.

Emily Bender addresses misconceptions about her 2021 paper on 'stochastic parrots' and large language models, explaining their limitations and impact on AI discourse.

Qualcomm's acquisition of Modular and SambaNova's $10B valuation reveal a seismic shift in AI. Dave Munichiello of GV explains how software is becoming as valuable as silicon in an era of hardware scarcity.

The 246th LWiAI podcast breaks down Google’s Gemini 3.5 flash model, the multimodal Gemini Omni video engine, Elon Musk’s lost lawsuit, and OpenAI’s breakthrough on an 80‑year‑old Erdős geometry problem. We unpack the technical specs, market ripples, and what developers should watch next.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.