
Meta researchers have introduced a novel neuro-symbolic agentic framework to significantly enhance the reasoning capabilities of Small Language Models (SLMs) like Gemma and Llama 3.2. This approach leverages knowledge graph grounding to overcome SLMs' historical struggles with complex, multi-hop logical tasks, offering a sustainable alternative to costly LLMs.

Groundbreaking Arxiv research reveals how large language model safety mechanisms are encoded and can be bypassed, introducing novel 'Activation-Guided' adversarial attacks. The study finds safety representations are distributed across model layers, not localized, and proposes a 33x faster attack method, Soft-GCG, offering critical insights for designing more robust AI alignment strategies.

Anthropic has unveiled groundbreaking research detailing its ability to 'read' the internal states, or 'thoughts,' of its Claude AI models. This pivotal study reveals the existence of a 'global workspace' within LLMs, offering unprecedented insights into their complex decision-making processes and significantly advancing the field of AI interpretability.

DeepSeek has unveiled a groundbreaking approach to abstract reasoning on ARC-AGI-1, leveraging an open-weight model (DeepSeek V3.2) in a 'non-thinking' mode, augmented by innovative agentic harnesses. This method achieves impressive generalization and pattern discovery, reaching up to 67.25% pass@2 with unprecedented cost-efficiency, sidestepping heavy compute or benchmark-specific fine-tuning.

DeepSeek introduces FirstResearch, a groundbreaking framework that tackles the auditability challenge in LLM-driven scientific discovery. By generating a structured 'Research Question Certificate,' FirstResearch ensures AI-proposed research questions are transparent, inspectable, and based on explicit mechanisms and assumptions, significantly enhancing trust in AI-powered scientific ideation.

Meta AI introduces ReContext, a groundbreaking training-free inference method that significantly boosts Large Language Model (LLM) performance on long contexts. By recursively replaying relevant evidence, ReContext enhances effective context utilization, bridging the gap between vast context windows and accurate reasoning without requiring retraining or external memory. This innovation promises to unlock more reliable and powerful LLM applications across industries.

Researchers from Anthropic have introduced a new diagnostic to evaluate the physics literacy of large language models (LLMs) in parallel physical worlds. The study tested three LLMs, including Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro, and found significant gaps in their ability to reason about unfamiliar physics frameworks. The results have important implications for the development and application of LLMs in scientific and technical domains.

Grok 4.5, a base model with 1.5 trillion parameters, has been further trained on Cursor data and is currently in beta testing at SpaceX and Tesla. This development marks a significant milestone in AI research and its applications in the tech industry. The model's capabilities and potential uses are being explored by these industry leaders.

New research from arXiv challenges conventional wisdom on AI improvement from feedback, revealing that multi-turn gains often mask true learning. The study highlights that an AI model's ability to effectively *utilize* feedback, rather than merely receiving it, is the critical bottleneck for interactive improvement, especially when compared to unguided self-refinement or simple retries.

The 246th LWiAI podcast breaks down Google’s Gemini 3.5 flash model, the multimodal Gemini Omni video engine, Elon Musk’s lost lawsuit, and OpenAI’s breakthrough on an 80‑year‑old Erdős geometry problem. We unpack the technical specs, market ripples, and what developers should watch next.
OpenAI introduces ATHENA-R1, an AI agent for treatment reasoning that outperforms language models and tool-use systems. Trained on 212 biomedical tools, ATHENA-R1 achieves 94.7% accuracy on open-ended drug reasoning and 82.9% on treatment reasoning. This breakthrough has significant implications for the healthcare industry and AI research.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.