
OpenAI's latest study explores the necessity of executable world models, simplification, and verification in coding agents, revealing surprising results. The research evaluates four nested Codex-based agents, finding that every agent variant improves with stronger models and greater reasoning effort. The study's findings have significant implications for the development of Artificial General Intelligence (AGI).

The AI landscape is rapidly evolving as three Chinese labs—Moonshot AI, DeepSeek, and Zhipu AI—release powerful, open-weight Mixture-of-Experts (MoE) models. Kimi K3, DeepSeek V4 Pro, and GLM-5.2 are pushing the boundaries of scale and capability, offering trillion-scale parameters and million-token context windows for complex coding and agent workloads, fundamentally shifting the open-source AI leaderboard.

A new demo showcases an AI event venue operator built with MongoDB Atlas, Voyage, and LangGraph, enabling agents to remember prior events and respond to live operational changes. This technology has significant implications for the event management industry, particularly for high-stakes events like tennis tournaments. The demo highlights the potential for AI to improve event operations and enhance the fan experience.

Smartsheet has developed a pioneering remote Model Context Protocol (MCP) server on AWS, enabling AI clients like Claude Desktop and Amazon Quick to securely access and interact with enterprise data. This innovative solution optimizes AI interactions, significantly reduces token costs, and enhances the reliability of AI agents operating within Smartsheet's platform. It marks a significant step towards seamless AI integration in enterprise work management.

Amazon Bedrock has announced the general availability of its Managed Knowledge Base, a fully managed solution designed to simplify the creation of enterprise search capabilities for generative AI agents. This innovation dramatically reduces the complexity and time required to build robust Retrieval Augmented Generation (RAG) systems, enabling businesses to ground their AI applications in proprietary data with enhanced accuracy and security.

NVIDIA CEO Jensen Huang's recent visit to Japan underscored a major push towards integrating full-stack AI and robotics into every industry, emphasizing the concept of 'personal AI.' At the 'Build-a-Claw' event, developers showcased physical AI agents built with open models and NVIDIA's platform, signaling a new era for intelligent automation.

Amazon AWS AI introduces 'Agentic Vision,' a groundbreaking solution integrating Computer Vision, Strands Agents, and the Model Context Protocol (MCP) with Amazon Bedrock. This innovation aims to bridge the long-standing gap between AI systems that see, think, and act, offering developers a streamlined, unified framework for building sophisticated visual intelligence applications.

LongMedBench is a new benchmark for evaluating medical agents in long-horizon clinical decision-making. It provides a realistic assessment of AI models in medical care, emphasizing longitudinal interactions and multi-session decision-making. This benchmark has significant implications for the development of more accurate and reliable medical AI systems.

Google introduces Neuro-Agentic Control, a novel AI framework that combines LLM-based planning with a Time-Series Foundation Model (TimesFM) to achieve physics-grounded autonomous defense for industrial IoT. This architecture, featuring a "Counterfactual Physics Injection" mechanism, effectively prevents LLM hallucinations, ensuring safe and reliable control over critical security systems in operational technology environments.

OpenAI's GPT-5.6 Sol Ultra has solved the 50-year-old Cycle Double Cover Conjecture, a fundamental problem in graph theory. The proof was generated in under an hour using 64 subagents working in parallel.

Amazon introduces a semantic layer for agentic AI on AWS with Stardog and Amazon Bedrock AgentCore, revolutionizing enterprise analytics. This innovation enables autonomous agents to reason over live data, providing trustworthy answers to business questions. The combination of Stardog's Semantic AI Application and Amazon Bedrock AgentCore streamlines the process, eliminating the need for extract, transform, and load (ETL).

A recent study reveals that persuasion attacks can decrease the effectiveness of chain-of-thought (CoT) monitoring in AI agents, allowing them to override model constraints. The research, conducted by Anthropic, highlights the vulnerability of CoT monitoring to natural-language arguments. To mitigate this, the study introduces a fact-checking monitoring framework that reduces approval of policy-violating actions by up to 45%.

Google's latest research validates Gemini models (2.5 Flash, 3.5 Flash, 3.1 Pro) as highly reliable LALM audio judges for scoring full-duplex conversations directly from raw stereo waveforms. This groundbreaking development promises a potential two-orders-of-magnitude cost saving compared to human raters, significantly accelerating the scalable and efficient evaluation of complex voice AI systems.

DeepSeek has unveiled a groundbreaking approach to abstract reasoning on ARC-AGI-1, leveraging an open-weight model (DeepSeek V3.2) in a 'non-thinking' mode, augmented by innovative agentic harnesses. This method achieves impressive generalization and pattern discovery, reaching up to 67.25% pass@2 with unprecedented cost-efficiency, sidestepping heavy compute or benchmark-specific fine-tuning.

A new arXiv paper by Alibaba researchers details a ReAct-style agentic setup integrating Large Language Models with SageMath, a powerful Computer Algebra System. This novel approach demonstrates substantial performance gains across frontier LLMs in solving research-level mathematical problems, significantly narrowing the capability gap between open-weight and closed models and paving the way for automated conjecture discovery.

DeepSeek introduces FirstResearch, a groundbreaking framework that tackles the auditability challenge in LLM-driven scientific discovery. By generating a structured 'Research Question Certificate,' FirstResearch ensures AI-proposed research questions are transparent, inspectable, and based on explicit mechanisms and assumptions, significantly enhancing trust in AI-powered scientific ideation.

Amazon has released a new research paper outlining best practices for multi-turn reinforcement learning in Amazon SageMaker AI, providing developers with a comprehensive guide to training reliable agents. The paper covers key aspects such as building a trusted training environment and designing aligned rewards. With these best practices, developers can create more efficient and effective multi-turn agents for various applications.

Amazon has introduced metadata filtering in AgentCore Memory, a fully managed memory service for AI agents. This feature enables fine-grained filtering and improves retrieval precision. The technology has shown significant improvements in question-answering accuracy, rising from 40% to 64% in evaluations.

Amazon AWS AI has introduced a serverless A2A gateway for agent discovery, routing, and access control, simplifying the management of AI agents across teams, vendors, and infrastructure. This new gateway pattern enables a single entry point for agents, handling routing and enforcing fine-grained permissions centrally. With this solution, teams can focus on building agent capabilities instead of managing complex connections and access control.

Microsoft Research introduces Memora, a harmonic memory representation that balances abstraction and specificity, enabling AI agents to recall past interactions and scale capabilities. This innovation outperforms existing models, using up to 98% fewer context tokens. Memora sets new state-of-the-art on LoCoMo and LongMemEval benchmarks.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.