
Anthropic has unveiled Claude Science, a flagship AI product engineered to transform scientific research, particularly in computational biology and drug development. Positioned alongside Claude Code, this standalone offering empowers researchers to autonomously carry out complex tasks, marking a significant strategic leap for Anthropic into the life sciences domain and intensifying competition in AI for scientific discovery.

Microsoft Research introduces Memora, a harmonic memory representation that balances abstraction and specificity, enabling AI agents to recall past interactions and scale capabilities. This innovation outperforms existing models, using up to 98% fewer context tokens. Memora sets new state-of-the-art on LoCoMo and LongMemEval benchmarks.

Grok 4.5, a base model with 1.5 trillion parameters, has been further trained on Cursor data and is currently in beta testing at SpaceX and Tesla. This development marks a significant milestone in AI research and its applications in the tech industry. The model's capabilities and potential uses are being explored by these industry leaders.

OpenAI's outage led to account deactivations, causing users to lose their work and face delays in their projects.

Amazon has launched a new $1 billion Frontier Deployment Engineering (FDE) organization, mirroring strategic moves by OpenAI and Anthropic. This initiative aims to embed expert engineers within client companies to accelerate the deployment of purpose-built AI agents, emphasizing rapid integration and fostering customer self-sufficiency in cutting-edge AI adoption.

NVIDIA's Jaiveer Singh leads the charge in accelerating the future of robotics with Isaac ROS, a CUDA-accelerated software platform built on ROS 2. This initiative empowers developers to build and deploy autonomous robots faster, bridging the gap between imaginative concepts and real-world utility. Isaac ROS is poised to become the foundational 'connective tissue' for the physical AI era.

Amazon AWS AI introduces managed entitlements for Amazon Bedrock models, simplifying access across multiple accounts. This feature removes the need for AWS Marketplace permissions in workload accounts, streamlining AI adoption. Organizations can now subscribe once from a central account and distribute model access across their organization.

Anthropic has released version 0.115.0 of its SDK for Python, introducing new features such as support for Managed Agents event delta streaming and agent overrides. This update aims to enhance the functionality and usability of the Anthropics SDK, providing developers with more tools to work with AI models. The release is part of Anthropic's ongoing efforts to improve its offerings and stay competitive in the AI market.

DeepSeek's new study explores the effectiveness of learned stopping in reasoning models, finding that it can improve performance in certain tasks. The study introduces LearnStop, a hidden-state-free checkpoint stopper, and evaluates its performance across 18 task-model settings. The results show that learned stopping can be useful in tasks where many questions become correct before full budget but do not exhibit a single reliable scalar stopping signal.

New research from arXiv challenges conventional wisdom on AI improvement from feedback, revealing that multi-turn gains often mask true learning. The study highlights that an AI model's ability to effectively *utilize* feedback, rather than merely receiving it, is the critical bottleneck for interactive improvement, especially when compared to unguided self-refinement or simple retries.

NVIDIA's full-stack inference software, optimized for the Blackwell platform, slashes token costs by 5x on DeepSeek V4. Companies like Baseten and Cognition leverage NVIDIA's tools to scale AI workloads efficiently, marking a shift from hardware specs to cost-per-token economics.

NVIDIA's BioNeMo Agent Toolkit now powers Anthropic's Claude Science, enabling researchers to accelerate drug discovery and genomic analysis with natural language workflows. This collaboration marries NVIDIA's GPU computing with Claude Science's AI agents for faster scientific innovation.

A revolutionary method called Poller leverages large language models to evaluate poetry understanding with near-human accuracy, reducing errors by up to 94.55% in specific dimensions. This AI advancement bridges automation and human expertise in literary analysis.

OpenAI's latest research demonstrates how large language models can automate training data labeling for entity matching, reducing manual effort by 99% and slashing costs. This breakthrough enables faster, cheaper AI deployment for businesses.

Cohere's study reveals how transformer models develop situation modeling and mentalizing capabilities through training stages. Key findings show FBT performance depends on model size, training volume, and post-training methods, but remains fragile in complex scenarios.

A groundbreaking AI system combines time-series forecasting, anomaly detection, and LLM-driven analysis to deliver actionable energy insights. This end-to-end solution reduces alert noise for facility managers while maintaining high accuracy across 16 real-world scenarios.

MedEvoEval introduces a groundbreaking framework for evaluating AI doctor agents in simulated clinical settings. By tracking cross-episode learning and decision-making, it addresses critical gaps in medical AI evaluation. This tool enables developers to measure knowledge retention, resource allocation, and behavioral adaptation over time.

Meta researchers achieved 87.69% accuracy in predicting primary ICD-10 diagnosis categories by combining frozen medical LLM representations with multimodal EHR data. Their approach outperformed existing models and demonstrated strong cross-dataset adaptability.

A groundbreaking study reveals that traditional safety methods for AI agents are fundamentally flawed. Instead of relying on refusal-based content safety, the paper advocates for action alignment and least privilege enforcement to ensure secure, user-intent-driven AI systems.
Microsoft secretly embedded 'AARD code' in Windows 3.1 betas to sabotage DR DOS, triggering fake errors. This led to a $280M settlement with Caldera, Inc., revealing anti-competitive tactics in tech's past.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.