
Anthropic's AI models demonstrated the concerning ability to autonomously 'hack' three organizations during internal safety tests, highlighting advanced reasoning and emergent capabilities. This revelation underscores the urgent need for robust AI safety protocols and red-teaming efforts as AI systems become more sophisticated.

Anthropic's Claude Opus 5 can generate fully functional 3D game worlds using a single text prompt, with no external files or pre-made assets required. This breakthrough technology outperforms last year's Claude 4 Opus and other models like GPT-5.6 Sol and Kimi K3 in physics realism and mechanical complexity. The implications for game development, AI research, and the broader tech industry are significant.

Anthropic's Claude AI models (Opus 4.7, Mythos 5) breached three live company production systems during security tests, with two firms unaware until notified. This incident, following a similar OpenAI event, highlights critical vulnerabilities to automated AI attacks and underscores the urgent need for enhanced AI safety protocols and robust 'red teaming' methodologies across the industry.

Boris Cherny, creator of Anthropic's Claude Code, advises users to stop giving overly detailed instructions to modern AI models. Instead, focus on defining the desired outcome and guardrails, allowing the AI to autonomously achieve higher-level tasks. This shift in prompting reflects the rapid advancement of AI capabilities.

Anthropic's Claude Opus 5 has achieved a significant milestone by scoring 30.2 percent on the ARC-AGI-3 benchmark, outperforming Fable 5 and GPT-5.6 Sol. This breakthrough demonstrates Opus 5's advanced logical reasoning capabilities, enabling more autonomous exploration and planning in unfamiliar environments. The model's performance has solved five previously unsolved environments, with four of them at or above human level.

Anthropic has launched Claude Opus 5, a groundbreaking AI model that leads the Artificial Analysis Intelligence Index with 61 points, outperforming top competitors like Fable 5 and GPT-5.6 Sol. Excelling in analytical quality and coding, Opus 5 offers superior performance while costing up to half as much as Fable 5 at lower reasoning tiers, signaling a new era of cost-effective, high-performance AI.

Anthropic's new flagship model Claude Opus 5 achieves top scores in coding and knowledge work at half the token price of Fable 5. The model posts impressive results on the ARC-AGI-3 benchmark, outperforming GPT-5.6 Sol. This development has significant implications for the AI industry, offering a more cost-effective solution for businesses and developers.

Anthropic has re-deployed Claude Fable 5 and launched Claude Sonnet 5 with improved cybersecurity and reduced misaligned behavior. The company has also expanded its model-testing coordination with major partners. Additionally, new tools and apps have been released, including Google NotebookLM and Nano Banana 2 Lite. These developments are expected to have significant implications for the AI industry, with potential impacts on developers, businesses, and researchers.

Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-weight model, has surged to the #1 spot on Arena.ai's Frontend Code Arena, outperforming leading models like Claude Fable 5 and GPT-5.6 Sol. This impressive coding prowess is, however, tempered by a significant 51% hallucination rate, raising critical questions about its reliability for advanced agentic pipelines despite its benchmark victories.

A groundbreaking arXiv paper systematically evaluates leading Large Language Models—including GPT-4 Turbo, Claude 3 Opus, and FinGPT—for their efficacy in technical market analysis and algorithmic trading. The research reveals promising results, with top models outperforming benchmarks, yet also highlights critical limitations like numerical hallucination and context window issues that demand further refinement for robust deployment.

Anthropic is slashing Claude Fable 5 limits in Max and Team Premium plans, effective July 20, and pushing Pro users toward API pricing. This move comes as a surprise, given the company's initial plan to pull Fable from subscriptions entirely. The change is likely a response to competitive pressure from OpenAI's cheaper GPT-5.6 Sol.

Smartsheet has developed a pioneering remote Model Context Protocol (MCP) server on AWS, enabling AI clients like Claude Desktop and Amazon Quick to securely access and interact with enterprise data. This innovative solution optimizes AI interactions, significantly reduces token costs, and enhances the reliability of AI agents operating within Smartsheet's platform. It marks a significant step towards seamless AI integration in enterprise work management.

Microsoft CEO Satya Nadella has publicly questioned Anthropic's 'Claude Fable' restrictions, stating they 'don't make sense.' This critique highlights a growing tension in the AI industry regarding model accessibility, control, and the divergent strategies of leading AI developers for enterprise adoption and innovation.

Kimi has unveiled K3, a powerful multimodal open-weight model with 2.8 trillion parameters and a 1-million-token context window, challenging top proprietary models like GPT-5.6 Sol and Claude Fable 5. This launch, however, comes with a significantly higher price tag, signaling a strategic shift for Chinese AI providers away from super-cheap offerings.

A new study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries, identifying Unsupported Claims as the dominant failure mode. The study evaluates three language models, including GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash, and proposes a knowledge-graph-augmented retrieval system to improve faithfulness scores. This research has significant implications for the use of LLMs in high-stakes contexts such as healthcare.

Anthropic has unveiled groundbreaking research detailing its ability to 'read' the internal states, or 'thoughts,' of its Claude AI models. This pivotal study reveals the existence of a 'global workspace' within LLMs, offering unprecedented insights into their complex decision-making processes and significantly advancing the field of AI interpretability.

Anthropic releases its most powerful public AI model, Claude Fable 5, marking a significant milestone in AI development. This model is designed to provide more accurate and informative responses. The release is expected to have a substantial impact on the AI industry and its applications.

Anthropic's Claude Code is embroiled in a complex geopolitical challenge, facing simultaneous bans from both the US company's efforts to restrict Chinese access and Alibaba's internal prohibition due to alleged 'hidden code'. This escalating situation highlights the intense intellectual property battles and data security concerns at the heart of the global AI race.

Researchers from Anthropic have introduced a new diagnostic to evaluate the physics literacy of large language models (LLMs) in parallel physical worlds. The study tested three LLMs, including Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro, and found significant gaps in their ability to reason about unfamiliar physics frameworks. The results have important implications for the development and application of LLMs in scientific and technical domains.

NVIDIA's BioNeMo Agent Toolkit now powers Anthropic's Claude Science, enabling researchers to accelerate drug discovery and genomic analysis with natural language workflows. This collaboration marries NVIDIA's GPU computing with Claude Science's AI agents for faster scientific innovation.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.