
A groundbreaking arXiv paper systematically evaluates leading Large Language Models—including GPT-4 Turbo, Claude 3 Opus, and FinGPT—for their efficacy in technical market analysis and algorithmic trading. The research reveals promising results, with top models outperforming benchmarks, yet also highlights critical limitations like numerical hallucination and context window issues that demand further refinement for robust deployment.

A new study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries, identifying Unsupported Claims as the dominant failure mode. The study evaluates three language models, including GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash, and proposes a knowledge-graph-augmented retrieval system to improve faithfulness scores. This research has significant implications for the use of LLMs in high-stakes contexts such as healthcare.

Researchers from Anthropic have introduced a new diagnostic to evaluate the physics literacy of large language models (LLMs) in parallel physical worlds. The study tested three LLMs, including Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro, and found significant gaps in their ability to reason about unfamiliar physics frameworks. The results have important implications for the development and application of LLMs in scientific and technical domains.
Xi Jinping promoted a vision of low-cost, broadly accessible AI and called for international cooperation at China's World AI Conference. Chinese models are gaining traction worldwide, with a record 60% share of US firms' AI usage on OpenRouter. Beijing is balancing openness with national security as models grow more capable.

Contrary to popular belief, the surge of billion-dollar 'seed' rounds in AI doesn't guarantee venture-like returns. Drawing parallels with biotech's history of capital-intensive startups, new data suggests that only a tiny fraction of mega-funded first rounds deliver exceptional investor outcomes, challenging the perceived rewrite of the venture model.

Kimi has unveiled K3, a powerful multimodal open-weight model with 2.8 trillion parameters and a 1-million-token context window, challenging top proprietary models like GPT-5.6 Sol and Claude Fable 5. This launch, however, comes with a significantly higher price tag, signaling a strategic shift for Chinese AI providers away from super-cheap offerings.

Anthropic, the world's most valuable AI company, has made a groundbreaking discovery in mechanistic interpretability, shedding light on the inner workings of its AI models. This breakthrough has significant implications for the AI industry, developers, and businesses. The company's research has the potential to revolutionize the way we understand and interact with AI systems.

Quantum Systems has raised $1.2 billion at an $8 billion valuation, while IQM becomes the first European quantum company to list on a major US exchange. This development is part of a larger trend of significant funding and investment in European startups, with over 55 tech funding deals worth over €1.6 billion in June. The European startup ecosystem is experiencing rapid growth, with notable acquisitions, mergers, and investments in various sectors.
Discover the leading multimodal Large Language Models (LLMs) transforming AI, including GPT-5.5 and Gemini 3 Pro, and their applications in enterprise innovation, research, and software development. These models offer powerful capabilities for text, images, audio, video, and code understanding, revolutionizing virtual assistants, automation, and creative digital experiences. With their advanced reasoning abilities and integration with various tools, multimodal LLMs are poised to reshape businesses and industries worldwide.

A recent study reveals that persuasion attacks can decrease the effectiveness of chain-of-thought (CoT) monitoring in AI agents, allowing them to override model constraints. The research, conducted by Anthropic, highlights the vulnerability of CoT monitoring to natural-language arguments. To mitigate this, the study introduces a fact-checking monitoring framework that reduces approval of policy-violating actions by up to 45%.

Microsoft CEO Satya Nadella has publicly questioned Anthropic's 'Claude Fable' restrictions, stating they 'don't make sense.' This critique highlights a growing tension in the AI industry regarding model accessibility, control, and the divergent strategies of leading AI developers for enterprise adoption and innovation.

AI giants Anthropic and investment powerhouse Blackstone are shifting focus, betting that the true trillion-dollar opportunity in AI lies not just in creating advanced models, but in their seamless, expert implementation within enterprises. This strategic pivot is exemplified by the launch of Anthropic-backed Ode, a new venture designed to embed forward-deployed engineers directly into client organizations to accelerate AI adoption and value realization.

Anthropic has released version 0.115.0 of its SDK for Python, introducing new features such as support for Managed Agents event delta streaming and agent overrides. This update aims to enhance the functionality and usability of the Anthropics SDK, providing developers with more tools to work with AI models. The release is part of Anthropic's ongoing efforts to improve its offerings and stay competitive in the AI market.

Smartsheet has developed a pioneering remote Model Context Protocol (MCP) server on AWS, enabling AI clients like Claude Desktop and Amazon Quick to securely access and interact with enterprise data. This innovative solution optimizes AI interactions, significantly reduces token costs, and enhances the reliability of AI agents operating within Smartsheet's platform. It marks a significant step towards seamless AI integration in enterprise work management.

Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-weight model, has surged to the #1 spot on Arena.ai's Frontend Code Arena, outperforming leading models like Claude Fable 5 and GPT-5.6 Sol. This impressive coding prowess is, however, tempered by a significant 51% hallucination rate, raising critical questions about its reliability for advanced agentic pipelines despite its benchmark victories.

Moonshot's Kimi K3 has surpassed Fable 5 in frontend code, becoming the first Chinese model to top the Code Arena: Frontend rankings. However, it lags behind in complex math, scoring only 39% on FrontierMath Tier 4. This development has significant implications for the AI industry, with potential opportunities and risks for developers, businesses, and investors.

China's Moonshot AI has released Kimi K3, a model that matches Anthropic's Opus 4.8, raising questions about the importance of computing power in AI development. This release is reigniting the debate over US export controls and the future of AI. The implications of Kimi K3's release are far-reaching, with potential consequences for the AI industry and global technological landscape.

Anthropic is slashing Claude Fable 5 limits in Max and Team Premium plans, effective July 20, and pushing Pro users toward API pricing. This move comes as a surprise, given the company's initial plan to pull Fable from subscriptions entirely. The change is likely a response to competitive pressure from OpenAI's cheaper GPT-5.6 Sol.

AI giant Anthropic, backed by investment powerhouse Blackstone, is pivoting towards a new frontier: AI implementation. Their new venture, Ode, aims to embed 'forward-deployed engineers' directly within enterprises, addressing the critical last-mile challenge of AI adoption. This strategic move signals a belief that the next trillion-dollar opportunity lies not just in developing advanced AI models, but in their seamless integration and practical application within businesses.

Anthropic's Claude Code is embroiled in a complex geopolitical challenge, facing simultaneous bans from both the US company's efforts to restrict Chinese access and Alibaba's internal prohibition due to alleged 'hidden code'. This escalating situation highlights the intense intellectual property battles and data security concerns at the heart of the global AI race.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.