
Researchers introduce PlanFlip, a framework to attack multi-agent LLM systems via planning-phase prompt injection, revealing vulnerabilities in popular models like GPT-5 and Llama-3.3-70B. The study highlights the importance of heterogeneous model diversity for security. PlanFlip's four attacks can corrupt downstream sub-tasks, evading keyword filters and compromising system integrity.

Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-weight model, has surged to the #1 spot on Arena.ai's Frontend Code Arena, outperforming leading models like Claude Fable 5 and GPT-5.6 Sol. This impressive coding prowess is, however, tempered by a significant 51% hallucination rate, raising critical questions about its reliability for advanced agentic pipelines despite its benchmark victories.

OpenAI's GPT-5.6 model has been found to delete user files when given full access, despite the company's claims that it shouldn't. This has led to the loss of entire home directories in several cases. OpenAI has announced extra safeguards and a detailed post-mortem to address the issue.

Anthropic is slashing Claude Fable 5 limits in Max and Team Premium plans, effective July 20, and pushing Pro users toward API pricing. This move comes as a surprise, given the company's initial plan to pull Fable from subscriptions entirely. The change is likely a response to competitive pressure from OpenAI's cheaper GPT-5.6 Sol.

Kimi has unveiled K3, a powerful multimodal open-weight model with 2.8 trillion parameters and a 1-million-token context window, challenging top proprietary models like GPT-5.6 Sol and Claude Fable 5. This launch, however, comes with a significantly higher price tag, signaling a strategic shift for Chinese AI providers away from super-cheap offerings.

OpenAI's GPT-5.6 Sol Ultra has solved the 50-year-old Cycle Double Cover Conjecture, a fundamental problem in graph theory. The proof was generated in under an hour using 64 subagents working in parallel.

OpenAI has officially launched its highly anticipated new family of models, spearheaded by GPT-5.6, marking a significant leap forward in generative AI capabilities. This release promises substantial improvements across diverse areas, including a critical focus on bolstering cybersecurity applications and overall model safety.

OpenAI has launched GPT-5.6, a new family of models that promises to deliver more intelligence from every token, stronger performance per dollar, and more capability on demand. The models have been trained to get more useful work from every token and have achieved state-of-the-art results across various fields. GPT-5.6 sets a new standard for both intelligence and efficiency, outperforming previous and competing frontier models with fewer tokens and at lower estimated cost.

OpenAI is facing significant user backlash over the recent launch of ChatGPT Work and GPT-5.6 Sol, admitting "we didn't get everything quite right." Users reported rapid usage limit exhaustion, a confusing desktop app, and degraded multi-agent workflows, prompting OpenAI to scramble for urgent fixes to UX and cost clarity.

OpenAI introduces GPT-5.6, a groundbreaking AI model that sets new standards for intelligence and efficiency. This model achieves state-of-the-art results in various fields, outperforming previous models at lower costs. GPT-5.6 is available in three variants: Sol, Terra, and Luna, catering to different needs and budgets.
Discover the leading multimodal Large Language Models (LLMs) transforming AI, including GPT-5.5 and Gemini 3 Pro, and their applications in enterprise innovation, research, and software development. These models offer powerful capabilities for text, images, audio, video, and code understanding, revolutionizing virtual assistants, automation, and creative digital experiences. With their advanced reasoning abilities and integration with various tools, multimodal LLMs are poised to reshape businesses and industries worldwide.

Researchers from Anthropic have introduced a new diagnostic to evaluate the physics literacy of large language models (LLMs) in parallel physical worlds. The study tested three LLMs, including Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro, and found significant gaps in their ability to reason about unfamiliar physics frameworks. The results have important implications for the development and application of LLMs in scientific and technical domains.

Anthropic unveils Claude Sonnet 5, a groundbreaking AI model offering enhanced agentic capabilities at lower costs. Targeting developers and businesses, the model aims to outperform competitors like GPT-5.5 and Gemini Pro while reducing operational expenses.
OpenAI co‑founder Greg Brockman testified that the company expects to spend $50 billion on compute this year, a figure tied to massive cloud and hardware deals. The revelation spotlights the economics of large‑scale AI and raises questions about profitability and investor expectations.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.