
Researchers introduce PlanFlip, a framework to attack multi-agent LLM systems via planning-phase prompt injection, revealing vulnerabilities in popular models like GPT-5 and Llama-3.3-70B. The study highlights the importance of heterogeneous model diversity for security. PlanFlip's four attacks can corrupt downstream sub-tasks, evading keyword filters and compromising system integrity.

Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-weight model, has surged to the #1 spot on Arena.ai's Frontend Code Arena, outperforming leading models like Claude Fable 5 and GPT-5.6 Sol. This impressive coding prowess is, however, tempered by a significant 51% hallucination rate, raising critical questions about its reliability for advanced agentic pipelines despite its benchmark victories.

MarkTechPost compared Qwen, Gemma, Mistral, and DeepSeek, the best local LLMs that can run on a single 24GB GPU in 2026. This comparison highlights the performance and capabilities of each model, providing insights for developers and businesses. The article discusses the key details, technical analysis, and industry impact of these LLMs.

Discover the top local LLMs that can run on a single 24GB GPU in 2026, including Qwen, Gemma, Mistral, and DeepSeek. Learn how to choose the right model for your needs and optimize performance. Get the latest insights on AI model development and deployment.

A groundbreaking arXiv paper systematically evaluates leading Large Language Models—including GPT-4 Turbo, Claude 3 Opus, and FinGPT—for their efficacy in technical market analysis and algorithmic trading. The research reveals promising results, with top models outperforming benchmarks, yet also highlights critical limitations like numerical hallucination and context window issues that demand further refinement for robust deployment.

Smartsheet has developed a pioneering remote Model Context Protocol (MCP) server on AWS, enabling AI clients like Claude Desktop and Amazon Quick to securely access and interact with enterprise data. This innovative solution optimizes AI interactions, significantly reduces token costs, and enhances the reliability of AI agents operating within Smartsheet's platform. It marks a significant step towards seamless AI integration in enterprise work management.

Microsoft CEO Satya Nadella has publicly questioned Anthropic's 'Claude Fable' restrictions, stating they 'don't make sense.' This critique highlights a growing tension in the AI industry regarding model accessibility, control, and the divergent strategies of leading AI developers for enterprise adoption and innovation.

Meta researchers have introduced a novel neuro-symbolic agentic framework to significantly enhance the reasoning capabilities of Small Language Models (SLMs) like Gemma and Llama 3.2. This approach leverages knowledge graph grounding to overcome SLMs' historical struggles with complex, multi-hop logical tasks, offering a sustainable alternative to costly LLMs.

Meta has released a new AI model that combines reinforcement learning with large language models to create a more transparent and reliable insulin pump controller for Type 1 Diabetes patients. The model, called LLM-T1D, has shown promising results in blood sugar control and safety verification. This breakthrough has the potential to revolutionize the treatment of Type 1 Diabetes and improve the lives of millions of people worldwide.

OpenAI has introduced GPT-Red, an advanced LLM designed as a 'super-hacker' to rigorously test and enhance the safety of its other AI models. This innovative system automates critical red-teaming evaluations, enabling OpenAI to proactively identify vulnerabilities and strengthen defenses against sophisticated cyberattacks. The move signifies a major leap in AI safety protocols, aiming to keep pace with evolving threats.

A new study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries, identifying Unsupported Claims as the dominant failure mode. The study evaluates three language models, including GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash, and proposes a knowledge-graph-augmented retrieval system to improve faithfulness scores. This research has significant implications for the use of LLMs in high-stakes contexts such as healthcare.

OpenAI introduces a new method for detecting model distillation in large language models, raising questions about fairness and policy violations. The approach uses reference-based membership inference to identify teacher models. This breakthrough has significant implications for the AI industry, developers, and businesses.

Google introduces Neuro-Agentic Control, a novel AI framework that combines LLM-based planning with a Time-Series Foundation Model (TimesFM) to achieve physics-grounded autonomous defense for industrial IoT. This architecture, featuring a "Counterfactual Physics Injection" mechanism, effectively prevents LLM hallucinations, ensuring safe and reliable control over critical security systems in operational technology environments.

Groundbreaking Arxiv research reveals how large language model safety mechanisms are encoded and can be bypassed, introducing novel 'Activation-Guided' adversarial attacks. The study finds safety representations are distributed across model layers, not localized, and proposes a 33x faster attack method, Soft-GCG, offering critical insights for designing more robust AI alignment strategies.

Anthropic has unveiled groundbreaking research detailing its ability to 'read' the internal states, or 'thoughts,' of its Claude AI models. This pivotal study reveals the existence of a 'global workspace' within LLMs, offering unprecedented insights into their complex decision-making processes and significantly advancing the field of AI interpretability.

Google researchers introduce HCC-STAR, a clinical-reasoning LLM for risk stratification and treatment guidance in hepatocellular carcinoma. This model achieves state-of-the-art performance in treatment recommendation and risk stratification. The study demonstrates the potential of AI in precision therapy for HCC patients.

A new arXiv paper by Alibaba researchers details a ReAct-style agentic setup integrating Large Language Models with SageMath, a powerful Computer Algebra System. This novel approach demonstrates substantial performance gains across frontier LLMs in solving research-level mathematical problems, significantly narrowing the capability gap between open-weight and closed models and paving the way for automated conjecture discovery.

DeepSeek introduces FirstResearch, a groundbreaking framework that tackles the auditability challenge in LLM-driven scientific discovery. By generating a structured 'Research Question Certificate,' FirstResearch ensures AI-proposed research questions are transparent, inspectable, and based on explicit mechanisms and assumptions, significantly enhancing trust in AI-powered scientific ideation.

New research reveals a critical vulnerability in advanced reasoning AI models, where logically inconsistent prompts can force them into 'overthinking,' leading to denial-of-service attacks. This 'Evolutionary Prompt Attack' significantly increases resource consumption and poses a serious threat to commercial LLM providers like OpenAI, Google, and DeepSeek.

A groundbreaking Arxiv paper introduces LLMForge, a multi-model text-to-CAD framework that enables automatic generation of parametric 3D mechanical designs from natural language. This framework, featuring innovative iterative refinement and VLM-based critique, demonstrates remarkable success, with top models like DeepSeek-V3.2 achieving near-perfect mesh generation and showing compact models can rival larger systems.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.