
Researchers introduce PlanFlip, a framework to attack multi-agent LLM systems via planning-phase prompt injection, revealing vulnerabilities in popular models like GPT-5 and Llama-3.3-70B. The study highlights the importance of heterogeneous model diversity for security. PlanFlip's four attacks can corrupt downstream sub-tasks, evading keyword filters and compromising system integrity.

DeepSeek has unveiled a groundbreaking approach to abstract reasoning on ARC-AGI-1, leveraging an open-weight model (DeepSeek V3.2) in a 'non-thinking' mode, augmented by innovative agentic harnesses. This method achieves impressive generalization and pattern discovery, reaching up to 67.25% pass@2 with unprecedented cost-efficiency, sidestepping heavy compute or benchmark-specific fine-tuning.

DeepSeek introduces FirstResearch, a groundbreaking framework that tackles the auditability challenge in LLM-driven scientific discovery. By generating a structured 'Research Question Certificate,' FirstResearch ensures AI-proposed research questions are transparent, inspectable, and based on explicit mechanisms and assumptions, significantly enhancing trust in AI-powered scientific ideation.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.