
OpenAI introduces RobustMAD, a benchmark for evaluating multimodal small language models' real-world robustness in anomaly detection. The study reveals promising capabilities of compact models but also critical robustness gaps. RobustMAD provides actionable guidance for designing next-generation industrial inspection assistants.

Research reveals AI models, including ChatGPT and DeepSeek's R1, exhibit stronger biases than humans when hiring, segregating candidates into jobs based on early observations. This discovery has significant implications for the use of AI in recruitment processes. The study suggests that newer models with higher reasoning capabilities show even more pronounced biases.

NVIDIA made significant strides at SIGGRAPH 2026, showcasing how agentic and physical AI are set to revolutionize graphics, simulation, and digital world creation. Key announcements included new tools for AI-driven content creation, advanced neural rendering techniques, and an open world model for local physical AI, redefining realism and automation across industries.

Researchers introduce PlanFlip, a framework to attack multi-agent LLM systems via planning-phase prompt injection, revealing vulnerabilities in popular models like GPT-5 and Llama-3.3-70B. The study highlights the importance of heterogeneous model diversity for security. PlanFlip's four attacks can corrupt downstream sub-tasks, evading keyword filters and compromising system integrity.

Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-weight model, has surged to the #1 spot on Arena.ai's Frontend Code Arena, outperforming leading models like Claude Fable 5 and GPT-5.6 Sol. This impressive coding prowess is, however, tempered by a significant 51% hallucination rate, raising critical questions about its reliability for advanced agentic pipelines despite its benchmark victories.

More than 60 tech funding deals worth over €2.7 billion were tracked in Europe last week, with artificial intelligence, healthtech, and software being the top industries. Germany, Sweden, and the UK led the countries with the most funding. The Tech.eu Funding Explorer provides deeper insights into funding data, investor activity, and market trends.

MarkTechPost compared Qwen, Gemma, Mistral, and DeepSeek, the best local LLMs that can run on a single 24GB GPU in 2026. This comparison highlights the performance and capabilities of each model, providing insights for developers and businesses. The article discusses the key details, technical analysis, and industry impact of these LLMs.

Discover the top local LLMs that can run on a single 24GB GPU in 2026, including Qwen, Gemma, Mistral, and DeepSeek. Learn how to choose the right model for your needs and optimize performance. Get the latest insights on AI model development and deployment.

DeepSeek's new AI model uses a unified multimodal learner for clinical prediction, simplifying the process and achieving state-of-the-art results. This approach converts all patient data into a single natural language sequence and fine-tunes a pretrained language model. The model outperforms task-specific multimodal baselines and a clinically deployed gradient boosting system.

A groundbreaking arXiv paper systematically evaluates leading Large Language Models—including GPT-4 Turbo, Claude 3 Opus, and FinGPT—for their efficacy in technical market analysis and algorithmic trading. The research reveals promising results, with top models outperforming benchmarks, yet also highlights critical limitations like numerical hallucination and context window issues that demand further refinement for robust deployment.

OpenAI's latest study explores the necessity of executable world models, simplification, and verification in coding agents, revealing surprising results. The research evaluates four nested Codex-based agents, finding that every agent variant improves with stronger models and greater reasoning effort. The study's findings have significant implications for the development of Artificial General Intelligence (AGI).

Renowned author Dave Eggers confronted OpenAI staff, stating ChatGPT was 'silencing an entire generation,' raising critical questions about AI's impact on human creativity and the future of artistic expression. This direct critique from a prominent cultural figure highlights growing tensions between generative AI developers and the creative community. The event underscores the urgent need for ethical considerations and sustainable models for creators in an AI-driven world.

Alibaba's Qwen team has released Qwen 3.8, a multimodal AI model with 2.4 trillion parameters, rivaling leading models and trailing only Fable 5. The model is available for preview now. This development is set to significantly impact the AI landscape, offering enhanced capabilities and potential applications across various industries.

HuggingFace has introduced a new AI model, DanielTobi0/afrique-qwen-8b-health-finetuned-2, with 259 downloads and capabilities in text generation and health-focused applications. This model utilizes transformers, safetensors, and qwen3, showcasing advancements in AI technology. The model's performance and potential applications are of significant interest to the AI community.

Moonshot's Kimi K3 has surpassed Fable 5 in frontend code, becoming the first Chinese model to top the Code Arena: Frontend rankings. However, it lags behind in complex math, scoring only 39% on FrontierMath Tier 4. This development has significant implications for the AI industry, with potential opportunities and risks for developers, businesses, and investors.

The AI landscape is rapidly evolving as three Chinese labs—Moonshot AI, DeepSeek, and Zhipu AI—release powerful, open-weight Mixture-of-Experts (MoE) models. Kimi K3, DeepSeek V4 Pro, and GLM-5.2 are pushing the boundaries of scale and capability, offering trillion-scale parameters and million-token context windows for complex coding and agent workloads, fundamentally shifting the open-source AI leaderboard.

HuggingFace has released a new AI model, EnzGamers/ABCDAI-R1-1.5b-SFT, with 4,126 downloads and 2 likes. This model is part of the transformers family and utilizes safetensors and qwen2 for text generation. The model has been generated from a trainer and features SFT technology.

OpenAI has announced a significant policy shift, committing to notify parents when their teenage children are removed from ChatGPT for violating its terms of service. This move underscores a growing commitment to user safety and responsible AI use, particularly among underage users, setting a new precedent for AI platform moderation.

China's Moonshot AI has released Kimi K3, a model that matches Anthropic's Opus 4.8, raising questions about the importance of computing power in AI development. This release is reigniting the debate over US export controls and the future of AI. The implications of Kimi K3's release are far-reaching, with potential consequences for the AI industry and global technological landscape.

OpenAI's GPT-5.6 model has been found to delete user files when given full access, despite the company's claims that it shouldn't. This has led to the loss of entire home directories in several cases. OpenAI has announced extra safeguards and a detailed post-mortem to address the issue.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.