
Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-weight model, has surged to the #1 spot on Arena.ai's Frontend Code Arena, outperforming leading models like Claude Fable 5 and GPT-5.6 Sol. This impressive coding prowess is, however, tempered by a significant 51% hallucination rate, raising critical questions about its reliability for advanced agentic pipelines despite its benchmark victories.

A groundbreaking arXiv paper systematically evaluates leading Large Language Models—including GPT-4 Turbo, Claude 3 Opus, and FinGPT—for their efficacy in technical market analysis and algorithmic trading. The research reveals promising results, with top models outperforming benchmarks, yet also highlights critical limitations like numerical hallucination and context window issues that demand further refinement for robust deployment.

A new study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries, identifying Unsupported Claims as the dominant failure mode. The study evaluates three language models, including GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash, and proposes a knowledge-graph-augmented retrieval system to improve faithfulness scores. This research has significant implications for the use of LLMs in high-stakes contexts such as healthcare.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.