
A groundbreaking arXiv paper systematically evaluates leading Large Language Models—including GPT-4 Turbo, Claude 3 Opus, and FinGPT—for their efficacy in technical market analysis and algorithmic trading. The research reveals promising results, with top models outperforming benchmarks, yet also highlights critical limitations like numerical hallucination and context window issues that demand further refinement for robust deployment.

The AI landscape is rapidly evolving as three Chinese labs—Moonshot AI, DeepSeek, and Zhipu AI—release powerful, open-weight Mixture-of-Experts (MoE) models. Kimi K3, DeepSeek V4 Pro, and GLM-5.2 are pushing the boundaries of scale and capability, offering trillion-scale parameters and million-token context windows for complex coding and agent workloads, fundamentally shifting the open-source AI leaderboard.

Meta has released a new AI model that combines reinforcement learning with large language models to create a more transparent and reliable insulin pump controller for Type 1 Diabetes patients. The model, called LLM-T1D, has shown promising results in blood sugar control and safety verification. This breakthrough has the potential to revolutionize the treatment of Type 1 Diabetes and improve the lives of millions of people worldwide.

Kimi has unveiled K3, a powerful multimodal open-weight model with 2.8 trillion parameters and a 1-million-token context window, challenging top proprietary models like GPT-5.6 Sol and Claude Fable 5. This launch, however, comes with a significantly higher price tag, signaling a strategic shift for Chinese AI providers away from super-cheap offerings.

OpenAI introduces a new method for detecting model distillation in large language models, raising questions about fairness and policy violations. The approach uses reference-based membership inference to identify teacher models. This breakthrough has significant implications for the AI industry, developers, and businesses.

OpenAI has officially launched its highly anticipated new family of models, spearheaded by GPT-5.6, marking a significant leap forward in generative AI capabilities. This release promises substantial improvements across diverse areas, including a critical focus on bolstering cybersecurity applications and overall model safety.

A new arXiv paper by Alibaba researchers details a ReAct-style agentic setup integrating Large Language Models with SageMath, a powerful Computer Algebra System. This novel approach demonstrates substantial performance gains across frontier LLMs in solving research-level mathematical problems, significantly narrowing the capability gap between open-weight and closed models and paving the way for automated conjecture discovery.
Discover the leading multimodal Large Language Models (LLMs) transforming AI, including GPT-5.5 and Gemini 3 Pro, and their applications in enterprise innovation, research, and software development. These models offer powerful capabilities for text, images, audio, video, and code understanding, revolutionizing virtual assistants, automation, and creative digital experiences. With their advanced reasoning abilities and integration with various tools, multimodal LLMs are poised to reshape businesses and industries worldwide.

Researchers from Anthropic have introduced a new diagnostic to evaluate the physics literacy of large language models (LLMs) in parallel physical worlds. The study tested three LLMs, including Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro, and found significant gaps in their ability to reason about unfamiliar physics frameworks. The results have important implications for the development and application of LLMs in scientific and technical domains.

A revolutionary method called Poller leverages large language models to evaluate poetry understanding with near-human accuracy, reducing errors by up to 94.55% in specific dimensions. This AI advancement bridges automation and human expertise in literary analysis.

OpenAI's latest research demonstrates how large language models can automate training data labeling for entity matching, reducing manual effort by 99% and slashing costs. This breakthrough enables faster, cheaper AI deployment for businesses.

Cohere's study reveals how transformer models develop situation modeling and mentalizing capabilities through training stages. Key findings show FBT performance depends on model size, training volume, and post-training methods, but remains fragile in complex scenarios.

Emily Bender addresses misconceptions about her 2021 paper on 'stochastic parrots' and large language models, explaining their limitations and impact on AI discourse.

A new arXiv study shows OpenEvidence's specialized clinical tool beats top general‑purpose models (Claude Opus 4.8, Gemini 3.1 Pro, GPT‑5.5) on 620 real‑world point‑of‑care questions. Physicians across 30 specialties rated the specialized tool higher on accuracy, utility, source quality, verifiability and completeness.
OpenAI introduces IMCBench, a benchmark for multimodal large language models in image-grounded medical conversations, evaluating safety, accuracy, and uncertainty in diagnosis. The benchmark tests eight models, with Claude Opus 4.6 achieving the highest overall score. The results highlight the need for multi-dimensional evaluation frameworks in medical AI.
Meta has introduced AnTenA, a novel AI system that leverages large language models to explain hidden patterns in human narratives. This system uses task-agnostic and task-specific prompts to analyze co-clustered latent patterns from tensor decomposition. AnTenA has the potential to revolutionize the field of explainable AI.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.