
A new study reveals that cost-effective automated judging of natural-language mathematical proofs is possible using cheap open-weight models, achieving similar accuracy to frontier models at a significantly lower cost. This breakthrough has significant implications for the development and evaluation of math-reasoning systems. The study's findings could lead to more efficient and affordable methods for grading mathematical proofs.

OpenAI's unreleased Astra model has achieved a monumental breakthrough, solving ten previously intractable open problems in mathematics and theoretical computer science. This landmark event, detailed in a new publication, showcases AI's accelerating capability to drive fundamental scientific discovery, promising to reshape research methodologies and human-AI collaboration.
Get the top AI stories in your inbox once a day, no spam.
New stories are added every couple of hours as they break, so the feed stays current throughout the day.
We pull from 100+ sources, including company blogs, research labs, and established tech publications, then fact check and summarize each story before it goes live.
Yes. Use the sidebar filters to narrow stories down by company (OpenAI, Anthropic, Google, and more), industry, or event type like funding and research.
Yes. AI Pulse is free for anyone who wants to keep up with AI news, no sign up required. The daily newsletter is optional if you want updates in your inbox.