{"id":46867,"date":"2026-03-16T15:43:10","date_gmt":"2026-03-16T15:43:10","guid":{"rendered":"https:\/\/naijaglobalnews.org\/?p=46867"},"modified":"2026-03-16T15:43:10","modified_gmt":"2026-03-16T15:43:10","slug":"as-ai-keeps-improving-mathematicians-struggle-to-foretell-their-own-future","status":"publish","type":"post","link":"https:\/\/naijaglobalnews.org\/?p=46867","title":{"rendered":"As AI keeps improving, mathematicians struggle to foretell their own future"},"content":{"rendered":"<p>\n<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">In the ongoing campaign by artificial intelligence companies to take over pure mathematics, another round is commencing.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">The team behind First Proof, an effort to benchmark the ability of large language models (LLMs) to contribute to research-level mathematics, has announced its next exam. For this second round, which it plans to roll out over the next few months, the team is requiring transparency from any AI company that wants to participate.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">This is occurring amid a sea change in mathematics research. In just the past few months, the best publicly available models have begun generating valid proofs for minor theorems of actual use for working mathematicians. To some experts, the opening round of First Proof was a pivotal moment in this ongoing story.<\/p>\n<h2>On supporting science journalism<\/h2>\n<p>If you&#8217;re enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">\u201cWe were quite impressed with how the AI models did,\u201d says Lauren Williams, a Harvard University mathematician and First Proof team member. \u201cThe problems that we proposed really are on the forefront of what AI models\u2014perhaps together with experts\u2014can solve.\u201d<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">First Proof grew out of its 11-person team\u2019s own eye-opening\u2014if sometimes frustrating\u2014experiences with AI. No preexisting benchmarks seemed sufficient for testing LLMs as a mathematician\u2019s assistant. In principle, an LLM could save time by proving smaller \u201clemmas\u201d\u2014intermediate propositions along a mathematician\u2019s path to developing larger theorems of greater interest. In practice, however, such AI assists have tended to go awry.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">So for their initial, \u201cexperimental\u201d test, the First Proof team decided on 10 lemmas from papers that members had written but not yet released and then set a one-week deadline for AI companies (and anyone else) to try proving these propositions using their favorite models.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">Groups from both OpenAI and Google posted their LLMs\u2019 responses to all of the problems. Five of the OpenAI model\u2019s proofs appeared to be correct. And Google Deepmind\u2019s Aletheia agent seemed to get six (although experts aren\u2019t unanimous on the validity of one of these proofs). Comparing the two models\u2019 performances, Williams was surprised to find each had solved multiple problems that the other couldn\u2019t. \u201cIt\u2019s interesting to see that their capabilities are different,\u201d she says.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">\u201cThe performance was higher than I expected,\u201d says Daniel Litt, a mathematician at the University of Toronto, who isn\u2019t directly involved in the First Proof effort. All in all, as many as eight of the 10 problems appear to have been solved at least partially by AI. \u201cIt\u2019s clear that capabilities have been improving really rapidly,\u201d Litt says.<\/p>\n<h2 id=\"a-hazy-but-hopeful-future\" class=\"\" data-block=\"sciam\/heading\">A Hazy but Hopeful Future<\/h2>\n<p class=\"\" data-block=\"sciam\/paragraph\">Litt isn\u2019t afraid of AI\u2019s growing mathematical prowess. \u201cI don\u2019t expect, five years from now, to be useless,\u201d he says. \u201cI actually expect to be doing the best work I\u2019ve ever done, because I\u2019ll have these amazing tools.\u201d In fact, the First Proof results inspired him to pen an essay, which was widely circulated among mathematicians over the past few weeks. It presents a speculative, optimistic view of the field\u2019s AI-infused future.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">For the sake of argument, Litt imagines a hypothetical library generated by superintelligent AIs and containing every proof possible in the mathematical universe. A mere human mathematician wandering among its innumerable shelves could peruse all its volumes but could create no novel proof themself.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">But that doesn\u2019t mean mathematicians would be crippled with ennui, Litt says. Far from it. \u201cThey would be unbelievably excited, and immediately get to work,\u201d he wrote in the essay. The mathematical universe is so vast, he says, that the joy is in exploring it, whether by reading and digesting a proof or writing a new one. \u201cMy job wouldn\u2019t even change at all,\u201d he says. \u201cThe job now is to try to understand things.\u201d<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">Even if all mathematicians agreed with Litt\u2019s decidedly utopian take on this thought experiment, the current situation is far from that lofty ideal\u2014as evidenced by First Proof\u2019s first round. \u201cCombined, the models solved maybe eight of the problems,\u201d he says. \u201cBut they also produced thousands and thousands of pages of garbage.\u201d<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">Current AIs, it turns out, are frequently wrong but convincingly confident. They\u2019ll cite a result in the literature but pretend it\u2019s stronger than it is. Or they\u2019ll bury a crucial mistake deep inside a tedious calculation, where it\u2019s easy to miss. \u201cStudents make errors, but they\u2019re definitely not trying to make errors,\u201d Litt says. \u201cThe models are not very honest.\u201d<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">This qualitative difference in the types of quantitative errors LLMs produce can make judging their answers very challenging. \u201cOne of the things we learned from this first round is how difficult it can be to check the correctness of the results,\u201d says Mohammed Abouzaid, a First Proof team member and mathematician at Stanford University. \u201cYou would almost say, \u2018No human who would know what all these words mean would make this mistake!\u2019\u201d<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">For round two, the team plans to outsource the task of evaluating each entry to mathematicians hired as anonymous reviewers, funded with a mix of grant money and donations from AI companies. But with no sign of the en masse mathematical onslaught slowing down, a deluge of LLM-written, subtly wrong proofs may soon overwhelm human resources. \u201cPeople need to start thinking about this,\u201d Litt says. \u201cOur institutions and the profession are not adapting to what\u2019s coming down the line.\u201d<\/p>\n<h2 id=\"an-unexplained-gap\" class=\"\" data-block=\"sciam\/heading\">An Unexplained Gap<\/h2>\n<p class=\"\" data-block=\"sciam\/paragraph\">The first round apparently revealed a glaring chasm between public and proprietary efforts. This would seem to challenge the notion that AI usurping human skills will democratize them\u2014for instance, by broadening who is able to contribute meaningfully to math\u2019s advancement. <\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">In the team\u2019s internal tests prior to posting the first round\u2019s lemmas, even the best publicly available models were only able to prove two. In the weeklong test period, various groups of amateurs and professional mathematicians tried to do better by building \u201cscaffolds,\u201d collaborative networks of LLMs that talked to one another to suss out mistakes. But all these efforts only solved one additional problem.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">A few different factors could explain why Google and OpenAI were able to (at least partially) solve eight problems versus the public\u2019s three. The companies could be using improved, unreleased versions of their LLMs or some more robust, internal scaffolds. Or the answers could rely on some undisclosed input from human mathematicians. (Google\u2019s team posted an explanation of its methodology. The team said this approach included \u201cabsolutely no human intervention\u201d\u2014a claim that First Proof\u2019s new requirements would verify.)<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">That\u2019s what the second round is meant to sort out, Williams says. \u201cThis was an experiment,\u201d she says, \u201cto get feedback from the community to figure out how to do a more formal round.\u201d<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">In addition to more robust human judging, this round will require that participants package models so the First Proof team can prompt them directly. \u201cIf it is not a public model, then we need to run it,\u201d Abouzaid says, \u201cbecause otherwise, it&#8217;s not clear what we&#8217;re testing.\u201d<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">It remains to be seen whether OpenAI and Google will comply, not to mention the many other LLM companies and AI-for-math start-ups that were conspicuously absent during the first round.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">In the coming months, First Proof and other AI benchmarks might help foretell the still-hazy fate of mathematics\u2014a tiny niche of the scientific world that suddenly has some of the Earth\u2019s wealthiest eyes trained upon it.<\/p>\n<p class=\"\" data-block=\"sciam\/paragraph\">\u201cOne of our main motivations is to make sure that we can tell young people what we expect the field to look like in a few years,\u201d Abouzaid says. \u201cAnd that requires understanding what these systems are actually capable of.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"<p>In the ongoing campaign by artificial intelligence companies to take over pure mathematics, another round is commencing. The team behind First Proof, an effort to benchmark the ability of large language models (LLMs) to contribute to research-level mathematics, has announced its next exam. For this second round, which it plans to roll out over the<\/p>\n","protected":false},"author":1,"featured_media":46868,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[50],"tags":[23832,2284,16405,8203,6734],"class_list":{"0":"post-46867","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-environment","8":"tag-foretell","9":"tag-future","10":"tag-improving","11":"tag-mathematicians","12":"tag-struggle"},"_links":{"self":[{"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=\/wp\/v2\/posts\/46867","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=46867"}],"version-history":[{"count":0,"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=\/wp\/v2\/posts\/46867\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=\/wp\/v2\/media\/46868"}],"wp:attachment":[{"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=46867"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=46867"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/naijaglobalnews.org\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=46867"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}