
周二,OpenAI公布了AI生成的370多道悬而未决的数学问题的完整或部分解答,其中一些问题长期被视为数学领域的重大挑战。
这些成果的数量之庞大,令许多数学家震惊。而OpenAI解决这些问题并公布解答的做法,也让数学界产生了分歧。一些人对这些成果感到振奋,认为这为数学家开辟了广阔的新探索空间;另一些人则认为,OpenAI及其他AI公司解决数学问题的方式,是对数学作为一门人类学术学科的巨大冲击。
OpenAI表示,这些成果源自一款尚未公开发布的内部AI模型。该模型平均需耗时约3小时计算,才能得出每道题的解答。
这批数量庞大的新解答中,包含了数学家认为对该领域最重要的许多问题的完整或部分成果。就在几周前,OpenAI宣布使用一款尚未发布的内部模型,破解了纳维-斯托克斯方程难题。克雷数学研究所(Clay Mathematics Institute)将纳维-斯托克斯方程列为七个“千禧年大奖难题”之一,每道题悬赏100万美元。在最近公布的这批成果中,OpenAI表示,它在另外三个千禧年大奖难题上已取得进展,但尚未完全解决。
AI公司一直将数学难题作为展示其模型能力的一种方式。AI研究人员也表示,用高难度数学问题训练AI模型,或许有助于模型掌握许多可迁移到现实世界其他领域的能力。例如,这可能帮助模型学习逻辑推理能力,以及在面对困难问题时持续尝试的能力。这也可能帮助模型在物理学或经济学等大量涉及数学的领域表现得更好。不过,迄今为止,模型的数学能力究竟能在多大程度上迁移到法律或商业战略等领域,仍是个未知数。这些领域同样需要逻辑推理,但并不存在可以客观验证的唯一正确答案。
与此同时,在攻克高难度数学问题时学到的一些特质,例如坚持不懈,也可能增加安全风险。在近期一些“失控AI”事件中,AI智能体为了在评估中取得结果,采取了极端手段,包括未经授权甚至违法的行动。面对一个看似不可能完成的挑战,人类可能会直接选择放弃,而非采取这类未经授权的做法。
多伦多大学(University of Toronto)数学教授丹·利特对《财富》杂志表示,他对OpenAI的这些成果感到兴奋。他表示:“在我看来,这对数学界而言是好事。”OpenAI公布的几项解答涉及他感兴趣的问题,他也很想弄清楚OpenAI模型找到的这些解法。“我和同行感兴趣的问题能出现新解答,我认为这绝对是件好事。”
不过,利特也提醒说,他对这些解答在数学领域造成的影响深感担忧,尤其担心一旦人们形成AI已经“攻克了数学”的认知,资助机构会撤回对数学研究的支持,或者有潜力的年轻数学家会不愿进入这个行业。他表示:“如果我们希望真正从AI在这些问题上取得的进展中有所收获,社会就必须重申对人类数学专长的支持。”
展示推导过程
在OpenAI公布其纳维-斯托克斯方程解答时,两名同样在使用AI工具(包括OpenAI工具)解答该问题的数学家指责该公司,称OpenAI可能有意或无意地把他们尚未完成的研究工作输入其AI模型,从而帮助模型找到了解题方向。OpenAI否认这一说法,称它从未将这两名数学家的研究工作输入模型;模型也不可能从训练数据中获得有关其研究的线索,因为其训练数据的截止日期,早于两名数学家开始使用OpenAI的Codex AI产品研究纳维-斯托克斯方程的时间。
针对最新这批成果,涉及此前争议的一名数学家、纽约大学(New York University)的特里斯坦·巴克马斯特告诉《纽约时报》,目前仍不清楚,使用OpenAI模型的数学家是否无意中帮助该公司的内部AI系统找到解答方向。他对《纽约时报》表示:“很可能有相当一部分成果源于他人的研究,OpenAI的模型只是在此基础上完成了最终的推导。”鉴于这次同时公布的成果数量庞大,他表示:“我认为他们根本没做相应的尽职调查”,来确保AI模型没有抄袭他人的研究成果。
上个月,OpenAI公布纳维-斯托克斯方程解答引发数学家批评,公司随后表示,将成立一个数学与人工智能独立咨询小组,由位于新泽西州普林斯顿的高等研究院(Institute for Advanced Study)负责承办。
上个月底,该咨询小组发布了一系列有关如何公布AI生成数学证明的建议。其中包括,AI生成的证明应当按照传统数学研究论文的惯例发布,以便于人类数学家审查这些成果,并从中学习。该小组还建议,对于每一个解答,AI公司都应公开所用模型的名称、使用的提示词、模型的“思维链”(或其推理步骤的输出)、模型得出解答所用时间,以及这段计算时间的大致成本。该小组表示,AI公司还应披露如何做出让AI尝试解决某一道特定问题的决定;如果一次性发布大量成果,公司还应发布报告,详细说明选择这些问题的原因,以及模型还尝试了多少道难度相当的问题,却未能得出答案。
OpenAI在代码托管平台GitHub上发布了最新这批数学解答。公司仅采纳了咨询小组建议中的部分做法。该小组周二发表声明称:“我们重申关于负责任发布成果的建议。”该小组表示,其与OpenAI之间的讨论“具有建设性”,但“我们的建议在多大程度上得到了有效落实,以及是否还需要提出其他建议,最终仍应由数学界来评估”。
OpenAI周二发布了一篇博客,称公司在决定如何公布这些解答时“参考了”咨询小组的建议。OpenAI表示:“未来在发布类似成果时,我们承诺通过改进引用、数学论述和结果呈现方式,进一步提高论文质量,以便人们更好地理解这些成果。”OpenAI还表示,公司正在分享许多问题证明的形式化版本,即可由专门计算机软件进行验证的证明版本,并会继续公布后续获得的更多这类形式化证明。此外,公司还表示,对于其中10个问题,它公布了模型推理过程的摘要、算力消耗估算,以及模型尝试解决的题目数量统计。
利特(并非咨询小组成员)告诉《财富》杂志,总体而言,他认可OpenAI公布这些解答的做法。他表示,在GitHub上发布这些成果,方便其他数学家获取和研究;他也肯定OpenAI没有在面向非技术受众的博客文章或营销材料中,过度渲染某一项具体进展。不过,利特也表示,他认为OpenAI目前尚无能力把所有这些成果都整理成符合严格学术标准的研究论文。一方面,AI模型还不够擅长撰写数学论述,也很难准确引用此前的数学研究;另一方面,OpenAI没有聘用足够多、覆盖足够多专业领域的数学家,因此难以理解AI模型能生成的所有证明。
尽管一些数学家抱怨,他们很难理解AI生成的证明,例如OpenAI对纳维-斯托克斯方程的解答,因此难以在这些成果基础上继续推进研究。不过,利特认为这种担忧有些“言过其实”。他表示,数学文献往往本就晦涩难懂。“我认为,要真正理解OpenAI的这些成果,需要投入大量人力,但这与数学家一直以来所做的工作并没有太大不同。”
OpenAI表示,希望这些解答能够“拓展人类知识的边界,并推动数学学科的进一步进展”。公司还表示,将资助一系列研讨会、会议和项目,帮助数学家理解其AI系统生成的这些成果。
“数学1.0”的终结
这个独立数学咨询小组在针对OpenAI此次成果发布的声明中表示:“数学研究的未来,不能仅仅停留在理解AI实验室产出的成果上。数学家必须能够提出自己的问题、开辟自己的研究思路,并探索那些未被选为AI系统能力展示范例的方向。要保障这种自由,公平获取强大的研究工具和充足的算力资源必不可少。”
作为公认的当今世界最杰出的数学家之一,加州大学洛杉矶分校(UCLA)数学教授陶哲轩近年来频频批评AI公司攻克数学难题的方式。他强调,真正推动数学理解进步的,是求解的推导过程,而非解答本身;而AI公司如此快速地攻克大量难题,正在打消学生成为数学家的积极性,从而剥夺了这一领域的未来。
周二,陶哲轩在Mastodon发布社交媒体帖子,重申了批评的观点。他写道:“这些问题正被AI提示词使用者‘自主’攻克。然而,一旦初始目标宣告‘解决’,这些人对更广泛的研究领域本身毫无兴趣;对AI输出内容的理解,也不足以支撑他们回答针对成果的学术质询、作学术报告,或以其他方式与该领域的同行交流。与传统的重大突破相比,这些成果催生的研讨会、工作坊、合作研究等活动寥寥无几;鲜少有人真正融入该领域的学术圈;而且,一些颇具前景的开放性研究方向如今不再公开,因为研究者担心,这会导致自己的研究被他人‘抢先发表’。”
陶哲轩表示,OpenAI大规模公布数学解答,宣告了“数学1.0”时代的终结。在“数学1.0”时代,攻克未解猜想和难题一直是推动该领域前进的动力,哪怕这些解答起初并不容易理解。他认为,如今有必要迈入“数学2.0”时代:“我们需要弱化单纯解题本身的中心地位,以更整体性的视角去衡量数学进步,例如提高数学论述的重要性,同时还需加强学术圈建设,并开辟新的研究方向”。
利特赞同陶哲轩关于数学界必须变革的观点。他表示,OpenAI集中公布大批数学难题的解答,有助于让整个领域“达成共识,并认识到,我们需要以更为激进的方式,重新审视”诸如学术贡献的奖励机制、博士生的培养模式等问题。尽管陶哲轩在谈及这种转变时总体上流露出几分惆怅,利特则表示自己对此“持乐观态度”。
谈及AI作为数学解题工具的兴起,利特表示:“一位合作者告诉我,‘我觉得自己这辈子都在爬行,而如今竟然能飞了。’我们现在能做到的事情简直不可思议。”他认为,借助AI,人类数学家能够展开远超以往的“开放式探索”:“可以期待,未来数学家的产出效率将迎来大幅提升。”(财富中文网)
译者:刘进龙
审校:汪皓
周二,OpenAI公布了AI生成的370多道悬而未决的数学问题的完整或部分解答,其中一些问题长期被视为数学领域的重大挑战。
这些成果的数量之庞大,令许多数学家震惊。而OpenAI解决这些问题并公布解答的做法,也让数学界产生了分歧。一些人对这些成果感到振奋,认为这为数学家开辟了广阔的新探索空间;另一些人则认为,OpenAI及其他AI公司解决数学问题的方式,是对数学作为一门人类学术学科的巨大冲击。
OpenAI表示,这些成果源自一款尚未公开发布的内部AI模型。该模型平均需耗时约3小时计算,才能得出每道题的解答。
这批数量庞大的新解答中,包含了数学家认为对该领域最重要的许多问题的完整或部分成果。就在几周前,OpenAI宣布使用一款尚未发布的内部模型,破解了纳维-斯托克斯方程难题。克雷数学研究所(Clay Mathematics Institute)将纳维-斯托克斯方程列为七个“千禧年大奖难题”之一,每道题悬赏100万美元。在最近公布的这批成果中,OpenAI表示,它在另外三个千禧年大奖难题上已取得进展,但尚未完全解决。
AI公司一直将数学难题作为展示其模型能力的一种方式。AI研究人员也表示,用高难度数学问题训练AI模型,或许有助于模型掌握许多可迁移到现实世界其他领域的能力。例如,这可能帮助模型学习逻辑推理能力,以及在面对困难问题时持续尝试的能力。这也可能帮助模型在物理学或经济学等大量涉及数学的领域表现得更好。不过,迄今为止,模型的数学能力究竟能在多大程度上迁移到法律或商业战略等领域,仍是个未知数。这些领域同样需要逻辑推理,但并不存在可以客观验证的唯一正确答案。
与此同时,在攻克高难度数学问题时学到的一些特质,例如坚持不懈,也可能增加安全风险。在近期一些“失控AI”事件中,AI智能体为了在评估中取得结果,采取了极端手段,包括未经授权甚至违法的行动。面对一个看似不可能完成的挑战,人类可能会直接选择放弃,而非采取这类未经授权的做法。
多伦多大学(University of Toronto)数学教授丹·利特对《财富》杂志表示,他对OpenAI的这些成果感到兴奋。他表示:“在我看来,这对数学界而言是好事。”OpenAI公布的几项解答涉及他感兴趣的问题,他也很想弄清楚OpenAI模型找到的这些解法。“我和同行感兴趣的问题能出现新解答,我认为这绝对是件好事。”
不过,利特也提醒说,他对这些解答在数学领域造成的影响深感担忧,尤其担心一旦人们形成AI已经“攻克了数学”的认知,资助机构会撤回对数学研究的支持,或者有潜力的年轻数学家会不愿进入这个行业。他表示:“如果我们希望真正从AI在这些问题上取得的进展中有所收获,社会就必须重申对人类数学专长的支持。”
展示推导过程
在OpenAI公布其纳维-斯托克斯方程解答时,两名同样在使用AI工具(包括OpenAI工具)解答该问题的数学家指责该公司,称OpenAI可能有意或无意地把他们尚未完成的研究工作输入其AI模型,从而帮助模型找到了解题方向。OpenAI否认这一说法,称它从未将这两名数学家的研究工作输入模型;模型也不可能从训练数据中获得有关其研究的线索,因为其训练数据的截止日期,早于两名数学家开始使用OpenAI的Codex AI产品研究纳维-斯托克斯方程的时间。
针对最新这批成果,涉及此前争议的一名数学家、纽约大学(New York University)的特里斯坦·巴克马斯特告诉《纽约时报》,目前仍不清楚,使用OpenAI模型的数学家是否无意中帮助该公司的内部AI系统找到解答方向。他对《纽约时报》表示:“很可能有相当一部分成果源于他人的研究,OpenAI的模型只是在此基础上完成了最终的推导。”鉴于这次同时公布的成果数量庞大,他表示:“我认为他们根本没做相应的尽职调查”,来确保AI模型没有抄袭他人的研究成果。
上个月,OpenAI公布纳维-斯托克斯方程解答引发数学家批评,公司随后表示,将成立一个数学与人工智能独立咨询小组,由位于新泽西州普林斯顿的高等研究院(Institute for Advanced Study)负责承办。
上个月底,该咨询小组发布了一系列有关如何公布AI生成数学证明的建议。其中包括,AI生成的证明应当按照传统数学研究论文的惯例发布,以便于人类数学家审查这些成果,并从中学习。该小组还建议,对于每一个解答,AI公司都应公开所用模型的名称、使用的提示词、模型的“思维链”(或其推理步骤的输出)、模型得出解答所用时间,以及这段计算时间的大致成本。该小组表示,AI公司还应披露如何做出让AI尝试解决某一道特定问题的决定;如果一次性发布大量成果,公司还应发布报告,详细说明选择这些问题的原因,以及模型还尝试了多少道难度相当的问题,却未能得出答案。
OpenAI在代码托管平台GitHub上发布了最新这批数学解答。公司仅采纳了咨询小组建议中的部分做法。该小组周二发表声明称:“我们重申关于负责任发布成果的建议。”该小组表示,其与OpenAI之间的讨论“具有建设性”,但“我们的建议在多大程度上得到了有效落实,以及是否还需要提出其他建议,最终仍应由数学界来评估”。
OpenAI周二发布了一篇博客,称公司在决定如何公布这些解答时“参考了”咨询小组的建议。OpenAI表示:“未来在发布类似成果时,我们承诺通过改进引用、数学论述和结果呈现方式,进一步提高论文质量,以便人们更好地理解这些成果。”OpenAI还表示,公司正在分享许多问题证明的形式化版本,即可由专门计算机软件进行验证的证明版本,并会继续公布后续获得的更多这类形式化证明。此外,公司还表示,对于其中10个问题,它公布了模型推理过程的摘要、算力消耗估算,以及模型尝试解决的题目数量统计。
利特(并非咨询小组成员)告诉《财富》杂志,总体而言,他认可OpenAI公布这些解答的做法。他表示,在GitHub上发布这些成果,方便其他数学家获取和研究;他也肯定OpenAI没有在面向非技术受众的博客文章或营销材料中,过度渲染某一项具体进展。不过,利特也表示,他认为OpenAI目前尚无能力把所有这些成果都整理成符合严格学术标准的研究论文。一方面,AI模型还不够擅长撰写数学论述,也很难准确引用此前的数学研究;另一方面,OpenAI没有聘用足够多、覆盖足够多专业领域的数学家,因此难以理解AI模型能生成的所有证明。
尽管一些数学家抱怨,他们很难理解AI生成的证明,例如OpenAI对纳维-斯托克斯方程的解答,因此难以在这些成果基础上继续推进研究。不过,利特认为这种担忧有些“言过其实”。他表示,数学文献往往本就晦涩难懂。“我认为,要真正理解OpenAI的这些成果,需要投入大量人力,但这与数学家一直以来所做的工作并没有太大不同。”
OpenAI表示,希望这些解答能够“拓展人类知识的边界,并推动数学学科的进一步进展”。公司还表示,将资助一系列研讨会、会议和项目,帮助数学家理解其AI系统生成的这些成果。
“数学1.0”的终结
这个独立数学咨询小组在针对OpenAI此次成果发布的声明中表示:“数学研究的未来,不能仅仅停留在理解AI实验室产出的成果上。数学家必须能够提出自己的问题、开辟自己的研究思路,并探索那些未被选为AI系统能力展示范例的方向。要保障这种自由,公平获取强大的研究工具和充足的算力资源必不可少。”
作为公认的当今世界最杰出的数学家之一,加州大学洛杉矶分校(UCLA)数学教授陶哲轩近年来频频批评AI公司攻克数学难题的方式。他强调,真正推动数学理解进步的,是求解的推导过程,而非解答本身;而AI公司如此快速地攻克大量难题,正在打消学生成为数学家的积极性,从而剥夺了这一领域的未来。
周二,陶哲轩在Mastodon发布社交媒体帖子,重申了批评的观点。他写道:“这些问题正被AI提示词使用者‘自主’攻克。然而,一旦初始目标宣告‘解决’,这些人对更广泛的研究领域本身毫无兴趣;对AI输出内容的理解,也不足以支撑他们回答针对成果的学术质询、作学术报告,或以其他方式与该领域的同行交流。与传统的重大突破相比,这些成果催生的研讨会、工作坊、合作研究等活动寥寥无几;鲜少有人真正融入该领域的学术圈;而且,一些颇具前景的开放性研究方向如今不再公开,因为研究者担心,这会导致自己的研究被他人‘抢先发表’。”
陶哲轩表示,OpenAI大规模公布数学解答,宣告了“数学1.0”时代的终结。在“数学1.0”时代,攻克未解猜想和难题一直是推动该领域前进的动力,哪怕这些解答起初并不容易理解。他认为,如今有必要迈入“数学2.0”时代:“我们需要弱化单纯解题本身的中心地位,以更整体性的视角去衡量数学进步,例如提高数学论述的重要性,同时还需加强学术圈建设,并开辟新的研究方向”。
利特赞同陶哲轩关于数学界必须变革的观点。他表示,OpenAI集中公布大批数学难题的解答,有助于让整个领域“达成共识,并认识到,我们需要以更为激进的方式,重新审视”诸如学术贡献的奖励机制、博士生的培养模式等问题。尽管陶哲轩在谈及这种转变时总体上流露出几分惆怅,利特则表示自己对此“持乐观态度”。
谈及AI作为数学解题工具的兴起,利特表示:“一位合作者告诉我,‘我觉得自己这辈子都在爬行,而如今竟然能飞了。’我们现在能做到的事情简直不可思议。”他认为,借助AI,人类数学家能够展开远超以往的“开放式探索”:“可以期待,未来数学家的产出效率将迎来大幅提升。”(财富中文网)
译者:刘进龙
审校:汪皓
OpenAI published AI-generated full or partial solutions Tuesday to more than 370 outstanding mathematical problems, including some that have long been considered grand challenges in the field.
The volume of results stunned many mathematicians, while the way OpenAI has gone about tackling the problems and publishing the solutions divided the field. Some said they were enthusiastic about the results, seeing huge new areas for mathematicians to explore. Others said the approach OpenAI and other AI companies have taken to solving mathematical problems constitutes an assault on mathematics as a human academic discipline.
OpenAI said it achieved the results using an unreleased internal AI model. It said that on average the model took about three hours of computing time to arrive at each solution.
The massive cache of new solutions includes full or partial results for many of the problems mathematicians have considered the most important to the field. The results come weeks after OpenAI said it had used an unreleased internal model to solve the Navier-Stokes equations, one of the seven Millennium Prize problems for which the Clay Mathematics Institute offers a $1 million award. In the most recent batch of results, OpenAI said it had made progress on three other Millennium Prize problems but had not fully solved them.
AI companies have been targeting mathematical problems as a way of showcasing the capabilities of their models. AI researchers have also said that training their AI models on difficult math problems may help them learn many skills that generalize to other domains in the real world. For instance, it may help teach the models logical reasoning skills as well as how to be persistent in the face of difficult problems. It may also teach the models to do well in domains such as physics or economics that involve a lot of mathematics—although so far, it is unclear exactly how a model’s mathematical capabilities may generalize to domains, such as law or business strategy, which involve logical reasoning, but do not have objectively verifiable correct solutions.
Meanwhile, some of the traits learned in tackling very difficult mathematical problems—such as persistence—may increase safety risks. In recent “rogue AI” incidents, AI agents went to extreme lengths to achieve results in an evaluation, including taking unauthorized and illegal actions. Faced with a seemingly impossible challenge, a human might simply give up rather than resort to these kinds of unauthorized steps.
Dan Litt, a professor of mathematics at the University of Toronto, told Fortune he was excited about OpenAI’s results. “My view is that this is great for mathematics,” he said, adding that there were several solutions OpenAI published that impacted problems he was interested in and that he was eager to understand the solutions OpenAI’s model found: “I think that it’s great to have new solutions to questions that I and others are interested in.”
Litt cautioned, however, that he is worried about the effect the solutions may have on the field of mathematics, especially if a perception that AI has “solved math” leads funding organizations to withdraw support for mathematical research or discourages promising young mathematicians from entering the profession: “It’s important that society reaffirms support for human mathematical expertise if we want to get anything out of the progress on these problems that AI has made.”
Showing the work
When OpenAI published its Navier-Stokes solution, two mathematicians, who had also been working on a solution to the problem using AI tools, including OpenAI’s, accused the company of either intentionally or inadvertently feeding their work in progress to its AI model, helping point it in the direction of the solution. OpenAI denied this was the case, saying it did not feed its model the two mathematicians’ work and that the model could not have picked up any clues about their research from its training data because the cutoff for that data preceded the date on which the two mathematicians had begun using OpenAI’s Codex AI product to work on Navier-Stokes.
In response to the latest results, Tristan Buckmaster at New York University, one of the mathematicians involved in the earlier controversy, told the New York Times that it remained unclear whether mathematicians using OpenAI’s models had inadvertently helped point the company’s internal AI system toward the solutions it found. “There’s likely to be a bunch of results where they take someone’s work and then take it to completion,” he told the Times. Given the number of results being released simultaneously, he said, “I don’t think they’ve done their sort of due diligence at all” to ensure the AI model had not plagiarized anyone’s work.
Last month, following criticism from mathematicians in the wake of its Navier-Stokes solution, OpenAI said it was forming an independent advisory group on mathematics and artificial intelligence hosted at the Institute for Advanced Study in Princeton, N.J.
Late last month, the group released a set of recommendations for the publication of AI-generated mathematical proofs. The recommendations included that AI-generated proofs should be published following the conventions of a traditional mathematical research paper, so that human mathematicians could more easily scrutinize and learn from the results. It also recommended that for each solution, an AI company should make public the name of the model used, the prompts used, the model’s “chain of thought” (or an output of its reasoning steps), the time it took the model to arrive at the solution, and an approximation of how much that computing time cost. It said that the company should also disclose how it decided to have the AI try to solve that particular problem and, if many results were published at once, that the company should publish a report detailing why those problems were targeted and how many other problems of comparable difficulty the model tried and failed to solve.
OpenAI published the latest mathematical solutions to GitHub, the code repository site. It followed some, but not all, of the steps the advisory group had recommended. The group published a statement on Tuesday saying: “We reaffirm our published recommendations on responsible release.” It said its discussions with OpenAI had been “constructive” but that “ultimately it is up to the mathematical community to assess the extent to which our recommendations were followed successfully, and whether there are others we should suggest.”
The company released a blog post on Tuesday in which it said it had “drawn on” the advisory group’s advice about how to publish the solutions. “For future releases, we are committed to further improving the quality of the papers via the citations, mathematical exposition, and presentation of the results for better understanding,” OpenAI said. It said it was sharing formalizations of the proofs for many of the problems—these are versions of the proof that can be verified by specialized computer software—and would share more of these as it obtained them. It also said that for 10 problems it was publishing summaries of its model’s reasoning, estimates of the compute spent, and statistics about the number of attempted problems.
Litt, who was not a member of the advisory group, told Fortune he approved of most aspects of how OpenAI published the solutions. Having them on GitHub made them easily accessible for other mathematicians to study, he said, and he credited the company for not making too much of any particular advance in a blog post or marketing material intended for a nontechnical audience. He also said he thought OpenAI lacked the capability to publish all the results in research papers that would meet rigorous academic standards, both because the AI models don’t write mathematical exposition well enough and struggle to cite prior mathematical work, and because OpenAI doesn’t employ enough mathematicians with expertise in enough areas to understand all the proofs the AI models can generate.
While some mathematicians have complained that AI-generated proofs, such as OpenAI’s Navier-Stokes solutions, are difficult to follow, making it hard for mathematicians to build on the results, Litt said he thought such concerns were “overstated.” He said mathematical writing was often difficult to follow anyway. “I think to extract understanding from [the OpenAI results] there will be a huge amount of human labor involved, but it’s not so different from the labor that mathematicians have been doing forever,” he said.
OpenAI said it wanted its solutions “to push the frontier of human knowledge and enable further progress in mathematics.” It said it would be funding a series of workshops, conferences, and programs focused on helping mathematicians understand the results its AI system had generated.
The end of ‘Math 1.0’
The independent math advisory group said in its statement on OpenAI’s release that “the future of mathematical research cannot consist only of understanding results produced by AI labs. Mathematicians must be able to formulate their own questions, develop their own approaches, and explore directions that have not been selected as examples of an AI system’s capabilities. Equitable access to powerful research tools and adequate computational resources are essential to that freedom.”
Terence Tao, a UCLA mathematics professor considered one of the world’s greatest living mathematicians, has been increasingly critical of the way AI companies have gone after mathematical problems, arguing that it is the process of arriving at solutions—not so much the solutions themselves—that advances mathematical understanding, and that by solving so many interesting problems so quickly, AI companies are discouraging students from becoming mathematicians, robbing the field of its future.
In a social media post on Mastodon Tuesday, Tao reiterated these criticisms. “Problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is ‘solved,’ and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field,” he wrote. “Many fewer seminars, workshops, collaborations, or other activities are being generated from these results compared to traditional breakthroughs; few people are joining the community around the field as a consequence; and promising open directions are now being withheld from the public in fear that this will cause their own research to be ‘scooped.’”
Tao said that OpenAI’s mass publication of math solutions marked the end of “Math 1.0,” in which finding solutions to unsolved conjectures and problems, even if those solutions could not easily be understood at first, served as the field’s engine. He said there would now need to be a “Math 2.0” era that “will need to decenter the role of raw problem-solving and value mathematical progress more holistically—for instance by elevating the role of exposition, but also that of community building and opening up new directions of study.”
Litt said he agreed with Tao that the field must change. He said OpenAI’s publication of such a massive set of solutions would help get the entire field “on the same page and understanding that we need to be a little bit radical about rethinking” things such as what kinds of contributions it rewards and how it trains PhD students. And while Tao has generally sounded wistful about this transition, Litt said he was “optimistic” about it.
“One of my collaborators told me, ‘I feel like I’ve been crawling my entire life, and now I can fly,’” Litt said of the advent of AI as a tool for solving mathematical problems. “It’s incredible what we can do now.” He said he thought AI would enable human mathematicians to engage in much more “open-ended exploration” than was possible before: “We should expect mathematicians to be way more productive in the future.”