
人工智能调用成本正大幅下降,然而企业人工智能开支仍在持续攀升。
这一核心悖论,正是麦肯锡高级合伙人唐吉·卡特林(Tanguy Catlin)与拉里·哈马莱宁(Lari Hämäläinen)近日在麦肯锡线上直播活动“提升代理式人工智能的经济效益”中探讨的议题。本场活动还重磅发布了公司《2026年人工智能现状》调研报告。
兼任麦肯锡数字业务负责人的哈马莱宁表示:“达到一定能力水平的人工智能服务成本正大幅下降。”
例如2023年推出的GPT-4,8K上下文API定价为每百万输出词元60美元。他解释称,如今,在主流基准测试中性能与GPT-4大致相当的模型,运行成本已降至当年的零头,但企业的人工智能投入仍在持续增加。
原因在于:单位智能产出的成本虽在暴跌,但企业对人工智能算力的消耗量呈爆发式增长。
哈马莱宁指出,虽然模型单位产出成本持续下降,但企业要求人工智能完成的推理任务与工作量大幅增加,尤其是借助智能体完成。与此同时,部分人工智能供应商也从效率提升中获益——实现利润率提升。
以软件开发为例,AI智能体能够反复检查、修改、重写整个代码库,其生成的代码量远超普通开发人员的常规工作量。
代理式人工智能的成本逻辑
哈马莱宁表示,企业才刚开始理解代理式人工智能的成本逻辑。传统软件执行单次任务的成本相对可预测,但智能体可通过不同路径达成同一目标,成本波动极大,同一任务不同运行的成本最高相差30倍。
成本大多消耗在生成最终结果前的推理与反复优化环节。系统设计尤为关键:哈马莱宁指出,选择单智能体还是多智能体架构,将极大地影响任务执行成本。
人工智能成本管理如今更侧重于从系统设计源头避免词元浪费。
他表示,管理者不应以单词元成本来衡量智能体表现,而应从任务维度进行整体评估:执行一项任务需投入多少成本、智能体的成功率如何、以及人工复核结果需耗费多少时间。
他指出,经验法则是:当核验智能体输出内容的时间,远低于人工从头完成这项任务的时间时,智能体才有落地价值。他补充道,如果人工完成一项任务需要一小时,而核验智能体结果只需六分钟,那么只要智能体的任务成功率高于10%,它就能创造价值。
他表示,更大的挑战在于重新设计配套工作流程,以充分利用智能体释放出的产能。
麦肯锡全球研究院董事卡特兰表示:“事实上,成本管控不存在单一抓手。”他梳理了企业管控人工智能支出的三大方面:
第一,企业需掌握开支明细,定位消耗资源的应用场景、业务部门、智能体、模型与人员;
第二,企业需优化工作流程:根据任务复杂度匹配适配模型,将请求分发至适配模型,缓存可复用上下文,减少不必要的工具调用与智能体循环;
第三,企业需强化采购管理:清理闲置授权、管理配额、与供应商协商合同条款,避免过度依赖单一模型或供应商。
但卡特兰提醒道,管控成本并不等同于一刀切。企业应确定人工智能投入回报最高的领域,并推进有的放矢的优化。(财富中文网)
译者:中慧言-王芳
人工智能调用成本正大幅下降,然而企业人工智能开支仍在持续攀升。
这一核心悖论,正是麦肯锡高级合伙人唐吉·卡特林(Tanguy Catlin)与拉里·哈马莱宁(Lari Hämäläinen)近日在麦肯锡线上直播活动“提升代理式人工智能的经济效益”中探讨的议题。本场活动还重磅发布了公司《2026年人工智能现状》调研报告。
兼任麦肯锡数字业务负责人的哈马莱宁表示:“达到一定能力水平的人工智能服务成本正大幅下降。”
例如2023年推出的GPT-4,8K上下文API定价为每百万输出词元60美元。他解释称,如今,在主流基准测试中性能与GPT-4大致相当的模型,运行成本已降至当年的零头,但企业的人工智能投入仍在持续增加。
原因在于:单位智能产出的成本虽在暴跌,但企业对人工智能算力的消耗量呈爆发式增长。
哈马莱宁指出,虽然模型单位产出成本持续下降,但企业要求人工智能完成的推理任务与工作量大幅增加,尤其是借助智能体完成。与此同时,部分人工智能供应商也从效率提升中获益——实现利润率提升。
以软件开发为例,AI智能体能够反复检查、修改、重写整个代码库,其生成的代码量远超普通开发人员的常规工作量。
代理式人工智能的成本逻辑
哈马莱宁表示,企业才刚开始理解代理式人工智能的成本逻辑。传统软件执行单次任务的成本相对可预测,但智能体可通过不同路径达成同一目标,成本波动极大,同一任务不同运行的成本最高相差30倍。
成本大多消耗在生成最终结果前的推理与反复优化环节。系统设计尤为关键:哈马莱宁指出,选择单智能体还是多智能体架构,将极大地影响任务执行成本。
人工智能成本管理如今更侧重于从系统设计源头避免词元浪费。
他表示,管理者不应以单词元成本来衡量智能体表现,而应从任务维度进行整体评估:执行一项任务需投入多少成本、智能体的成功率如何、以及人工复核结果需耗费多少时间。
他指出,经验法则是:当核验智能体输出内容的时间,远低于人工从头完成这项任务的时间时,智能体才有落地价值。他补充道,如果人工完成一项任务需要一小时,而核验智能体结果只需六分钟,那么只要智能体的任务成功率高于10%,它就能创造价值。
他表示,更大的挑战在于重新设计配套工作流程,以充分利用智能体释放出的产能。
麦肯锡全球研究院董事卡特兰表示:“事实上,成本管控不存在单一抓手。”他梳理了企业管控人工智能支出的三大方面:
第一,企业需掌握开支明细,定位消耗资源的应用场景、业务部门、智能体、模型与人员;
第二,企业需优化工作流程:根据任务复杂度匹配适配模型,将请求分发至适配模型,缓存可复用上下文,减少不必要的工具调用与智能体循环;
第三,企业需强化采购管理:清理闲置授权、管理配额、与供应商协商合同条款,避免过度依赖单一模型或供应商。
但卡特兰提醒道,管控成本并不等同于一刀切。企业应确定人工智能投入回报最高的领域,并推进有的放矢的优化。(财富中文网)
译者:中慧言-王芳
Good morning. Intelligence is getting radically cheaper, but enterprise AI bills keep climbing anyway.
That was the central paradox McKinsey senior partners Tanguy Catlin and Lari Hämäläinen tackled on Tuesday during a McKinsey Live virtual session, “Improving the Economics of Agentic AI,” which highlighted the firm’s State of AI in 2026 survey.
“Intelligence at a certain capability level is getting a lot more affordable,” said Hämäläinen, who is also a leader in McKinsey Digital.
For example, GPT-4 launched in 2023, with the 8K-context API priced at $60 per million output tokens. Today, models that perform at roughly the same level as GPT-4 on established benchmarks can be run for a fraction of the cost—yet companies are still spending more and more on AI, he explained.
But as the cost of producing a unit of intelligence collapses, the amount businesses consume is exploding.
Models are getting cheaper per unit of capability while enterprises ask them to perform vastly more reasoning and work, especially through autonomous agents. AI vendors are also capturing some of those efficiency gains through higher margins, Hämäläinen said.
In software development, for example, AI agents can repeatedly inspect, modify and rewrite entire codebases, generating far more code than a human developer would typically touch, he said.
The economics of agentic AI
Companies are only beginning to understand the economics of agentic AI, Hämäläinen said. Unlike traditional software, where the cost of running a task is relatively predictable, agents can take different paths to the same result, making costs highly variable. The same task can cost up to 30 times more from one run to another, he said.
Much of that cost comes from the reasoning and repeated refinement behind the final output. System design matters: choices such as using a single agent versus multiple agents can dramatically affect the cost of completing a task, Hämäläinen said.
AI cost management has become more about engineering systems that don’t waste tokens in the first place.
Rather than measuring agents by cost per token, he said leaders should evaluate them at the task level: how much a task costs to execute, how often the agent succeeds, and how much human time is required to verify its work.
As a rule of thumb, an agent can make sense when the time required to verify its output is a small fraction of the time it would take a human to complete the task from scratch, he said. If a task takes a human an hour, but an agent’s work can be verified in six minutes, an agent with a success rate above 10% could already begin to create value, he added.
The bigger challenge, he said, is redesigning the surrounding workflow to take advantage of the capacity the agent frees up.
Meanwhile, “The truth is, there is no single cost lever,” said Catlin, who is also a director of the McKinsey Global Institute. He identified three areas where companies can manage AI spending:
First, companies need visibility into which use cases, business units, agents, models and users are driving spend.
Second, they need to optimize workflows by matching model complexity to the task, routing requests to appropriate models, caching reusable context and limiting unnecessary tool calls and agent loops.
Third, they need greater sourcing discipline, including removing unused licenses, managing quotas, negotiating provider terms and avoiding excessive dependence on a single model or vendor.
But Catlin cautioned that the goal should not be indiscriminate cost-cutting. Companies should determine where AI spending delivers the highest returns and optimize toward those returns.