
没人愿意收到一笔突如其来的巨额账单。但今年,许多企业在部署Claude Code、Codex这类热门AI编码智能体,执行耗时越来越长的自主任务时,恰恰遭遇了这种情况。与聊天机器人对话不同,这类智能体可持续运行数小时,反复调用前沿大模型,在不知不觉中消耗数百万词元(token)。开发人员出去吃顿午饭,回来就发现智能体已在模型推理(即模型输出)上烧掉数千美元。最终结果就是:账单高得令人咋舌。
一项最新研究揭示了这一影响:62%的企业表示,过去一年中,AI方面的意外支出实实在在改变了企业的商业决策。其中40%的企业称该问题需上报董事会层级处理,33%的企业紧急冻结相关开支,25%的企业直接推迟或取消AI项目。
这就是AI模型路由器突然成为企业科技领域最热门赛道的原因。该软件可帮助企业根据任务需求,以合理成本选用对应的AI模型。企业无需将所有请求都发送给成本最高的前沿大模型,而是依据任务需求,在智能体工作流程的各环节中,通过人工配置或自动化机制,遴选成本、响应速度与性能综合表现最优的模型。企业表示,这类智能模型路由可实现推理成本两位数降幅,部分场景降幅可达30%。
一大批初创公司纷纷涌入模型路由赛道,但采取的技术路线各不相同。OpenRouter等公司提供可接入数百种AI模型的市场和统一网关——据报道,该公司正与Stripe进行收购谈判,估值最高可达100亿美元;Not Diamond则自动为请求匹配最适配任务的模型。LiteLLM在内的其他公司,则帮助企业构建并管理自有路由基础设施。赛富时(Salesforce)、Databricks这类大型厂商也正在自家AI平台内置路由功能。据悉,Cursor、Ramp以及Meta都在自研模型路由器,就连视频初创公司Runway也推出了相关产品。
高词元消耗型智能体推高需求
OpenRouter联合创始人兼首席运营官克里斯·克拉克(Chris Clark)在接受《财富》杂志采访时表示,Anthropic旗下Claude Code等工具(2025年年中发布)产生的巨大输出需求,推动了这股热潮。他解释道,2024年至2025年间,企业高管层“力推”AI落地,但直到今年,相关技术才真正落地,工具链和智能体工作流实现突破,不再局限于简单对话。“AI模型使用工具、调用不同系统,并执行操作。”他说道。但与此同时,智能体的词元消耗量也大幅攀升,对于按词元付费的企业而言,成本压力与日俱增。
“此前几乎没有企业为此划拨预算,”克拉克称,“大家普遍抱着一种尽可能多用的心态。”
但这对OpenRouter而言却是业务利好。“我们一直坚信,AI词元将成为所有企业的一项重大运营开支。”他表示,这将促使几乎所有公司寻求降本方案。
克拉克解释,并非所有任务都需要性能最强大、最昂贵的AI模型,按需挑选模型确实能带来实实在在的益处。“我们向客户传递智能饱和这一概念:当一款新前沿模型问世,你让智能体切换到这款最新、最强大的模型,性能真的会提升吗?”他解释道。很多任务已经处于“智能饱和”状态,也就是说,即便使用最新、最强、最智能的模型,性能也不会提升,因为旧模型就足以完成任务。
Not Diamond(SAP等大型企业客户的合作方)首席执行官托马斯·埃尔南多·科夫曼(Tomás Hernando Kofman)表示,如今绝大多数企业不分场景,默认调用性能最强大的模型,造成资源浪费。不过他补充道,如果在错误的时间选用能力不足的小模型,同样会推高成本,因为小模型完成相同工作量耗时更长。自动化模型路由系统正是为了解决这一痛点。
Not Diamond将这一自动化调度类比为机器人技术。正如机器人需协调数千个小动作才能完成一项任务,AI编码智能体在长周期工作流的不同阶段,也需要调用不同模型。其路由系统不会孤立优化每个提示词,而是根据请求复杂度、对话历史及其他信号,预判下一步最合适的模型与推理级别。
不止于成本控制
尽管企业如今部署AI模型路由器是为了遏制推理成本飙升,但赛富时总裁兼首席架构师戴维·沃德(David Ward)认为,这仅仅是第一阶段。
“先解决账单过高的问题,”他说,“这当然是优先事项,但还有其他优先事项也需要实现。”
沃德预判,随着时间推移,路由决策将超越模型选择,扩展到信任、合规、治理以及可量化的业务成果。路由器除了选用适配的模型外,还可界定智能体的企业数据访问权限,同时评估微调后的开源模型能否胜任特定任务。
“这不只是为哪项工作选哪个模型,而是哪种工具、哪种技能能带来我想要的可量化成果,"他说,“行业将围绕这些议题展开讨论。”(财富中文网)
译者:中慧言-王芳
没人愿意收到一笔突如其来的巨额账单。但今年,许多企业在部署Claude Code、Codex这类热门AI编码智能体,执行耗时越来越长的自主任务时,恰恰遭遇了这种情况。与聊天机器人对话不同,这类智能体可持续运行数小时,反复调用前沿大模型,在不知不觉中消耗数百万词元(token)。开发人员出去吃顿午饭,回来就发现智能体已在模型推理(即模型输出)上烧掉数千美元。最终结果就是:账单高得令人咋舌。
一项最新研究揭示了这一影响:62%的企业表示,过去一年中,AI方面的意外支出实实在在改变了企业的商业决策。其中40%的企业称该问题需上报董事会层级处理,33%的企业紧急冻结相关开支,25%的企业直接推迟或取消AI项目。
这就是AI模型路由器突然成为企业科技领域最热门赛道的原因。该软件可帮助企业根据任务需求,以合理成本选用对应的AI模型。企业无需将所有请求都发送给成本最高的前沿大模型,而是依据任务需求,在智能体工作流程的各环节中,通过人工配置或自动化机制,遴选成本、响应速度与性能综合表现最优的模型。企业表示,这类智能模型路由可实现推理成本两位数降幅,部分场景降幅可达30%。
一大批初创公司纷纷涌入模型路由赛道,但采取的技术路线各不相同。OpenRouter等公司提供可接入数百种AI模型的市场和统一网关——据报道,该公司正与Stripe进行收购谈判,估值最高可达100亿美元;Not Diamond则自动为请求匹配最适配任务的模型。LiteLLM在内的其他公司,则帮助企业构建并管理自有路由基础设施。赛富时(Salesforce)、Databricks这类大型厂商也正在自家AI平台内置路由功能。据悉,Cursor、Ramp以及Meta都在自研模型路由器,就连视频初创公司Runway也推出了相关产品。
高词元消耗型智能体推高需求
OpenRouter联合创始人兼首席运营官克里斯·克拉克(Chris Clark)在接受《财富》杂志采访时表示,Anthropic旗下Claude Code等工具(2025年年中发布)产生的巨大输出需求,推动了这股热潮。他解释道,2024年至2025年间,企业高管层“力推”AI落地,但直到今年,相关技术才真正落地,工具链和智能体工作流实现突破,不再局限于简单对话。“AI模型使用工具、调用不同系统,并执行操作。”他说道。但与此同时,智能体的词元消耗量也大幅攀升,对于按词元付费的企业而言,成本压力与日俱增。
“此前几乎没有企业为此划拨预算,”克拉克称,“大家普遍抱着一种尽可能多用的心态。”
但这对OpenRouter而言却是业务利好。“我们一直坚信,AI词元将成为所有企业的一项重大运营开支。”他表示,这将促使几乎所有公司寻求降本方案。
克拉克解释,并非所有任务都需要性能最强大、最昂贵的AI模型,按需挑选模型确实能带来实实在在的益处。“我们向客户传递智能饱和这一概念:当一款新前沿模型问世,你让智能体切换到这款最新、最强大的模型,性能真的会提升吗?”他解释道。很多任务已经处于“智能饱和”状态,也就是说,即便使用最新、最强、最智能的模型,性能也不会提升,因为旧模型就足以完成任务。
Not Diamond(SAP等大型企业客户的合作方)首席执行官托马斯·埃尔南多·科夫曼(Tomás Hernando Kofman)表示,如今绝大多数企业不分场景,默认调用性能最强大的模型,造成资源浪费。不过他补充道,如果在错误的时间选用能力不足的小模型,同样会推高成本,因为小模型完成相同工作量耗时更长。自动化模型路由系统正是为了解决这一痛点。
Not Diamond将这一自动化调度类比为机器人技术。正如机器人需协调数千个小动作才能完成一项任务,AI编码智能体在长周期工作流的不同阶段,也需要调用不同模型。其路由系统不会孤立优化每个提示词,而是根据请求复杂度、对话历史及其他信号,预判下一步最合适的模型与推理级别。
不止于成本控制
尽管企业如今部署AI模型路由器是为了遏制推理成本飙升,但赛富时总裁兼首席架构师戴维·沃德(David Ward)认为,这仅仅是第一阶段。
“先解决账单过高的问题,”他说,“这当然是优先事项,但还有其他优先事项也需要实现。”
沃德预判,随着时间推移,路由决策将超越模型选择,扩展到信任、合规、治理以及可量化的业务成果。路由器除了选用适配的模型外,还可界定智能体的企业数据访问权限,同时评估微调后的开源模型能否胜任特定任务。
“这不只是为哪项工作选哪个模型,而是哪种工具、哪种技能能带来我想要的可量化成果,"他说,“行业将围绕这些议题展开讨论。”(财富中文网)
译者:中慧言-王芳
No one likes a surprise sky-high bill. But that’s exactly what many companies have faced this year as they deploy popular AI coding agents like Claude Code and Codex for increasingly long-running, autonomous tasks. Unlike a chatbot conversation, these agents can work for hours, repeatedly calling frontier models and quietly racking up millions of tokens. A developer can go to lunch and return to find that his agent has spent thousands on the inference, or output of the model. The result? Sticker shock.
A new study illustrates the impact: 62% of organizations said an unexpected AI expense materially altered a business decision over the past year. Among them, 40% said the issue required board-level escalation, 33% implemented emergency spending freezes, and 25% delayed or canceled an AI initiative outright.
That’s why AI model routers, software that allows organizations to choose the AI model for the right task at the right cost, have suddenly become one of the hottest areas in enterprise tech. Rather than sending every request to the most expensive frontier model, companies can either manually define or automatically choose the model that offers the best combination of cost, speed, and performance for each step of an agent’s work. Companies say this kind of intelligent model routing can reduce inference costs by double-digit percentages – in some cases up to 30%.
A slew of startups have rushed into the model router space, though they’re taking different approaches. Companies like OpenRouter, which has reportedly been in acquisition talks with Stripe at a valuation of up to $10 billion, provide a marketplace and unified gateway to hundreds of AI models, while Not Diamond automatically routes requests to the model best suited for each task. Others, including LiteLLM, let enterprises build and manage their own routing infrastructure. Large vendors such as Salesforce and Databricks are also building routing capabilities into their AI platforms. Cursor, Ramp and Meta are reportedly working on their own model routers, and even video startup Runway has launched one.
Token-hungry agents have driven demand
OpenRouter co-founder and chief operating officer Chris Clark told Fortune that the vast demand for output from tools like Anthropic’s Claude Code, which was released in mid-2025, has driven the demand. Through 2024 and 2025, the C-suite was “pounding the table” for companies to adopt AI, he explained, but it wasn’t until this year that it all started falling into place because harnesses and agentic flows began advancing capabilities beyond chat. “It’s an AI model using tools, calling out to different systems, taking action,” he said. But agents are also more token-hungry, which for companies paying per-token, became more and more expensive.
“No one had any budgets in place,” Clark said. “It was sort of this maximalist attitude.”
But for OpenRouter, it’s been good for business. “Our durable belief is that AI tokens “will be a massive line item for every business’ operating expense,” he said, which will lead nearly every company to seek relief.
Not every task requires the smartest—or most expensive—AI model, Clark explained, so there are real benefits to being choosy. “We try to talk to people about the idea of intelligent saturation: When a new frontier model shows up and you switch that agent to the newest, smartest model, does the performance improve or not?” he explained. Many tasks are already ‘intelligence-saturated,’ meaning if agents use the latest and greatest, smartest model, the performance doesn’t improve because an older model can already tackle that task.
Still, most companies today simply default to the most powerful model for everything, which becomes wasteful, said Tomás Hernando Kofman, CEO of Not Diamond, which works with large enterprise clients such as SAP. However, he added that choosing the wrong small model at the wrong time can also be expensive because smaller models can take longer to get the same amount of work done. That’s where an automated model routing system can help.
Not Diamond likens the problem of automating this to robotics. Just as a robot must coordinate thousands of small actions to complete a task, an AI coding agent needs different models at different stages of a long-running workflow. Rather than optimizing each prompt in isolation, its routing system predicts the best model and reasoning level for the next step based on the complexity of the request, the conversation’s history, and other signals.
More than cost control
But while enterprises are adopting AI model routers today to rein in soaring inference costs, Salesforce president and chief architect David Ward argues that’s only the first phase.
“Let’s solve the big bill,” he said. “That’s certainly one priority, but there are other priorities that need to come into place as well.”
Over time, Ward expects routing decisions to expand beyond model selection to include trust, compliance, governance, and measurable business outcomes. Rather than simply choosing the right model, routers could help determine what enterprise data an agent can access and whether a fine-tuned open-source model is sufficient for a particular task.
“It’s not just which model for which job; it’s which tool, which skill gives me the measured outcome I want,” he said. “The industry will be having those conversations.”