Most companies eventually hit the same question in their AI programs: which model are we standardizing on?
The question looks like a technology choice, but the answer shapes your architecture.
The model you choose will not hold still. It will be updated, repriced, overtaken or retired, and new kinds of model keep appearing beside it.

Microsoft sells a model router, a product that picks a model for each request. In one update in August 2026, Microsoft removed five retired models from it and told customers who had chosen them to pick replacements. Model retirement is becoming ordinary infrastructure maintenance.
If replacing a model means changing twenty applications, reworking prompts, repeating security reviews and hoping the output still behaves the same way, you created a dependency the day you chose it.
I would ask a leadership team a different question. What should we own so the models can change underneath us?
One front door, many models
Imagine twelve AI pilots running across a company. Each team picked its own vendor, and each application connects straight to a model with its own credentials, prompts and logging. That setup feels manageable while they are pilots.
Then one becomes a customer-facing product. Three become internal workflows. Another starts making decisions that affect revenue. Security finds twelve different access patterns and twelve sets of logs nobody reads. Finance wants to understand the bill. A model is retired, and another provider gets much better at one of the tasks.
Now every connection is its own migration.
Uber ran into this early. Its engineers wrote that teams had each found their own way of connecting to models, which created redundant work. Uber built what it calls a GenAI Gateway, one consistent way for every team to reach OpenAI’s models, Google’s and the models Uber hosts itself. JPMorgan made the same call. “We don’t want to architect ourselves around one particular model,” said Derek Waldron, its chief analytics officer.
The industry calls this a gateway. I think of it as the company’s AI front door. Applications connect to the front door, and behind it, routing rules decide which model handles the work.

Standardize the connection, and let the models behind it change. Routine work goes to a cheaper model, hard work to a more capable one, and a workload moves when its model is retired. The pattern works inside one vendor too. With a multi-year contract, routing between that vendor’s small and large models is the same idea.
The front door becomes critical infrastructure of its own, so its security, its availability and the exceptions for features only one vendor offers all need an owner too.
What slows a swap down
The clearest test of switching happened at the Pentagon. After barring Anthropic’s Claude from its contract work in early 2026, the Pentagon moved about 90% of that work to other models in roughly six months. Its testers found the models responded differently to the same prompts. Joe Saunders, CEO of the security company RunSafe, named the cost: each model “requires validation, and in many cases, re-authorization before it can be used.”
A gateway makes models easier to replace. It does not make them equivalent.
Companies report the same thing. In a 2025 survey of 100 CIOs by the venture firm Andreessen Horowitz, one leader said “all the prompts have been tuned for OpenAI.”
Your ability to switch vendors is only as good as your ability to evaluate the replacement. Intuit, the maker of TurboTax, built a leaderboard that scores models on benchmarks drawn from its own work in tax, personal finance and accounting, and uses it to judge whether a newly released model is worth switching to.
For an important use case, I would start with 50 to 100 examples of the actual task before choosing a model. Run them through several candidates and compare quality, failure modes, speed and cost. A full evaluation program can come later. Fifty good examples are enough to keep model selection from turning into a demo contest. The business owner for the task should define what good looks like.
Model choice becomes a routing decision
Once a company owns its examples and separates applications from models, it stops making one model decision. It starts making model decisions task by task, each rechecked against the examples whenever a model changes.
Airbnb’s customer service agent runs on 13 models. In October 2025, Brian Chesky, Airbnb’s CEO, explained why OpenAI’s newest models see relatively little production use: “We use OpenAI’s latest models, but we typically don’t use them that much in production because there are faster and cheaper models.” A simple classification task may never need the most capable model. A hard reasoning task might.
A market has formed around that behavior. In August 2026, Stripe agreed to acquire OpenRouter, a service that routes requests across more than 400 models, citing “the pace at which models are released and repriced.”
Model selection turns into an ongoing optimization problem. The decision you used to make at procurement and revisit at renewal comes back every quarter as an operating decision. Model optionality has to be designed in before you need it.
How far toward owning
Picture a range. At one end, a company routes between one provider’s smaller and larger models. A step further, it puts its own front door in front of several providers. Further still, it runs open-weight models on its own cloud or hardware. Each step adds control and adds work, and the CFO and COO weigh that trade.

For most companies, owning the model itself will not be the important distinction. Owning the ability to replace it will. A company that stays with one vendor can still own its examples, prompts, routing rules and logs.
The same principle applies to gateways, evaluation platforms and agent memory. Rent them if that is the right trade, and keep the underlying tests, prompts, routing logic and logs portable.
Three questions for every AI use case
I would put three questions into every architecture and vendor review.

1. What is the lowest-cost model that meets the quality, speed and risk requirements for this task on our own examples?
You can see what one of your own tasks would cost on different models with the AI model cost calculator.
2. If this model disappeared tomorrow, what would we have to change?
One routing rule is a manageable dependency. Twenty applications is not.
3. If we changed vendors tomorrow, what would we lose?
Your data, test examples, prompts, routing rules and logs should not be on that list.
The CIO owns the front door. The CISO defines where models and data are allowed to go. Finance challenges the economics. The business owner defines what good looks like. Someone must also own the recurring decision to change models.
Even if one vendor pulls ahead for good, the test set tells you whether its next version still does your work.
Build for the change, not the model.
Sources
- Microsoft Learn, “What’s new in model router” (August 2026 update, retired models removed). learn.microsoft.com
- Uber Engineering, “Navigating the LLM Landscape: Uber’s Innovation with GenAI Gateway,” 2024-07-11. uber.com
- American Banker, “How JPMorganChase democratized employee access to gen AI,” 2025-05-22. americanbanker.com
- DefenseScoop, 2026-09-11. defensescoop.com
- Yahoo Finance / Bloomberg, Pentagon testing models to replace Claude, May 2026. ca.finance.yahoo.com
- Federal News Network, 2026-03-19. federalnewsnetwork.com
- a16z, “How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025,” 2025-06-10. a16z.com
- Intuit Engineering, “Intuit’s Custom LLM Leaderboard”. medium.com
- Customer Experience Dive, Airbnb’s AI customer service agent, 2025-08-08. customerexperiencedive.com
- Fortune, Brian Chesky on OpenAI’s tools, 2025-10-21. fortune.com
- Stripe, “Stripe agrees to acquire OpenRouter,” 2026-08-19. stripe.com

