For years, choosing an AI model was mostly a vendor decision. Pick a model, negotiate access and build around it. That choice is becoming less permanent.

Organisations can now put several models behind the same application, download weights, host them in different locations and route work according to the task. The practical consequences are showing up in places that have little to do with the rhetoric around AI sovereignty.

Japan is testing models against each other inside its government AI system. In India, Sarvam's platform sends different tasks to different models. In Hong Kong, Animoca says switching its primary model cut inference costs by as much as 95%.

These are different implementations of the same change: the model is becoming a replaceable component that can be evaluated, routed and swapped.

Japan: evaluate before you buy

Japan's Digital Agency built Gennai (源内) as an in-house generative-AI system for government staff. Between May and July 2025, about 950 of roughly 1,200 Digital Agency staff used it, for more than 65,000 sessions. A large-scale trial for all ministries began in May 2026. As of 29 May 2026, about 100,000 government staff could use Gennai. The target is about 180,000. Full use is planned from fiscal 2027.

Staff were already working on capable models before the domestic trial. On 10 July 2026, Nikkei xTECH and ITmedia named the selectable menu as Amazon Nova Lite and Anthropic's Claude Haiku 4.5, Sonnet 4.6 and Opus 4.8. The Digital Agency's notice the same day confirms that domestic models will be compared with "other foundation models already in use" on Gennai, though those vendors were unnamed.

The comparison is built into live chat. A Digital Agency technical blog on 22 June 2026 sets out the method. When a staff member sends a first message, Gennai produces answers from both the frontier model already in the system and the domestic model under test. The user picks which answer is useful for work without being told which model wrote it. The session then continues on the frontier model. A second stage trains a preference model on those choices so domestic models can be ranked against one another. The results are meant to feed paid procurement from April 2027.

The Agency is inserting a comparison at the start of a real session, then returning the user to the model already in production. Evaluation happens on traffic that already exists.

In March 2026, the Agency selected seven domestic foundation models from 15 bids. Two later dropped out after vendor circumstances, leaving five under contract. Three of those models, from NTT Data, Fujitsu and Preferred Networks, are due to run inference on Sakura Cloud, the only Japanese vendor on the government's cloud. Side-by-side tests are scheduled for September to November 2026.

Gennai is where those models have to prove they can do the same work.

The Agency describes the trial in the language of 国産, or domestically produced, models and 自律性, or autonomy. The operation is more pragmatic: the state is learning how to judge models while continuing to use the ones already in the building.

India: route the task, not everything

As Japan uses live staff traffic to compare models before it buys, a Bengaluru company is already sending different jobs to different models.

Sarvam hosts inference in India for its own 105-billion-parameter model, Sarvam-105B, and for downloadable weights including GLM 5.2 and Gemma 4 31B. Its documentation also lists DeepSeek V4 Flash and Kimi K3. The processing can happen in Indian data centres even when the model weights come from elsewhere.

What happens after a request arrives is the more distinctive move. In the configuration Sarvam uses for its own agents, 36% of tasks go to Sarvam-105B and 64% go to GLM 5.2. The company says that mix cuts serving cost by roughly 40%. It also publishes a price comparison: Sarvam-105B at USD 0.80 per million blended tokens, against USD 4.50 for GPT 5.4 Mini and USD 9.00 for Gemini 3.5 Flash. These are company figures from its own agent setup, not an independently audited cost comparison.

Product documentation tells developers which model to pick by job. Indic-language work goes to Sarvam-105B. Long-context work goes to GLM or DeepSeek. Images go to Gemma 4.

That map is application architecture. The product no longer assumes a single lab. It classifies the incoming task, chooses a model and keeps the interface the user sees stable. Replacing a component can then be a configuration change. A cheaper or more capable model can take a class of jobs without a rebuild around it. The organisation still has to know which jobs are which, and still has to measure whether the cheaper path is good enough. Selection has moved from the contract into the software.

Sarvam sells India-resident inference alongside that mix. The 36/64 split is the company's own configuration, but it is a clear picture of what task routing looks like when the application can choose between models.

Hong Kong: change the model, change the economics

Evaluation and routing assume several models stay in play. The bill can also move when the primary one changes.

Animoca Brands, the Hong Kong-headquartered company, said in a 6 August 2026 shareholder update that its Minds platform, in alpha since February 2026, had moved its primary cognition model to MiniMax M3 and internally assessed a 90–95% reduction in inference costs.

Animoca's FAQ still lists OpenAI, Google and xAI as third-party processors. Minds can change the model that does most of the thinking and still call other providers. The economic event is the swap of the primary model.

If the primary model can be swapped, the surrounding software has to let the organisation make that change. Animoca says it did, and that the unit economics moved with it.

What changes when models become interchangeable?

Four things move at once. They are easier to miss if each market is read only as a local anecdote.

Cost

Inference is priced by volume and by which model receives the volume. When several capable models can do overlapping work, the dearest option is a choice. Routine jobs can take a cheaper path while harder jobs can keep a more expensive one. The bill can move while the product the user sees stays still.

The cheaper path only works if quality is checked on the same task. Evaluation and cost are the same operating problem.

Inference location

A prompt has to be processed somewhere. Residency rules, latency and procurement can all decide which building is allowed to produce the answer. Downloadable weights and regional hosts make it possible to move that processing onto a government cloud, an in-country cluster or an overseas inference service without rewriting the application. Where data is stored and where the answer is generated can be specified as separate conditions.

Japan is tying its trial to designated government infrastructure. Sarvam is selling inference that stays inside India. The reasons differ, but the operational choice is similar: the application can stay while the infrastructure handling the request changes.

Vendor lock-in

Lock-in does not disappear. It moves.

Dependence used to concentrate in the model contract. Prompts, adaptations and application glue were built around one lab. When several models can sit behind the same interface, that glue has to live above any single model.

The new dependence is on the evaluation, orchestration and deployment layer that makes the swap possible. Application teams are putting more of the product in the tests, the routing layer and the deployment environment. Labs will still compete to be one of several components.

Task routing

Not every request needs the same capability. Sending every request to the most capable model runs up the bill. Sending every request to the cheapest model fails on the jobs that need more.

The scarce knowledge is which model is good enough for which job, and when that changes. Building that map is engineering work. Keeping it current is operational work.

What remains locked

A model used to be something an organisation built around.

Increasingly, it can be something an organisation builds around the ability to change.

Japan is building a way to compare models. Sarvam is building a way to route between them. Animoca says changing the primary model can materially change the economics.

The model may be interchangeable. The machinery that decides when, where and why to use it is not.

What to watch

• Japan: whether Gennai turns evaluation into procurement. The September to November 2026 tests should show whether staff preference on live traffic translates into paid domestic-model contracts from April 2027. The more revealing signal is whether frontier models remain on the production path alongside the domestic ones.

• India: whether routing reaches regulated customers. A named bank, telco or public-sector customer using India-hosted inference across multiple models would turn the 36/64 configuration from a product demonstration into an operating approach adopted outside Sarvam.

• Hong Kong: whether the MiniMax saving holds as Minds scales. Later reporting should show whether MiniMax remains the primary cognition model and whether the surrounding processor mix continues to change. The useful signal is whether the economics stay attractive once the model sits inside a larger production system.

---