Why Doesn’t Buying a Better Model Create a Better AI System?

After AI Enters the Workflow · Season Two: “From Personal Tool to Organisational Capability” · Article 6

A customer-service AI repeatedly cites an obsolete policy. Management concludes that the model is not capable enough and buys a stronger, more expensive version. The new model writes more naturally and explains complex questions more fully, but it still cites the obsolete policy. The reason is simple: the retrieval collection was never repaired, and the old document still ranks first.

The organisation upgraded its most visible component without repairing the work chain that determined the result. The new model may make the same error more persuasive.

A model is an important component of an AI system. System outcomes also depend on task definition, data, retrieval, tools, permissions, interface, human review and operating governance. Equating a stronger model with a better system is like assuming that a more powerful engine will make a vehicle with defective brakes safer.

Model capability covers only part of the system

A stronger model may improve language understanding, complex reasoning, code generation or use of long context. Those gains are real. Yet work can fail before and after the model.

The input layer may select the wrong customer, an expired contract or an incomplete record. A tool may execute the wrong action. An interface may hide sources and uncertainty. Human reviewers may approve mechanically because volume is excessive. The organisation may never have defined completion.

System quality resembles a chain: if a critical link falls below the required standard, the whole use case fails. An excellent model cannot repair an incorrect permission, and it cannot confer institutional authority on a document.

Procurement demonstrations focus on what models do best

Vendor demonstrations normally use clean material, clear questions and prepared success cases. Buyers compare fluency, speed and general capability. They are less likely to test conflicting data, tool failure, unusual users and exit arrangements.

The UK Government AI Playbook calls for the right tool for the job and places goals, teams, procurement, lifecycle management, meaningful human control and assurance within the same development process. UK Government, “AI Playbook”

The procurement object should therefore be more than an API or licence. It is a service capability: which outcomes it must achieve in which context, which data it uses, how changes are notified, which logs are available, how it is tested and how the organisation can migrate at contract end.

A third-party model creates a supply chain; it does not assume institutional responsibility

Buying a model reduces the need for foundational research, but increases dependence on an external component. A provider may change the model, price, limits, data location or support policy. Output behaviour may change even when an API name remains stable.

The NCSC secure-AI guidance recommends assessing and monitoring the AI supply chain, documenting assets including models, data, software and external APIs, and preparing alternatives for mission-critical systems. NCSC, “Secure development”

The provider is responsible for the service and commitments it supplies. The adopting institution remains responsible for deciding whether that component suits the use case, how it is configured, which data may enter and how outputs will be used.

“This is the best model on the market” is not a risk-acceptance rationale.

Lock-in usually develops around the model

Organisations worry that they will be unable to replace a model. In practice, switching cost comes from the surrounding ecosystem: proprietary tool calls, instruction formats, vector stores, evaluations, monitoring, identity systems, contracts and employee habits.

OECD research on AI markets notes that the ability to switch models supports competition, while technical and other switching costs can create lock-in. OECD, “Developments in Artificial Intelligence Markets”

Reducing lock-in does not require operating several suppliers at all times. At minimum, the organisation should retain its knowledge, test cases, business rules and activity records in portable forms; isolate model calls behind defined interfaces; and specify data, logging, change-notification and exit obligations in contracts.

The organisation’s core asset should not be a conversation history held only in a supplier account.

Locate the bottleneck before upgrading

Quality failures can be diagnosed along the work chain.

First confirm that inputs are correct, complete and current. Then test whether retrieval found the authoritative material. Observe whether the model used that material correctly. Check whether tool calls respected permissions. Finally, examine how people interpreted and adopted the output.

An upgrade is justified when the error actually arises from model capability and a candidate model produces a material improvement on the organisation’s test set. Otherwise, replacement may simply repackage the same failure.

The NIST AI RMF puts contextual mapping before measurement and management: organisations first establish the purpose, affected people and risks, then choose controls. NIST AI RMF Core This is more dependable than selecting a use case backwards from a leaderboard.

A good system allows the model to remain replaceable

A mature architecture respects model differences without attaching every business rule to one model’s accidental behaviour.

It uses structured inputs and outputs, independent permission layers, authoritative knowledge, repeatable evaluations, version records and failure degradation. When a model changes, the organisation can rerun tests, compare cost and quality, and revert when criteria are not met.

Some tasks are better handled by deterministic rules or ordinary search and need no generative model. Others may use a model to propose candidates while a verifiable program performs the final calculation. Keeping non-AI components in the right places is often more reliable than asking a larger model to perform everything.

A stronger model can change the meaning of existing controls

An upgrade does not merely make the same answer more accurate. A new model may use tools more readily, process longer records, return a different structure or respond differently to constraints that worked before. A length limit, keyword filter or review interface designed around the old model may no longer control the new behaviour.

Replacement should therefore be treated as controlled change. Freeze the current configuration and baseline, compare critical tasks and failure cases in an isolated environment, inspect data flows and permissions, and determine whether training and human review must change. An improvement in average quality should not justify release if a new severe failure appears.

Stronger capability will also generate requests for new uses. Those require separate purpose and risk decisions. Purchasing a more capable component does not automatically authorise the original use case to expand its objective, data or permissions.

Every upgrade should also name a rollback window and the person authorised to restore the earlier configuration when anomalies appear.

System design must also include the capacity to exit. If a model supplier changes price, data terms, functionality or service region, the organisation should know which processes stop, which records can be exported and how long an alternative would take to validate. Without portable prompts, evaluation sets, interface contracts and knowledge assets, the “best model” can become an expensive dependency. Procurement should therefore compare not only present quality but migration cost, contractual protection and the means of sustaining critical work during supplier disruption.

Conclusion: organisations buy components but build capabilities

A stronger model can provide an important improvement. It does not automatically curate records, define tasks, set permissions, establish evaluation or assign responsibility.

The organisational rule should be:

A model upgrade is a system upgrade only when it addresses an identified capability bottleneck and improves adoptable outcomes in an evaluation of the complete workflow. Otherwise, it is a component replacement.

A strong AI system is not permanently attached to the most powerful model. It creates a verifiable, maintainable and replaceable structure among models, knowledge, tools and people. The provider supplies one part of the capability; the organisation must build the rest.

Primary sources and further reading

Continue reading: Explore the After AI Enters the Workflow series.


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.