Before Giving Organisational Knowledge to AI, What Must Be Put in Order?

After AI Enters the Workflow · Season Two: “From Personal Tool to Organisational Capability” · Article 4

A company connects its shared drives, policy library and project folders to an AI assistant so employees can ask questions directly. The demonstration is impressive. In production, the assistant encounters three versions of the same policy. A file marked “final” was never approved. A departmental guide conflicts with a signed contract. Instructions left by a former employee rank above newer records.

The AI did not create these contradictions. It made visible the knowledge problems that experienced staff had quietly navigated. A colleague once knew to say, “Ignore that file and ask this person.” When AI becomes the doorway, those oral corrections are absent. Faster retrieval can make the wrong answer easier to adopt.

Before giving organisational material to AI, the organisation must first decide what makes that material knowledge rather than merely stored text.

Document volume is not knowledge quality

A shared drive contains facts, drafts, conversations, templates, old versions, personal notes and formal decisions. All contain information; they do not carry equal authority.

AI retrieval can select material by words, semantic similarity and other signals. Similarity is not applicability. An old policy may be more detailed than its replacement and therefore produce a better-looking answer. A meeting transcript may contain an unambiguous sentence that records only a rejected proposal.

At minimum, organisational records need six kinds of structure: provenance, owner, scope, effective date, approval status and supersession. Without them, the system uses textual relevance as a poor substitute for institutional order.

Identify authoritative sources before building retrieval

Not every kind of knowledge belongs in one undifferentiated repository. Customer identity, contracts, policy, project status and professional guidance may each have a different system of record and a different authorised maintainer.

A better approach begins with an authority map. Which system controls each class of fact? Who may change it? What prevails when sources conflict? How long are historical versions retained? Which roles may access the material?

AI can retrieve across systems, but its answers should preserve their boundaries. It can say, “According to the approved policy effective July 2026.” It should not merge a contractual obligation and internal advice into one unsupported sentence.

The NIST AI RMF asks organisations to understand data and system context, document risk and manage it throughout the lifecycle. Its governance function also covers roles, documentation and third-party data and software supply chains. NIST AI RMF Core

Clean current relationships before trying to clean everything

“Organise all information before using AI” is impractical. Large organisations will always have legacy material and incomplete records. The priority is to repair the domains that feed frequent or consequential work.

Five actions provide a workable beginning: mark the current authoritative version; label superseded records as historical; assign an owner to orphaned material; record a review date; and expose significant conflicts as unresolved issues.

Old records do not always need deletion. Audit, research and dispute resolution may require history. The system must tell AI that a record is historical rather than allowing history to impersonate the present.

Documents should describe their own purpose and limitations

Machine-learning researchers proposed datasheets and model cards so that users could understand how a dataset or model was created, where it applied and what limitations it carried.

“Datasheets for Datasets” proposes documenting motivation, composition, collection process and recommended uses to improve communication between dataset creators and consumers. Microsoft Research, “Datasheets for Datasets”

“Model Cards” recommends recording intended uses, evaluation conditions and performance across relevant groups, reducing the chance that a model is applied in an unsuitable context. Google Research, “Model Cards for Model Reporting”

The same idea can be applied to organisational knowledge. A policy collection needs a knowledge card describing coverage, authoritative owner, update schedule, permitted uses, known gaps and the method for resolving conflict.

Access control must exist before retrieval

If an employee cannot open a personnel file, an AI system should not disclose its contents merely because it can retrieve them semantically. A subtler problem is inference: the system may derive sensitive information about a person from several individually accessible records.

The OAIC notes that generating or inferring personal information can itself amount to collection of personal information, and that organisations must consider necessity, accuracy and whether a product is appropriate to the intended purpose. OAIC, “Guidance on privacy and commercially available AI products”

Permissions cannot be checked only when a user opens the original file. Retrieval, summarisation, caches, logs and generated answers must preserve the relevant boundary. Otherwise, AI becomes a natural-language route around the source system’s controls.

A knowledge base requires continuing maintenance, not a one-time import

On the day integration is complete, the knowledge base begins to age. Policies change, contacts move, projects end and exceptions are approved.

Each knowledge domain needs a content owner, review cycle and change triggers. The system should identify authoritative records that have not been reviewed, record which downstream answers depend on a changed source, and rerun relevant tests after material updates.

The UK NCSC’s secure-AI guidance recommends identifying and version-controlling assets such as models, data, prompts, documentation and evaluations, with the ability to restore a known good state. NCSC, “Secure development”

An AI knowledge service is therefore not merely a search project. It combines records management, access governance and operational maintenance.

Prove the knowledge chain in a bounded domain

An organisation can begin with a well-defined domain—approved employee travel policy, for example—rather than connecting every shared drive. It can build a test set containing ordinary questions, exceptions, obsolete documents and users with different permissions. The evaluation should establish whether the system cites the current version, states the applicable date, stops when evidence conflicts and returns an answer appropriate to the user’s access.

The test is not merely whether an answer sounds correct. Reviewers should inspect what it cited, what it omitted, whether it disclosed restricted material and whether a later knowledge update makes old answers identifiable. Exceptions found in real use can then enter the test set. Curation becomes a maintenance cycle linked to actual questions rather than a one-time clean-up before launch.

If the organisation cannot demonstrate these properties in one domain, broader connection will only multiply invisible error sources. The aim of a bounded pilot is not to show how much the AI appears to know. It is to prove that the organisation can maintain the chain from record to authority, answer and correction.

Putting knowledge in order also means deciding what must no longer be retrieved. Superseded policy, withdrawn advice, unverified meeting notes and personal information due for lawful deletion should not remain in scope merely because it might someday be useful. Retention, archival, deletion and legal-hold rules must propagate into the AI index. Knowledge governance is not only the work of making correct material easier to find. It is also the capacity to remove expired or impermissible material from the decision process in a way that can be demonstrated.

Conclusion: establish responsibility for knowledge before expanding AI access

An organisation does not need to finish an endless clean-up of every file before obtaining value from AI. It does need to give its most important knowledge identifiable authority, time, scope, permissions and ownership.

The minimum rule is:

Material should support formal AI-assisted work only when it can answer: who maintains it, when does it apply, why is it authoritative, who may use it, and what supersedes it?

AI can help detect duplicates, conflicts and gaps. It can lower the cost of curation. It cannot decide on the organisation’s behalf which record carries institutional authority. The wider the access, the less that decision can remain hidden in the memories of a few experienced employees.

Primary sources and further reading

Continue reading: Explore the After AI Enters the Workflow series.


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.