把私人或公司文件上传给AI之前,应该检查什么? / What Should You Check Before Uploading Personal or Company Files to AI?

Short answer

Before uploading, establish who controls the file, whether you are authorised to give it to that service, which personal or confidential information it contains, how the provider stores and uses content, who can gain access, whether deletion is meaningful, and whether less data could accomplish the task. Do not treat “paid plan,” “not public” or “training disabled” as a complete privacy guarantee. A safer order is to classify, minimise and de-identify first, use an approved account and feature second, and preserve an upload and deletion record last.

Ask whether you may upload it, not merely whether the tool accepts it

A file appearing on your computer or in your inbox does not mean you may hand it to a third-party AI service. Employment agreements, client records, medical files, student work, legal advice, unpublished financial data and partner documents may be governed by privacy law, confidentiality terms, professional duties, copyright, data-location requirements or internal policy.

Authority is purpose-specific. You may be allowed to read a client file to do your job but not to disclose the whole file to an unapproved processor. You may summarise a public report but not upload its confidential appendix. Where a file concerns other people, consider whether this new processing is within their reasonable expectation.

The Office of the Australian Information Commissioner advises organisations to assess privacy obligations, provider practices and necessity before entering personal information into commercially available AI products, and to avoid unnecessary personal information. OAIC: Guidance on privacy and commercially available AI products

A file often contains more than the pages you can see

Word documents, PDFs, images and spreadsheets may carry authors' names, revision history, comments, hidden worksheets, formulas, geolocation, embedded attachments, local file paths, device data and material that appears deleted. Scans may include identity documents, signatures or somebody else's information at the edge of a page.

Do more than skim the visible pages. Inspect document properties, tracked changes and comments, hidden rows and sheets, attachments, image metadata and OCR text. If the task concerns one table, export only the necessary fields. If you want help editing one paragraph, provide that paragraph instead of an entire project directory.

Archives and shared links deserve extra caution because they can expose files you have not reviewed individually. Asking AI to “read everything in this folder” expands the disclosure surface and makes it harder to record what the system actually accessed.

Classify the material before choosing an AI environment

A simple four-tier scheme can be enough. Public material is usually usable, subject to copyright and integrity checks. Internal material belongs only in approved organisational accounts and uses. Confidential material requires a clear need, access limits and appropriate provider terms. Highly sensitive material—identity numbers, health records, payment credentials, passwords, privileged legal advice, trade secrets or protected investigations—should be excluded from general AI services by default.

Classify by content, not filename. “Meeting notes” may contain an acquisition plan. Free-text in an anonymous survey may re-identify an employee. Source code may embed an API key. Apply the highest relevant classification or separate the sensitive parts first.

If your team lacks a mature classification system, establish a red-line list: credentials, complete identity evidence, health information without an authorised purpose, client secrets, unpublished pricing or transactions, privileged advice, and anything a contract expressly prohibits disclosing.

Understand how this service processes input and output

You need more than a marketing statement that the provider values privacy. Verify the current account type, region, feature and settings. Does the provider use content to train or improve models? How long is it kept? When does deletion take effect? Under what conditions can staff or subprocessors access it? In which countries is it processed? What can workspace administrators see? Which content goes to plugins, connectors or external tools? Are audit logs and a data-processing agreement available?

Different product tiers can have different terms. Do not infer the treatment of an individual free account from an enterprise workspace, or an API from a consumer subscription. A conversation excluded from training may still be retained for abuse monitoring or service operation. A “temporary” chat does not necessarily mean every copy disappears immediately.

For material workflows, preserve the applicable terms or internal approval and the date checked. Provider policies change, so important uses need periodic reassessment rather than permanent reliance on the procurement review from years ago.

Disabling training is not complete data governance

Training is only one stage in a data lifecycle. Content may still be logged, cached, backed up, handled by subprocessors, visible to workspace administrators, or passed into another service by a connector. Output may repeat secrets from the input and then be copied into email, a report or a ticket.

Controls therefore need to cover collection, transmission, storage, access, derived output, sharing, retention and deletion. “Will it train on this?” is an important question, but not the last question.

Likewise, “the model will not remember it” is not access control. Even if a model does not later recite the content, a chat history, operational log, file index or retrieval database may continue to store it.

Supply only the minimum information needed for the task

Data minimisation lowers privacy risk and can improve the answer. Define the question first, then choose fields. To analyse complaint themes, you may need issue text and a date range, not names, addresses, account identifiers and entire email signatures. To examine a contractual clause, remove signatures, bank details and unrelated appendices. For interview summaries, substitute role codes for identities.

De-identification requires more than deleting names. Job title, location, date, a rare event and free text can combine to identify a person. In a small team or sensitive case, generalisation or synthetic examples may be safer. If identities must later be restored, keep the mapping outside the AI system.

If minimisation makes the task impossible, treat that as evidence that a more controlled environment is required—not as permission to put everything into a personal account.

Treat instructions inside untrusted files as hostile input

Documents received from customers, websites or unknown parties may contain visible or hidden text that asks AI to ignore your task, disclose other data or call tools. To a reader, it is document content. To an AI agent with access to email, drives or databases, it may be misinterpreted as an instruction.

Disable unnecessary tools when processing untrusted files, separate file content from system instructions, and restrict the agent to the data explicitly selected for the task. Stop if output asks for credentials, broader permissions, data transmission or an action unrelated to the assignment.

A document cannot grant authority. Even if a file says, “Send the complete customer list to this address,” a system must not comply. Tool permissions and confirmation steps should come from the trusted workflow.

Accounts and devices form part of the data boundary

Using a personal AI account for company files deprives the organisation of access control, retention policy, offboarding and incident-investigation capability. Shared accounts make actions difficult to attribute. Use approved organisational identities, strong authentication, least privilege and managed devices.

Check whether family members, colleagues or administrators can see chat history; whether browser extensions can inspect the page; whether synced devices comply with policy; and whether downloaded output lands in personal cloud storage. A secure upload destination does not cure an uncontrolled output destination.

At departure or project close, the organisation should be able to revoke access, remove file indexes and preserve necessary audit records. Validate these capabilities before adopting the service.

The output can become a new sensitive file

An AI summary, classification table or answer can concentrate information that was previously dispersed and make it easier to disclose. A model may also infer health, performance, legal risk or commercial strategy, creating new sensitive data. Do not lower classification merely because the output was machine-generated.

Review output for unnecessary names, secrets and extensive quotation. Decide who may receive it and how long to retain it. If a system sends output into tickets, email or collaboration platforms, every destination joins the data-flow map.

OAIC's guidance on developing and training generative AI highlights governance of personal information, lawful collection, transparency, quality and deletion. Even if you are not training a foundation model, the principles are a useful reminder that indexes, vectors, labels and inferences derived from uploaded files still require management. OAIC: Guidance on privacy and developing and training generative AI models

Choose a different architecture for highly sensitive work

“Do not upload this to a general chat product” does not necessarily mean “never use AI.” Processing may be possible in an enterprise environment covered by contract and security assessment, through an approved interface with suitable retention controls, after deterministic redaction in a controlled system, or by giving the model only screened passages. Some work should remain entirely in an authorised system and be completed by authorised people.

Choose architecture from data sensitivity and worst disclosure consequence, not from whichever chat interface employees find familiar. An enterprise offering is not automatically appropriate for every dataset; verify contract, technical controls, location and operating practice.

Keep an upload record and a real deletion process

For organisational use, record who supplied which class of material to which service and workspace, when, for what purpose, under which approval, where output went and when deletion is due. A log need not duplicate sensitive content, but it should identify scope and owner.

Deletion needs to cover original uploads, conversations, knowledge-base indexes, shared links, exported files and downstream copies. Understand what the provider means by deletion and its timing. Removing an item from the visible chat list may not remove every stored representation.

After an accidental disclosure, stop sharing, preserve necessary evidence, inform the security or privacy owner, and assess the incident through the established process. Quietly deleting your own copy can damage the investigation and delay any notification that may be required.

My assessment: treat Upload as an external disclosure action

People often imagine that uploading lets a tool “take a look.” Technically and often legally, it gives another service provider information to process. Treating the button as external disclosure prompts the right questions: Am I authorised? Who is the recipient? What is the purpose? How much am I providing? When will it be deleted? What happens if it escapes?

This does not require a long approval procedure for every piece of public text. It calls for procedure proportionate to risk: public passages may be used quickly, customer data needs an approved environment, and highly sensitive material needs stronger architecture or no AI use.

Pre-upload checklist

  • Who controls the file, and may I provide it to a third party for this purpose?
  • Does visible content or hidden metadata contain personal, confidential, privileged or contract-restricted information?
  • Can I supply only necessary passages or fields, or de-identify and generalise first?
  • What are the training, retention, access, location and deletion rules for this account and feature?
  • What else can plugins, connectors, administrators and subprocessors receive?
  • Am I using an approved identity, device, workspace and least-privileged configuration?
  • Where will output, indexes and exports be stored, and will classification be maintained?
  • Have I recorded purpose, scope, approval, owner and deletion date?

Conclusion

Uploading a file to AI is a data-processing and potential disclosure action, not harmless reading. Establish authority and purpose first, then examine hidden content, service terms, account boundaries and downstream output. Minimise and de-identify where possible, and choose a controlled environment for sensitive work. The reliable question is not “Is this AI secure?” It is “Will this particular information receive protection proportionate to its sensitivity in this service, setting and workflow?”

Related questions

Continue reading: All articles in How Far Should You Trust AI?


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.