AI Philosophy Observations | Precaution Does Not Turn Private Evidence into Public Knowledge

On 1 September 2026, OpenAI announced that its unreleased Astra model had reached the “Critical” cybersecurity capability threshold in its Preparedness Framework. Under the company’s definition, a model at this level can, given suitable tools and access, find previously unknown vulnerabilities and develop functional exploits across well-protected systems without step-by-step human direction, or devise and execute a novel attack strategy from a high-level goal. Astra is the first OpenAI model formally classified at this level.

OpenAI said its evidence combined automated public and private benchmarks with expert-led assessments against hardened browsers and operating systems. The company reported a perfect score on ExploitBench. On an internal dataset containing 20 recently disclosed, high-severity V8 vulnerabilities, Astra reportedly achieved higher arbitrary-code-execution rates than GPT-5.6 Sol while using fewer output tokens, and it found and used two zero-day vulnerabilities in an exploit chain. OpenAI also reported that Astra built browser-escape and local privilege-escalation chains. The company said the vulnerabilities were being disclosed, but it did not publish the internal dataset, complete evaluation records or details of the zero-days. It plans to publish a full system card when Astra launches. Reuters independently confirmed the Critical classification, the plan for limited access and the absence of a specific release date, but did not independently reproduce the capability claims.

This creates a question that cannot be resolved simply by deciding whether to trust the company. When the decisive evidence remains inside the developer, while outsiders receive definitions, selected results and institutional testimony, has Astra’s Critical capability become public knowledge? If it has not, do restrictions and stronger safeguards lack an epistemic basis?

Four states need to be kept distinct. The first is the model’s actual possession of a capability, which is a fact about the object. The second is a developer’s judgement based on internal experiments, which may be well supported even though access to the supporting evidence is unevenly distributed. The third is public knowledge, in which a claim enters a wider structure of inspection, criticism, reproduction and correction. The fourth is a decision under risk, which asks not whether the claim has been conclusively established but which actions are acceptable under present uncertainty. These states are connected, but none can substitute for the others.

OpenAI may have strong internal grounds for managing Astra as a critical-capability system. It can inspect the model, evaluation environments, failed trials and expert records, and it bears direct exposure to failures in its infrastructure or misuse of the model. For the public, however, the same conclusion is available mainly through institutional testimony. Outsiders can inspect the threshold definition and selected aggregate results, but cannot determine whether the test set is representative, how success criteria were applied, how failures were distributed, or whether the capability remains stable in other environments. “OpenAI has classified Astra as Critical” is therefore a publicly confirmed fact. “Astra has been independently shown to possess Critical capability” is not.

That distinction does not require inaction until every item of evidence is public. Cybersecurity evaluation contains a practical conflict: publishing enough detail to reproduce zero-day findings and exploit chains may also amplify the risk being assessed. Philosophy of risk distinguishes decisions made with reasonably characterised probabilities from decisions in which relevant probabilities remain incomplete. Where the possible harm is severe and the cost of a false negative is high, stricter access controls, monitoring and refusal mechanisms can be justified on precautionary grounds. The basis is that the existing evidence is sufficient to change a security decision, not that the capability claim has already become uncontested public knowledge.

Precautionary action cannot, however, prove the original capability judgement. Restricting access may be reasonable, but it does not automatically make an internal evaluation more accurate. A company’s willingness to delay training or deployment and accept operational costs adds weight to the sincerity and practical seriousness of its assessment; it does not replace external scrutiny. Otherwise, a closed circle forms: an institution withholds evidence because it is dangerous, treats its costly precautions as confirmation that the danger is real, and then cites the reality of the danger as the reason secrecy must continue. Such a circle may sustain a management regime, but it cannot independently produce public knowledge.

Work on the social dimensions of scientific knowledge emphasises that objectivity does not arise only from the honesty of individual investigators. It also depends on a shared structure in which criticism can enter, objections can be formed, methods can be examined and conclusions can be revised. Research on trust in scientific expertise similarly distinguishes mere disclosure, public accessibility and mechanisms that permit meaningful critical engagement. Transparency need not mean releasing every item at once. In a security-sensitive field, it can be layered: independent evaluators operating under confidentiality can inspect private results; auditable methodological accounts can be published; versions and counterexamples can be recorded; and more technical material can be released after vulnerabilities are repaired. The essential issue is not the volume of information but whether external correction can genuinely affect the judgement.

In Sustenesis Theory, Difference first requires these epistemic states to remain distinguishable: internal evidence is not public evidence, institutional classification is not independent reproduction, and a precaution is not proof of fact. If these differences disappear, the public is left with a false choice between total acceptance and total rejection. Constraint refers to the conditions under which a claim can form, circulate and be tested. Secrecy may be a necessary constraint against harm, but without controlled external review, explicit thresholds and later disclosure, it also becomes a constraint that prevents correction of knowledge.

Sustained Coherence here is not the internal consistency of a report. It is the capacity of a judgement to remain viable across different evaluators, environments, counterexamples, model versions and deployment feedback. Astra currently has a strong institutional judgement supported by developer-held evidence and enacted through concrete safeguards. It does not yet have a public epistemic structure of equal strength. A system card, independent evaluations, vulnerability disclosures and deployment records could strengthen, narrow or revise that judgement.

My judgement is that OpenAI’s present evidence is sufficient to warrant stronger safeguards for Astra, particularly because underestimating cyber capability may produce harm that is difficult to reverse. It is not yet sufficient to turn the Critical classification into independently established public knowledge. The reasonable position is neither to demand complete publication before any action nor to stop asking for evidence because a risk has been declared. Precautionary constraint and epistemic confirmation should remain separate. The former may proceed first; the latter must continue to form through relations that support scrutiny, comparison and revision.

The boundary of this conclusion is equally important. It does not deny the truthfulness of OpenAI’s internal tests, and it does not infer that Astra’s capability has been exaggerated. It identifies the present limit of what the public can confirm. If the eventual system card still offers only summary conclusions, independent evaluators cannot access the decisive material, or results vary substantially across environments, public knowledge will not become complete merely because the model has launched. Conversely, controlled independent review, technical disclosure after remediation and continuing deployment evidence can reduce the evidential asymmetry without directly distributing offensive capability.

References

OpenAI, “Path to Astra: critical capabilities and frontier safeguards”, 1 September 2026
https://openai.com/index/path-to-astra/

Reuters, “OpenAI says upcoming model is so capable it requires stronger guardrails”, 1 September 2026
https://www.reuters.com/business/openai-says-upcoming-model-is-so-capable-it-requires-stronger-guardrails-2026-09-01/

Stanford Encyclopedia of Philosophy, “The Social Dimensions of Scientific Knowledge”
https://plato.stanford.edu/entries/scientific-knowledge-social/

Hanna Metzen, “Trust in Scientific Expertise and the Varying Demands of Value Transparency”, Canadian Journal of Philosophy, published online 28 May 2025
https://www.cambridge.org/core/journals/canadian-journal-of-philosophy/article/abs/trust-in-scientific-expertise-and-the-varying-demands-of-value-transparency/AA7B03CDB9EEEC77DF02CCEEFF426C3B

Stanford Encyclopedia of Philosophy, “Risk”
https://plato.stanford.edu/entries/risk/


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.