Choosing the Right AI Tool · Article Four
Producing three thousand words with AI is easy. Enter a subject and, within minutes, an introduction, a set of sections and a conclusion appear. The grammar will generally be sound and the paragraphs will seem to connect. The difficult question is whether those three thousand words deserve to be published under an author’s name.
An article can avoid an obvious factual error while presenting uncertain evidence with too much confidence. Its paragraphs may each be plausible but fail to advance a shared question. A particularly common sign is excessive regularity. Five ideas at different conceptual levels are compressed into five short sentences with the same structure. The prose gains rhythm while the thought is rearranged for the sake of syntax. Another request to “polish” the draft often strengthens this artificial completeness.
The comparison of AI for long-form writing cannot be reduced to which system has the most attractive prose. Publishability emerges from a chain of work: how material is confirmed, how an argument is formed, how the draft is subjected to contrary examination, how language preserves the distinctions the author intends, and how the final version enters a traceable file and publishing process.
“Writing” contains several different assignments
Long-form production includes source discovery, close reading, argument design, drafting, fact checking, stylistic editing and publication preparation. A conversational product may participate in every stage. That does not mean asking it to perform all stages in one conversation is the most reliable arrangement.
Research should look for material the writer has missed. Drafting should usually remain within a selected evidentiary boundary. A fact checker must distrust the draft rather than continue finding more elegant language in its support. A style editor has another concern: rhythm, repetition, semantic density and authorial voice. A prompt that merely asks a tool to “research this subject and write an authoritative, natural and profound article” collapses those competing assignments into one generation. The result is often a fluent synthesis rather than a tested manuscript.
The meaningful unit of comparison is therefore a workflow, not a model name. ChatGPT Projects or Claude Projects can maintain a continuing writing context. Gemini Notebook can help inspect a selected source collection. Deep Research expands discovery. A file-based agent such as Codex can manage formats, versions, links and pre-publication checks. They do not perform the same work and need not compete for a single title of best writer.
The first step is to stabilise evidence, not generate prose
Stabilising the evidence does not prohibit new sources. It means that when an argument begins to form, the writer knows what the present evidence consists of, which items have been verified and which are merely leads.
A source register can be simple: title, institution or author, date, link, source type, the proposition it can support and any important limitation. Law, official statistics, product features, prices and current policy should lead back to primary pages. News is useful for discovering a case, but it should not carry a central institutional claim by itself when an official record is available.
At this stage, the most valuable output from deep research may be a candidate source map rather than a finished report. Screen the material it discovers. Remove content farms, undated pages and paraphrases whose originals cannot be reached. Personal PDFs and notes can enter a Project or Notebook, but primary documents, old drafts and unverified leads should be clearly distinguished.
If this stage is skipped, AI fills gaps with general knowledge or web material while drafting. The article looks well informed, yet the writer can no longer say where particular claims came from. Verification then has to reconstruct sources from sentences, which is far more expensive than establishing the boundary before writing.
An argument needs structure, but not a standardised template
Once the material is ready, state one question the article will answer and one provisional judgement. Each part should then perform a different argumentative task: establishing the factual scope, confronting the strongest counterexample, explaining an institutional structure or returning an abstraction to its practical consequence. Headings should arise from those tasks.
AI templates are recognisable for reasons deeper than a few repeated phrases. A system tends to adjust thought into a linguistically symmetrical shape. Four sections have similar length, every opening contains a story and every conclusion says that “this reminds us” of something. Different articles gradually acquire the shape of one product.
A blacklist of expressions does not solve the underlying problem. Editing must return to conceptual relationships. Some sections need a conclusion before explanation. Another should retain uncertainty. One difficult concept may deserve much more space than its neighbours. If two points belong to the same level, they can be combined. If a turn changes the direction of the argument, it should have room of its own. Sentence length may vary because thought does not arrive at equal intervals.
AI is useful here as a structural opponent. Ask it to identify leaps, repetition and unsupported causation, then to formulate the strongest objection. The author need not accept the unified outline the system proposes. The system supplies pressure; the actual question continues to determine the structure.
Let the draft and a claim register grow together
A dependable long-form workflow maintains a claim register. It need not record every ordinary sentence. It captures the facts and judgements that bear the argument: the proposed claim, its source, the writer’s interpretation, its scope and whether it has been checked.
This separates items that generated prose commonly merges. A source may explicitly state a fact. A relationship formed by reading several sources is an inference. A claim that an institution ought to change is a normative judgement. AI can state all three in the same confident voice. If the register does not mark their difference, syntax alone will not make it visible to the reader.
Drafting can proceed by section rather than by requesting the entire article at once. Each section receives relevant sources, its structural task and the factual boundary it must not cross. After generation, test whether the section advances the argument before considering style. This takes longer, but it prevents a vague statement early in the article from becoming an assumed fact later.
ChatGPT Projects and Claude Projects both preserve project instructions and source material, making them plausible environments for terminology, readership and continuity across a series. Gemini Notebook’s inline source references are useful when asking what basis a sentence actually has in the collection. None of them removes the need for a claim register. The product maintains context; the register preserves the evidence relationship that the author is prepared to accept.
Fact checking and style editing should be separate passes
Once a draft reads smoothly, it is tempting to edit only the sentences. A better sequence begins with a cold pass that ignores elegance and checks names, dates, numbers, jurisdiction, causal wording, quotations and links. AI can identify statements that appear to require support; a person then opens the primary sources. The system can help find the issues, but it cannot prove its earlier output by reviewing itself.
A second pass looks for omissions and counterexamples. Has the article treated product design as demonstrated effect? Has it generalised from one case? Does it omit evidence that cuts against the conclusion? When a source conflicts with the thesis, a reliable workflow modifies the thesis rather than generating a paragraph that merely “balances” the conflict away.
Style editing begins only then. Remove generic openings, synonymous repetition and conclusions that add no information. Look for paragraphs automatically arranged into the same rhythm, repeated “not this but that” constructions and polished sequences of structurally identical claims. Do not simply replace connective phrases with synonyms. Restore the actual logical relationship between the points.
Removing the signs of AI writing does not mean inserting mistakes or randomly disturbing sentences. The aim is to remove false order created for the appearance of completeness. The article should again follow the writer’s judgement: what can be stated firmly, what remains provisional and which question matters more than the rest.
Bilingual writing cannot treat translation as the last mechanical step
A Chinese article passed through AI translation can easily become a grammatical English summary. Qualifications disappear, institutions receive informal names, and a cautious judgement in Chinese becomes too certain in English. The reverse also occurs. Similar length is not evidence that the argument remains equivalent.
A more reliable method stabilises the factual scope and structure of one language, then writes an independent version in the other. Institutions, laws and measures use official names. Central numbers, dates and sources are checked again. The final comparison concerns the task performed by each section and the qualifications attached to the conclusion, not sentence-by-sentence correspondence.
Each language should also sound natural in its own way. Chinese may sustain some relationships through context where English benefits from a more explicit subject. A sequence of short English paragraphs should not be reproduced as mechanically regular Chinese sentences. Bilingual equivalence means that both versions accept responsibility for the same judgement. It does not require identical syntax.
Why the final article cannot exist only in a chat history
Conversation is good for developing thought but a poor final source of publication truth. Threads branch, files are updated and system memory may change. The final article should enter a stable Markdown or other formal file with title, slug, source status, checking date and publication status. Significant revisions should be preserved through version control or at least an explicit file history.
A local file environment also makes it possible to run mechanical checks without asking AI to rewrite the prose. Scripts can verify bilingual markers, link format, minimum length and consistency between headings and metadata. An agent such as Codex can operate over a directory, run those checks and produce a diff for review. Its ability to manipulate files should not imply authority to publish them. The more capable the file operations, the more clearly publication approval needs to remain separate.
A sensible division of work follows. Conversation helps with thinking and drafting. A source-grounded environment supports close checking. A local project preserves, validates and prepares for publication. One product may cover more than one stage, but the responsibilities should remain distinguishable. If the same system writes, declares itself correct and publishes immediately, an error can cross the entire chain without a meaningful pause.
Comparing two real writing environments
Prepare the same screened source collection and choose a subject whose result can be checked manually within a week. Do not provide only one prompt. Give each environment the same intended reader, factual boundary, central question and delivery requirements, while permitting use of its native Project, Notebook or research workflow.
Afterwards, ask five questions. Which important sources did the article omit? How many load-bearing claims lead quickly to primary material? How much human time was needed for correction? Did continuing work require the background to be explained repeatedly? Could the final file enter your storage and publishing process without being reconstructed?
Style belongs in the comparison, but it should not be read first. Begin with the paragraphs most likely to contain error and see whether the conclusion exceeds the evidence. Only then examine whether the prose has a thought rhythm of its own. Two runs are usually more informative than one, because a working environment must repeatedly produce acceptable material rather than occasionally deliver a surprising success.
Conclusion: reliability lies in whether the article can be questioned
When a writing project begins with a fixed source collection and traceability is the first priority, a source-grounded environment such as Gemini Notebook can be a useful starting point. When the project contains ongoing discussion, research, writing and varied tools, ChatGPT Projects or Claude Projects is a more natural home. If many local article files, bilingual markers and publishing checks have to be managed, a Codex-style file environment belongs to another stage of the process.
No combination automatically produces a publishable article. A reliable workflow makes every important judgement open to questions. Where is the evidence? Why does this inference follow? How was the counterexample treated? Do the Chinese and English versions carry the same conclusion? If an error is discovered tomorrow, is it clear which parts must change?
Once AI makes drafts inexpensive, more of an article’s value gathers around those questions. Fluent prose shows that language generation succeeded. Publication under an author’s name still requires sources, judgement, editing and responsibility to arrive together.
Primary sources
- OpenAI: Projects in ChatGPT
- OpenAI: Deep research in ChatGPT
- Anthropic: What are projects?
- Google: Learn about Gemini Notebook
- Official OpenAI documentation: Custom instructions with AGENTS.md
Continue reading: Choosing the Right AI Tool
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.