Can ChatGPT, Claude and Gemini Really Be Compared Directly?

Choosing the Right AI Tool · Article One

Suppose we give the same question to ChatGPT, Claude and Gemini, then place their three answers side by side. We have certainly performed a comparison. The harder question is what, exactly, the comparison has measured.

It may have measured three passages generated on one day, under three particular account tiers and sets of default settings. A different model could change the result. Turning on search or deep research could change it again. Moving the task into a Project or a Notebook would alter the information available to the system. Subscription plans affect both capabilities and usage limits. Even if every visible condition remained unchanged, another run would not necessarily produce the same answer.

Such a comparison is not useless. It can show how three configurations performed on one task. It cannot comfortably support a much larger claim such as “which AI is best”. That question sounds as though it concerns three fixed objects. In practice, it concerns three expanding product systems.

How many things are concealed by one name?

Everyday discussion often merges a company, its models and its consumer product. ChatGPT is a user-facing product that offers different models and tools. Claude can refer to Anthropic’s model family, its chat product and working environments such as Claude Code. Gemini likewise covers models, a conversational application, AI features within Google services and specialised research tools.

At least five layers need to be separated:

Layer The question to ask Why it changes the result
Model Which model and reasoning setting were actually used? Capability, speed and usage cost differ
Mode Was this ordinary chat, search, deep research or an agent task? The system obtains information and acts in different ways
Workspace Was it a single chat, a Project or a Notebook? Files, instructions and history may or may not persist
Plan Was the account free, individually paid or organisational? Limits, models, connectors and controls differ
Entry point Was the work done on the web, mobile, desktop, in an IDE or at a terminal? The system can see and change different parts of the working environment

This is not pedantry. A person trying to organise twenty interview transcripts needs to know whether the files remain attached to a continuing project, how the system retrieves them, whether a citation leads back to the source and whether future conversations inherit the project’s instructions. A single general-knowledge question reveals almost none of those differences.

If the task is merely to make an email more concise, elaborate research capacity may add little. Startup friction, response speed and the tool’s ordinary location may matter more. A supposedly superior capability acquires practical meaning only after it enters a particular kind of work.

The same prompt is not always a fair test

A standard prompt looks like the fairest comparison. It controls the input words, but it can also suppress the way a product was designed to work.

A deep-research tool will ordinarily plan, search multiple sources and produce a cited report. Standard chat is better suited to rapid clarification. Requiring both to complete the task in one response under a single short prompt creates equal-looking conditions that favour one task structure. Gemini Notebook is designed to answer in relation to sources selected by the user. Asking it an open question without providing material does not test its central design.

Fairness does not require every tool to follow an identical path. A better method fixes the intended outcome and evidentiary requirements while allowing each product to use its native workflow. Imagine preparing a brief on the Australian market for adult piano learning. The brief could be required to distinguish official statistics from industry estimates, give an accessible source for each major claim, identify unresolved questions and end in a document that can still be edited. One product may use an editable research plan, another a web research process, and another a notebook. The comparison has moved from “which one wrote the best opening paragraph?” to “which one completed the whole assignment more reliably?”

A good answer is not identical to a good tool

The quality of a response matters, but a long-term tool choice includes costs that are often left outside the score.

There is preparation cost: how much material must be uploaded, how often must the background be explained again, and how much instruction has to be maintained? There is verification cost. Can citations be opened directly? Can a number in a table be traced to its original source? Has the system presented an inference as an established fact? Then there is the cost of continuing. Does the result enter a document, repository or team process cleanly, or must it be copied out of the chat and reconstructed elsewhere?

Recovery from error matters even more when consequences increase. A mistaken line of promotional copy is different from an unwanted change across production code. The second case calls for a visible diff, testing, constrained permissions and a reliable route back. A score based only on the final text removes the very conditions that determine whether an AI tool belongs in real work.

A product can therefore produce a less dazzling response on one occasion and still be the more dependable daily environment because its sources are visible, context remains stable and changes are easy to inspect. Another may occasionally return a remarkable draft, yet be a poor fit for a high-risk task if the process cannot be reconstructed. Output quality is one component of tool quality; it is not the whole.

Human time belongs in the comparison

AI reviews frequently record the number of seconds before a response appears. They rarely count the user’s work afterwards. If a research report is generated in four minutes but its citations require two hours of repair, the research was not completed in four minutes. If a coding agent runs for half an hour and a developer then spends half a day understanding an unnecessarily broad change, agent runtime alone is a misleading measure of efficiency.

Human time appears before, during and after generation. Material has to be prepared. Some decisions must be supplied while the tool is working. The result then needs checking, correction and integration. A useful product does not necessarily remove the person from this sequence. It puts human attention where judgement is valuable and reduces needless transport and repetition.

Familiarity with an ecosystem can consequently change the result. For someone whose documents already live in Google Drive, connections between Gemini and Google services may remove genuine handling work. A development team centred on GitHub may place more value on issues, pull requests, permissions and review occurring in one environment. A reviewer with different material and habits will assign those features different value.

Many claims of objective ranking simply hide these background conditions. Stating them does not make comparison arbitrary. It makes the conclusion usable by someone else.

A practical personal comparison

Most users do not need to construct a large benchmark. Choose one genuine task that recurs every month and whose result can be checked. It might be a summary drawn from a fixed source pack, a short piece of research, a data-cleaning job or a small code change covered by tests.

Before beginning, define an acceptable result. Must it cite nominated sources? May it introduce material from the web? What final format is required? Which errors are unacceptable? Record the plan, mode and date. Run the task at least twice in each environment, because a single success or failure may be accidental.

Do not begin by scoring style. Check factual accuracy and omissions, then count the human time needed to reach a usable result. The ease of continuing the work deserves separate attention. Finally ask a very ordinary question: if this task returns next week, in which environment would I prefer to resume it?

The method will not produce a dramatic “number one AI” for social media. It is more likely to produce an answer that fits an actual life.

How the three products should be understood

At the time of checking, ChatGPT Projects brings chats, files, project instructions, memory and a range of tools into a continuing workspace. Claude Projects centres on self-contained project histories, knowledge bases and project instructions. Google has renamed NotebookLM as Gemini Notebook, continuing its source-centred research environment while connecting it more closely to the Gemini application.

The products are absorbing capabilities from one another, but they have not become the same thing. ChatGPT extends towards a multi-tool workspace and agent action. Claude connects chat, Projects, Research and environments such as Claude Code. Gemini’s product structure is closely related to Google Search, Drive, Gmail, Docs and Notebook. These are descriptions of product form, not declarations that any one of them wins a particular task.

The comparisons worth making are narrower. Which environment gives a user appropriate control over sources during deep research? How do Projects and a Notebook treat a collection of personal evidence? During long-form writing or code modification, which workflow is easiest to inspect? The remaining articles in this series address those questions separately.

Conclusion: fix the task before judging the answer

ChatGPT, Claude and Gemini can be compared. What cannot safely be done is to expand the outcome of one chat into a verdict on three entire products. A useful comparison fixes the final task, source boundary, tolerance for error and delivery requirements. It also records the model, mode, plan and entry point that were actually used.

The conclusion should not be a ranking presented as permanent. It should answer a conditional question: given where my material lives, how I work and what an error would cost, which tool reaches a result that is easier to verify, continue and correct with a reasonable amount of human handling?

Once the question is framed this way, choosing a tool becomes less dependent on brand loyalty or a single impressive answer. Products will continue to change quickly. The judgement now has a more stable point of support.

Primary sources

Continue reading: Choosing the Right AI Tool


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.