Build Practical AI Systems
How to Compare Two AI Tools on the Same Marketing Task
Compare two AI tools fairly on the same marketing task using matched inputs, evidence, criteria, versions and a recorded review.
To compare two AI tools fairly, give them the same representative task, input, evidence, output shape and review criteria. Record the versions and conditions, then describe the differences you observed without claiming a universal winner.
The aim is a decision you can explain for one workflow. It is not a product ranking or a promise that the same result will appear in every task.
Start with one task
Choose a task that occurs often enough to matter and is safe to test. For example, ask both tools to turn the same fictional product notes into a 400-word briefing for a named B2B reader.
Define the required output before testing:
- audience and purpose;
- source material;
- length and structure;
- evidence boundary;
- tone and accessibility needs; and
- what the tool must not invent.
If the task changes between tools, you are comparing tasks as well as tools.
Fix the input and conditions
Keep the source text, prompt, attachments and output format the same. Record whether each tool has browsing, memory, retrieval or other capabilities enabled. Do not give one system extra context and then call the result a fair comparison.
Also record the date, model or application version and any relevant settings. A later rerun may produce a different answer even if the product name is unchanged.
Use a fictional protocol
| Field | Protocol example |
|---|---|
| Task | Draft a brief from one fictional product note |
| Input | Same 700-word source for both tools |
| Output | 400 words, three headings and a source note |
| Criteria | Source fidelity, reader fit, structure, unsupported claims |
| Conditions | Same prompt, no browsing, same review window |
| Record | Version, date, output and reviewer notes |
| Stop rule | Pause if either tool exposes sensitive data or invents evidence |
The protocol is fictional and tool-neutral. It is a starting record, not a claim about any named product.
Agree the criteria before reading outputs
Write a simple rubric before you see the drafts. For each criterion, define what acceptable work looks like.
For example:
- Source fidelity: important facts remain aligned with the supplied note.
- Reader fit: language and next action suit the defined audience.
- Structure: the requested sections are present and usable.
- Evidence: unsupported claims are absent or clearly marked.
- Review effort: a human can check the result without excessive repair.
Use descriptive notes rather than pretending that a small test produces a precise universal score.
Keep observations separate from explanations
An observation might be “Tool A retained the source date; Tool B omitted it.” An interpretation might be “The prompt or output shape may have made date retention easier for one tool.” The second statement is a hypothesis, not a proven cause.
Record both, but label them. If the result matters, repeat the test or change one condition at a time.
Watch for cherry-picking
Do not choose the most flattering output from one tool and the weakest output from the other. Save the full run and review the same sections. If one tool produces several candidates, define in advance how the candidate will be selected.
Avoid changing the prompt after seeing one result unless you start a new, clearly labelled test. Otherwise the comparison becomes a sequence of adjustments that cannot be explained later.
Include a human reviewer
The reviewer should understand the task and the criteria, not just prefer one writing style. Ask the reviewer to identify:
- supported observations;
- missing context;
- invented or overstated claims;
- accessibility or clarity issues; and
- changes needed before use.
If possible, have a second reviewer look at disagreements. A human review is part of the workflow, not proof that one tool is objectively better.
Set a stop rule
Decide what will pause the test. Examples include exposure of private material, an unsafe instruction, a tool action outside scope or a result that cannot be checked against the source.
Stopping protects the task and keeps the comparison honest. Do not continue just to complete a spreadsheet of scores.
State the limitations
Your comparison may be limited by a small sample, one prompt, one date, one reviewer or fictional input. State those limits next to the observations.
Do not use the result to claim that one tool is best for every marketer, every model version or every task. A local test supports a local decision.
Decide what to do next
The outcome might be:
- continue with one tool for this task;
- run a larger test;
- redesign the prompt or source format;
- add a human review step; or
- stop using both tools for this workflow.
A tie or an inconclusive result is a valid result. It tells you that the evidence is not strong enough for a broader decision.
Save a comparison note
Keep the task, inputs, prompts, versions, outputs, criteria, reviewer notes and decision together. If a model changes, rerun the same protocol or state clearly why the comparison is no longer current.
This is the practical value of a fair comparison: a colleague can see how the decision was made and challenge the assumptions.
Further Reading
- OpenAI Evals API documentation, as a provider-specific evaluation example.
- NIST AI RMF Core, for a broad frame for measuring and managing AI risk.
Final FAQ
Is the best tool the one with the best writing?
Not necessarily. Your task may prioritise source fidelity, permissions, accessibility, review effort or another criterion.
Can I change the prompt between tools?
Only if the comparison is explicitly about prompt design. For a tool comparison, keep the task and prompt matched.
Should I score every output numerically?
Not always. Descriptive observations can be more honest for a small, qualitative test.
What if the outputs are both poor?
Record the result and inspect the brief, input and criteria. The task may need redesign rather than a different tool.
What makes the result reusable?
Keep the protocol, versions, evidence, reviewer notes and limitations together so the test can be repeated or challenged.
A fair comparison does not promise certainty. It gives you a traceable way to learn what two tools did on one task under one set of conditions.