Understand AI Marketing

What Should You Measure When Testing AI in Marketing?

Choose a baseline, task measure and quality check for an AI marketing workflow without confusing time saved with business impact.

11 September 2026By Michael Sweenie7 min read

Measure the thing you are deciding about. For a small AI marketing workflow, start with one task measure, one quality check and a baseline. Include human review time, record rework and failures, and keep any downstream business outcome separate. A faster draft is not automatically a better result, and a change in a metric does not by itself prove that AI caused it.

Four layers answer different questions

LayerMain questionExample
Task effortDid the work take less or more effort?Minutes from brief to reviewed outline, including review
Output qualityDid the result meet the agreed criterion?A checklist for source accuracy and required sections
Rework and failureWhat needed correction, stopping or escalation?Revision count, missing-source flags or stop decisions
Business outcomeDid a downstream indicator change?An approved next action or another pre-agreed measure

These layers are related but not interchangeable. Time can change while quality falls. Rework can rise because a reviewer is more thorough. A business indicator can change for reasons unrelated to the workflow. Name the question before choosing the metric.

Start with a baseline

A baseline gives the proposed AI-supported process something to be compared with. It might be a small set of recent manual runs, the current workflow or a clearly defined before period.

Write down:

  • the task and starting input;
  • who reviews the output and what “ready” means;
  • the dates or cases included;
  • the tools and steps used; and
  • what changed in the proposed version.

If the baseline is vague, the comparison will be vague too. Three cases can help you learn about a small workflow, but they do not establish a universal rate or guarantee.

Measure task effort, including review

If your question is “does this help with the work?”, measure the whole task. Include preparing the input, waiting for the output, checking sources, correcting the draft, escalating questions and recording the decision.

Tool response time alone may look impressive while a person spends longer repairing the result. Use a consistent start and end point, such as “approved brief received” to “reviewed outline ready for the next step”. Record interruptions or unusual cases that make a run unlike the baseline.

ILLUSTRATIVE EXAMPLE: The following values are fictional teaching material, not a measured trial.

VersionCasesMean task time including reviewObservation
Current manual workflow342 minutesOne case needed a source check
AI-supported draft plus review335 minutesTwo cases needed wording rework

Even this simple table needs caution. A mean from three fictional cases is not a forecast. It does not explain the difference or tell you whether the quality threshold was met.

Define one quality check

Choose a criterion a reviewer can inspect. Examples include:

  • every important claim has an approved source;
  • the outline includes the required reader question and next action;
  • no private or unsupported material appears; or
  • the final draft keeps the approved meaning and qualification.

Write the criterion before looking at the proposed outputs. Use pass, fail or needs review when that fits the task, and record the reason. A quality check should match the risk. One checklist cannot certify every aspect of an article, campaign or customer-facing decision.

Record rework and failure patterns

Count more than successful outputs. Record corrections, missing inputs, stop decisions, source gaps, privacy flags and the kind of rework required.

RecordWhy it matters
Revision countShows how much work remained after the first output
Rework typeDistinguishes wording, factual, structural and source problems
Stop or escalationShows where the process correctly refused to proceed
Missing-input flagShows whether the workflow had what it needed

Rework is not automatically bad. A careful reviewer may identify issues that the old process missed. The useful comparison is whether the workflow meets the agreed quality and risk standard at a reasonable total effort.

Keep business outcomes separate

A business outcome may matter, but it is further away from the AI step. A website interaction, enquiry or sale can be influenced by audience, offer, timing, channel, seasonality, distribution and many other changes.

Measurement tools can record interactions. For example, Google Analytics describes events as measurements of user interactions that feed reports. That can help you define an instrumented observation, but an event count does not show that AI caused it.

If you track a downstream measure, define it before the trial, keep the comparison meaningful and record other material changes. Do not turn “more clicks” into “AI improved marketing” without a suitable design and evidence.

Match the measure to the decision

Ask what you will do with the result:

DecisionUseful first measureAdditional check
Should we keep this workflow?Total task effort including reviewOne defined quality criterion
Where does it fail?Rework and stop categoriesSource or privacy check
Is the output fit for use?Quality pass/fail with reasonsHuman reviewer decision
Did a downstream indicator change?Pre-agreed business measureComparison, context and alternative explanations

Do not collect metrics simply because a tool makes them available. A measure without a decision can create reporting work without improving the workflow.

Document limitations and context

NIST's AI RMF Measure guidance emphasises selecting methods and metrics that fit the risks and context, and documenting what cannot be measured. Apply that modestly here. Record sample size, task definition, reviewer, dates, tool version if relevant, missing data and any change between baseline and proposed runs.

If the quality criterion changed halfway through, say so. If the cases were unusually easy, say so. If a business outcome has too many confounders, leave it as an observation rather than a causal claim.

Avoid common interpretation errors

  • Time saved equals value: Less task time may be useful, but quality, rework and reviewer effort still matter.
  • A pass equals reliability: A few passing cases do not cover every future input.
  • More activity equals better marketing: More clicks, drafts or outputs may not help the intended reader.
  • Correlation equals causation: A downstream change can have several explanations.
  • A dashboard equals evidence: A number needs a definition, baseline, context and decision owner.

How to Test and Improve an AI Marketing Workflow describes a small test and one visible change. U16 adds the measurement question, not a formal benchmark or ROI calculator.

Your next step: choose two measures

Choose one recurring AI-supported task. Write its baseline, start and end points, sample or observation period and reviewer. Then choose:

  1. one task measure, such as total minutes including review; and
  2. one quality check, such as source accuracy or required-section completeness.

Record rework and stop decisions alongside the result. State what the measures cannot show, then decide whether to keep, revise, investigate or stop the workflow. Add a business outcome only when you can define a suitable comparison and context.

Further Reading

You Might Still Be Wondering...

Frequently asked questions

Back to Blogs