Build Practical AI Systems
How to Build a Small Test Set for an AI Marketing Workflow
Build a small, reusable AI marketing workflow test set with normal, incomplete and ambiguous inputs, expected human decisions and a clear rerun record.
Use more than one kind of input when you check an AI marketing workflow. A small test set can include a normal case, an incomplete case and an ambiguous case, each with an expected human decision. Run the same cases again after a visible workflow change and record what happened.
This makes review more consistent. It does not prove that the workflow is reliable for every future task, model or audience. Three cases are a practical learning set, not a benchmark.
What is a small test set?
A small test set is a short collection of inputs used to check a recurring workflow in a consistent way. Each case has a purpose, an expected review decision and a record of what the workflow actually did.
For a marketing workflow, the set might check an article outline, a source summary, a campaign brief or a list of approved message options. The set is about the workflow and its inputs, not a public league table of models.
OpenAI describes evaluations as tests of model outputs against style and content criteria that you specify. NIST also notes that measurement and evaluation methods depend on context. Those points lead to a simple rule: design the test cases around the task you actually want to review.
Start with one workflow and criterion
Choose a recurring task with a clear next step. For example: “Turn approved generic product notes into a short B2B article outline.”
Then write one observable criterion. It might be: “The outline contains the required sections from the approved brief and adds no unsupported product claim.” A person should be able to inspect the output and explain whether the condition is met.
Do not begin by trying to test every quality dimension. A small set is easier to interpret when it has one clear purpose. You can add criteria later if the first record shows a gap.
Add a normal case
The normal case represents the input the workflow is meant to receive. It should contain the expected fields, authorised material and enough context for the task.
For the illustrative outline workflow, the normal case could include a topic, intended reader, approved facts and three required sections. Record the exact input version and the expected decision, such as “run the workflow, then review the outline against the criterion”.
Normal does not mean perfect. It means the case is within the workflow's intended shape. If you quietly repair it before every run, you lose information about what the workflow actually received.
Add an incomplete case
The incomplete case is missing something the workflow needs. The missing item could be the audience, a required source, an approval decision or a product detail.
The expected decision should be explicit. The workflow may need to pause and request the missing material rather than produce a confident outline. This case checks whether the process exposes a gap instead of treating absence as permission to invent.
Keep the missing item visible in the record. Do not fill it with a plausible detail just to make the example run smoothly.
Add an ambiguous case
The ambiguous case contains information that can reasonably be read in more than one way. Two approved notes might use different descriptions for the same feature, or an instruction might allow two possible audiences.
The expected decision is usually to pause for clarification or ask a named reviewer to choose the interpretation. An ambiguous case tests whether the workflow makes uncertainty visible before it becomes polished copy.
Ambiguity is not the same as an error. The case is useful because it identifies a decision the workflow cannot safely make on its own.
Write expected decisions before running
Record the expected action for each case before you run the workflow. This reduces hindsight and gives the reviewer something specific to compare.
| Case | Input shape | Expected decision |
|---|---|---|
| Normal | Required topic, reader, approved facts and sections are present. | Continue to the AI step, then review the output against the criterion. |
| Incomplete | Audience or one required source is missing. | Pause and request the missing material. Do not invent it. |
| Ambiguous | Two notes describe one product detail differently. | Pause for clarification or named human selection before drafting. |
ILLUSTRATIVE EXAMPLE: These cases are fictional teaching material for an outline workflow. They are not a customer record, model result or benchmark.
Keep one run record per case
For each run, keep:
- the case label and input version
- the workflow version or visible change
- the output being checked
- the criterion and expected decision
- the actual decision: pass, fail or uncertain
- the reviewer's reason
- the next action
The output record matters even when the expected decision is “pause”. A refusal to continue can be the correct behaviour when the input is incomplete or ambiguous.
Run the same set after one change
If you change an instruction, context file, review step or tool, rerun the same three cases before adding new ones. Keeping the cases stable makes it easier to see what changed in this observation.
Do not change the input, prompt, model and review method together if you want to learn which change mattered. Record the change, the date and any new uncertainty. A different result is an observation about these conditions, not proof of causation.
For the broader small-test loop, see How to Test and Improve an AI Marketing Workflow. B6 covers the before point, one visible change and the next decision. B13 adds the normal, incomplete and ambiguous input spread.
Know what the set cannot tell you
Three cases do not establish reliability across all inputs, users, models or future versions. They do not certify safety, prove accuracy or predict leads, rankings, conversion or time saved.
As a workflow becomes more varied or higher risk, you may need more representative examples, more criteria, specialist review, monitoring or a formal evaluation programme. NIST's guidance is broader than this starter method and keeps measurement tied to context.
The test set is useful when it makes the next decision clearer. It is not useful when its labels are treated as a guarantee.
Protect privacy and permissions
Use authorised material only. Replace personal, confidential, employer-sensitive or access-restricted input with a fictional or generalised case when the real material is not needed. A test record should document the workflow without becoming a new place to store unnecessary private information.
If a case reveals a permission, source or technical uncertainty, pause and route it to the appropriate person. Do not hide the issue to keep the set tidy.
Your next step: create three cases
Choose one recurring AI marketing workflow and write one criterion. Create a normal input, an incomplete input and an ambiguous input. For each, record the expected human decision before running anything.
Run the same set once, keep the outputs and review notes together, then rerun it after one visible workflow change. Record what changed and what remains uncertain. Treat the set as a learning tool, not proof that the workflow will behave the same way everywhere.
Further reading
You Might Still Be Wondering...
Frequently asked questions
They give a small set more range than an easy, complete example alone. The labels help you check whether the workflow can proceed, pause for missing material or surface uncertainty.
No. Three cases support a small, repeatable learning check. They do not represent every input or prove accuracy, safety, business impact or reliability across future uses.
Only when it is authorised, necessary and appropriate for the tool and task. Otherwise use a fictional or generalised case that preserves the decision you need to test.
Record that decision. Pausing for a missing source, permission or clarification can be the correct workflow behaviour. A test set should make that visible, not reward the workflow for continuing.
Add cases when the workflow, audience, input types or risk changes, or when a failure reveals a condition the current set does not cover. Keep the original cases so you can still compare observations.