Build Practical AI Systems
How to Test and Improve an AI Marketing Workflow
Learn a small way to test an AI marketing workflow: define one criterion, review one task, record uncertainty and change one thing at a time.
To decide whether an AI marketing workflow is worth reusing, choose one thing a useful output must contain, run the workflow on one small task, review the result and change one thing at a time. Record what you observed and what remains uncertain.
That is a learning step, not proof that the workflow caused an improvement. One before-and-after comparison can show what happened in that instance, but it cannot establish a general gain in quality, efficiency, leads, rankings or performance.
What does “worth reusing” mean?
For a starter workflow, “worth reusing” should mean only that it met the criterion you chose on the task you observed, and that no unresolved issue makes another use inappropriate.
It does not mean that the workflow is generally better, reliable in every situation or proven to create a business result. The input, instructions, tool behaviour, reviewer and task conditions can all affect an output.
The goal of a small test is to make the next decision clearer. You might decide to repeat the workflow on another suitable task, change one part, keep it for human-assisted use or stop using it until an uncertainty is resolved.
1. Define what a useful output must contain
Start with one criterion that a person can inspect. For example, a useful output might:
- include the required sections from an approved brief
- answer the stated reader question
- keep important claims tied to supplied or approved sources
- follow an agreed format
These are examples of possible criteria, not a universal scorecard. Choose one for the first test. If you try to judge everything at once, it becomes harder to see what the workflow is actually being tested for.
Write the criterion before running the task. Phrase it so that another person could understand what would count as present, missing or uncertain.
2. Record the before point
Decide what “before” means in your small comparison. It could be the result from your current manual or AI-assisted approach, recorded before one change is made. It could also be an initial version of the same workflow before you adjust it.
Record enough context to make the comparison understandable:
- what task was attempted
- what input was authorised
- what the workflow was asked to produce
- which one criterion you chose
- what you noticed before changing anything
Keep the comparison fair enough to interpret, but do not describe it as a controlled experiment unless you have designed and run one. A different task, input or reviewer can change the result.
3. Run one small task
Choose a task that is narrow enough to review. An illustrative example could be turning a short set of approved generic notes into a draft outline. The task should not require private customer information, confidential employer material, credentials or access to a live publishing system.
Use the workflow as it currently exists and record the output. Do not quietly change several instructions, inputs and review steps during the same run. If the workflow includes an AI-supported action, keep a person responsible for deciding whether the output can be used.
The purpose of this run is to learn about one use of the workflow. It is not to prove that the workflow will behave the same way on every future task.
4. Review the result against the criterion
A person should inspect the output and mark the criterion as met, not met or uncertain. Then record the reason in plain language.
The reviewer should also notice issues that could make reuse inappropriate, including:
- an important fact or technical detail that needs checking
- an output that does not serve the task's purpose or reader
- private or unnecessary information in the input or output
- an unsupported claim or missing source
- a step that the workflow did not have permission to take
This is a focused test, not a replacement for human review. Passing one criterion does not show that every other aspect of the output is correct, relevant, appropriate or safe to use.
5. Change one thing at a time
If the result is weak or uncertain, choose one change for the next run. You might narrow the input, clarify the output requirement, add an approved source, change the order of one workflow step or strengthen a review question.
Make the change visible in your record. Avoid changing the prompt, context, model, input and review method together, because you may not know which change affected the next output.
Changing one thing at a time is a practical learning recommendation. It does not guarantee that the next version will be better, and it does not remove the need for human judgement.
6. Decide what to do next
Use the record to choose one proportionate next step:
- Repeat carefully: the criterion was met, no material issue was found and the next task has a similar low-risk shape.
- Change one thing: the criterion was missed or the reason is clear enough to address.
- Pause: an important fact, source, privacy issue, permission or technical detail remains uncertain.
- Stop: the workflow is outside its intended purpose or cannot be reviewed responsibly.
These are decisions about the next use of the observed workflow. They are not claims that the workflow has been validated for every task.
A small illustrative test record
ILLUSTRATIVE: The following is a fictional blank record for a workflow that turns approved generic notes into a draft outline. It contains no test result or performance claim.
| Field | Record |
|---|---|
| Task | Turn approved generic notes into a short draft outline. |
| One success criterion | The outline includes the required sections from the approved brief. |
| Before point | Record the initial output and whether the criterion was met, not met or uncertain. |
| After point | Record the output after one visible workflow change. Do not infer causation from the comparison. |
| Human review | Check the criterion, purpose, important claims, sources, relevance and privacy. |
| Uncertainty | Record anything that remains unknown, unsupported, inconsistent or outside scope. |
| One change | Record the single workflow change made between the two observations. |
| Next decision | Repeat carefully, change one thing, pause or stop, with the reason. |
The blank fields are deliberate. Filling them with invented results would turn an illustration into an unsupported case study.
Keep privacy and uncertainty visible
Use the minimum information needed for the test and confirm that it is authorised for the AI tool and task. Remove or generalise personal, confidential, employer-sensitive, private customer or access-restricted material unless the relevant process allows its use.
If the output includes an uncertain claim, an invented detail, a source that has not been checked or a decision outside the workflow's purpose, mark it clearly and pause. A test record does not make private information safe, and a review note does not prove that every error has been found.
Your next step: run one small comparison
Choose one recurring, low-risk task and write down one success criterion before you run it. Record the current or initial output as the before point. Then make one clearly described workflow change, run the task again and record the after point.
Ask a person to review both outputs against the same criterion. Write down what changed, what did not change and what remains uncertain. If the record does not support a responsible next use, pause rather than treating the workflow as ready to repeat.
Further reading
- How to Map a Simple AI Marketing Workflow
- How to Build a Simple AI Skill for One Marketing Task
- How to Plan a Safe First AI Agent for a Marketing Task
- What Does Human Review Mean in AI Marketing?
- OpenAI: Working with evals
- Google Search Central: Guidance on using generative AI content on your website
- UK Government Data and AI Ethics Framework
You Might Still Be Wondering...
Frequently asked questions
Test one visible criterion for one small, low-risk task. For example, check whether the output includes the required sections from an approved brief. Do not try to judge every possible quality dimension in the first test.
It is a simple record of an initial or current output and an output after one visible workflow change. It helps you describe what you observed. One comparison does not prove that the change caused the difference.
No. It means only that the criterion was met in the task you observed. You still need to consider purpose, important facts, sources, relevance, privacy, permissions and any uncertainty before another use.
Changing one thing makes the next observation easier to interpret. It does not guarantee improvement, but changing many parts at once makes it harder to know what affected the result.
Record the uncertainty and pause if it could affect accuracy, privacy, permissions, purpose or responsible use. Seek the appropriate human decision before repeating the workflow.
No. It is a small starter method for learning from one workflow task. A larger or higher-risk system may need more representative data, more criteria, repeat runs, specialist review and formal evaluation.