Build Practical AI Systems
How to Keep a Useful AI Marketing Experiment Log
Keep a useful AI marketing experiment log with baseline, inputs, review time, result, limitations and a clear next decision.
Keep a trial record that another person can understand later. Write down the baseline, inputs, tool and version if known, active time including checks, result, limitations and next decision. If baseline data is missing, write unknown rather than filling the gap. One trial can teach you about a workflow, but it cannot prove general efficiency or business value.
A test case is not a trial record
A test case is a prepared example you can rerun. A trial record says what actually happened on a particular date, with a particular input, process and review.
| Record | Main purpose |
|---|---|
| Test case | Preserve a known input, expected handling and review criterion |
| Trial log | Capture the conditions, observations, limitations and next decision from a run |
B13 covers preserved cases in the cycle package. B16 is the notebook for the trial itself, without adding a public link to the held-out article.
Start with the decision
Write what the trial should help you decide. Examples include:
- Should we repeat this workflow with a timed baseline?
- Which step needs a clearer instruction or source?
- Did the quality check pass often enough to justify another small trial?
- Should we stop because the process creates too much rework or risk?
A decision gives the log a purpose. Without it, the record can become a list of timings that nobody knows how to interpret.
Record the task and baseline
Name the task, starting input, intended output and review criterion. Then state what you are comparing the trial with. The baseline may be a timed manual run, the current process or a prior version of the same workflow.
If the baseline was not measured, record unknown. Do not treat an assumption as zero minutes or claim a percentage change from a missing comparison.
Capture the inputs and version
Record the versions or identifiers that could change the result:
- brief or source-note version;
- instruction, template or context-file version;
- tool and model name and version if visible;
- date and person running the trial; and
- any material change from the baseline.
You do not need to preserve private content in the log. Use an authorised reference, a redacted description or a controlled location. The point is to make the conditions traceable without copying unnecessary sensitive material.
Count active time, including checks
If you measure time, define the start and end. Include preparing the input, the AI interaction, checking sources, correcting the output, asking a question and recording the decision.
Tool response time is only one part of the task. A draft that arrives quickly but needs extensive factual repair may take longer overall. Note interruptions and unusual cases rather than smoothing them out.
Record the result and quality decision
Describe what the trial produced and how it was judged. Use a defined quality criterion, such as “all required outline sections present and every factual point traceable to the approved notes”. Record pass, fail or needs review with a short reason.
Do not write “worked” without saying what worked. A result can be structurally complete but unsuitable because it contains an unsupported claim or misses the reader question.
A filled illustrative trial log
ILLUSTRATIVE EXAMPLE: This entry is fictional teaching material. No model was run and no customer, client, employer or efficiency result is claimed.
| Field | Trial record |
|---|---|
| Decision | Decide whether to repeat a brief-to-outline workflow with better baseline data |
| Date and runner | 9 September 2026; fictional marketer |
| Task | Produce a reviewable outline from one approved B2B article brief |
| Baseline | Unknown. The prior manual run was not timed. |
| Inputs | Brief v0.2; source notes v1.1; tone reference v0.3 |
| Tool and version | Fictional AI tool; version not recorded |
| Active time including checks | 31 minutes, including source check and two wording revisions |
| Result | Outline completed; one unsupported phrase removed |
| Quality criterion | Required reader question, structure and source traceability |
| Quality decision | Needs review. Structure passed; source trail for one point was incomplete. |
| Limitations | One case, unknown baseline, no business outcome measure and no comparison group |
| Next decision | Repeat with a timed manual baseline and a second case, or stop if the source gap remains |
The useful fact here is not “31 minutes”. It is the combination of conditions, quality decision and limitations. A later reader can see what to repeat and what remains unknown.
Separate observation from explanation
Write what you saw before writing why you think it happened.
| Observation | Possible explanation | What to check next |
|---|---|---|
| Two wording revisions were needed | The instruction may not have stated the claim boundary clearly | Test a scoped instruction on another case |
| Baseline is unknown | The old run was not timed | Time the same manual task before comparing |
| One source point remained unresolved | The source notes may be incomplete | Ask the source owner or remove the point |
The possible explanation is a hypothesis, not a finding. Keep it labelled until the next trial or review supports it.
Record limitations without apologising for them
Limitations make a log useful. State the sample size, missing baseline, unusual input, reviewer change, quality-criterion change, tool-version uncertainty and anything else that affects comparison.
If the trial is intentionally small, say so. A small record can guide a next step without pretending to be a benchmark. How to Test and Improve an AI Marketing Workflow describes a small test and one visible change; B16 adds the conditions and limitations record.
Choose one next decision
End with one action, not a vague conclusion.
| Next decision | When it fits |
|---|---|
| Repeat | The workflow is promising but the baseline or case set is too small |
| Revise | A specific instruction, source or review step needs changing |
| Keep for this scope | The quality and effort are acceptable for the defined task |
| Stop | Risk, rework, missing source or unclear ownership remains material |
The next decision should name the owner and the condition for reopening the trial. That turns the log into a learning record rather than a post-hoc success story.
Protect the record and its sources
Keep personal, customer, employer-confidential and credential material out unless the authorised process requires it. Record a reference to the approved source instead of copying sensitive text into a shared log.
Limit access to the people who need the trial record. If the tool or workflow changes, start a new version or make the change visible so a later reader does not combine incomparable runs.
Your next step: complete one log entry
Choose one low-risk AI-supported task. Record its decision, baseline, inputs, tool and version if known, active time including checks, result, quality decision, limitations and next action.
Write unknown for missing baseline data. Then ask another person whether they can understand what happened and what should happen next without asking you to reconstruct the trial from memory.
Further Reading
You Might Still Be Wondering...
Frequently asked questions
Record the decision, task, baseline, inputs, tool and version if known, active time including checks, result, quality criterion, limitations and next decision. Add rework, stop decisions and source references when they affect interpretation.
Mark it unknown. You can still record the trial and use it to improve the next run, but you should not claim a percentage change or general efficiency result without a meaningful comparison.
Yes, when the question concerns total task effort or whether the workflow helps. Include preparation, checking, corrections, escalation and decision recording, not only the AI tool's response time.
It may justify a cautious next step, especially when the task is low risk, but one trial does not establish reliability or business value. Record limitations and decide whether to repeat with a stronger baseline or more cases.
A trial log records one defined run and its context. An evaluation uses stated criteria and a method to assess performance across selected cases. A log can support later evaluation but is not automatically a formal evaluation.
Record what failed, why it matters, the owner and the condition for another attempt. Revise the instruction or source, repeat with a comparable case or stop if the risk and uncertainty remain material.