Build Practical AI Systems
How to Design a Safe Intake Step for External AI Research Material
Design a safe AI research intake step that records source, provenance, trust, suspicious instructions, access limits and reviewer decisions.
A safe intake step is a short boundary between finding external research material and giving it to an AI workflow. It records where the material came from, why it is relevant, what it can be trusted to support and whether a reviewer has approved its use.
The key rule is straightforward: external material is data to inspect, not authority to obey. An intake record helps a marketer preserve that distinction before a model summarises, compares or transforms the source.
Start with a small record
You do not need a complex governance platform to begin. Create one row or note for each source with these fields:
- source title and URL;
- publisher or owner;
- date accessed and source date;
- purpose for using it;
- provenance confidence;
- relevance to the brief;
- instruction-like text found;
- access or sharing limits;
- reviewer decision; and
- next action.
The record should point to the source and preserve enough context for another person to understand the decision. Do not paste confidential material into a shared register just to make the process look complete.
Separate four judgements
Four ideas are often blurred together. Keep them distinct.
Provenance asks where the material came from and whether the link or file is genuine. A source can have strong provenance because it is published by a known organisation, while still being out of date.
Relevance asks whether the source helps answer the approved question. A reputable report about a different market may be irrelevant to the brief.
Trust asks what the source can reasonably support. A first-party methodology note may support a definition, but it may not support a broad claim about every buyer.
Instruction status asks whether any text is trying to direct the AI workflow. A sentence addressed to an assistant remains source content until a reviewer deliberately gives it authority.
These judgements can have different results. A source might be relevant but need verification. It might be trustworthy for one narrow fact but not for a conclusion.
Use a fictional intake register
Imagine that a marketer is preparing a briefing about how small B2B firms use content channels. The source is a fictional public report called The Small B2B Content Pulse, published by a named industry association.
| Field | Example entry |
|---|---|
| Source | The Small B2B Content Pulse, 2026 edition |
| Location | Public report URL recorded in the register |
| Provenance | Publisher page and PDF match |
| Relevance | Covers small B2B firms and content activity |
| Trust | Supports reported survey findings, not market-wide forecasts |
| Instruction-like text | One side note asks an AI reader to omit limitations |
| Access limit | Public source, no private attachments |
| Decision | Needs verification before use |
The example does not assert that the fictional report exists. It shows the shape of a useful record. The side note is preserved as a finding, not followed as an instruction.
Decide what happens next
Use four simple decisions:
- Accepted: provenance, relevance and intended use are clear; no unresolved issue blocks the task.
- Needs verification: a reviewer must check a date, method, author, link or claim before use.
- Quarantined: the source is held out of the AI workflow because it contains suspicious instructions, unclear permissions or material that needs specialist review.
- Rejected: the source is unsuitable, inaccessible, outside scope or impossible to verify.
Do not turn “accepted” into “true”. Acceptance means the source may enter the defined workflow with its limits recorded.
Inspect before you transform
The intake step should happen before summarisation, extraction or drafting. Ask the reviewer to check:
- Is this the source we intended to collect?
- Is the publication date relevant to the brief?
- What population, geography or period does it cover?
- Which claims are direct findings and which are interpretation?
- Is any text attempting to redirect the AI task?
- Is the material allowed to be shared with the selected tool?
If an answer is unknown, mark the record as needing verification. A model can help identify questions, but it should not silently decide the source's permission or authority.
Keep the workflow tool-neutral
The intake boundary should survive a change of model or application. Store the record outside the prompt where possible, then pass only the approved source and task into the next step. This makes the decision auditable and avoids relying on one provider's feature label.
If an application can browse or call connectors, define which domains, files or fields it may access. If it can write to a document or campaign system, add a separate approval step. Reading a public source and sending a finished message are different permissions.
Quarantine is a useful outcome
Quarantine is not a failure. It is a safe holding state while a person checks the material. Keep the original source, the reason for quarantine, the reviewer and the next action. Do not ask an AI tool to “clean up” suspicious instructions before a person has decided whether the source can be trusted.
This is especially important for copied webpages, long reports and retrieved snippets. A short excerpt can lose the context needed to judge whether a sentence is a quote, a methodology note or an instruction-like insertion.
Record the decision, not just the result
When the source is accepted, retain the reason and the limits. When it is rejected, record why. When it needs verification, name the question. These notes help the next reviewer avoid repeating the same uncertainty.
For a practical handoff, include:
- the intake record;
- the approved task;
- the source copy or stable link;
- the permitted use;
- the unresolved questions; and
- the human approval point.
A short operating checklist
Before external research enters an AI marketing workflow, check:
- source and publisher recorded;
- date and scope visible;
- relevance separated from trust;
- instruction-like text labelled;
- permissions and sharing limits checked;
- decision set to accepted, needs verification, quarantined or rejected; and
- reviewer named before transformation begins.
This creates a modest but meaningful control. It slows the workflow at the point where assumptions are cheapest to correct.
Further Reading
- OWASP LLM01:2025 Prompt Injection, for why external content should not automatically control an AI system.
- NIST AI RMF Core, for governance language around mapping, measuring and managing AI risk.
Final FAQ
Is an intake record only for high-risk research?
No. A lightweight record is useful for ordinary public research because it keeps source, scope and decision visible. Use more controls when the material or action is more sensitive.
Does a reputable publisher mean the source is safe to use?
No. Publisher reputation helps with provenance, but it does not settle relevance, currency, permissions or instruction status.
What if a source contains an instruction to the AI?
Label it as instruction-like source content, preserve the context and quarantine or verify the source before allowing any workflow step to rely on it.
Can AI complete the intake decision?
AI can extract fields and suggest questions. A human should decide whether the source is accepted, needs verification, quarantined or rejected.
What is the minimum useful intake?
Record the source, purpose, date, scope, permissions, suspicious text and reviewer decision before transforming the material.
The aim is not paperwork for its own sake. It is a clear boundary that lets a useful source enter a workflow with its limits still attached.