Content Agents: From Brief to Draft to Review Without Publishing Filler
What an agent needs in a brief to produce usable content, why context files beat prompts, who reviews, and how to measure quality with an eval set.
in this article
- 01The brief is the product, not the prompt
- 02Context files beat clever prompts, because prompts do not survive
- 03Agents draft, outline, vary and repurpose. They do not decide
- 04Review is editing, and it needs the person who holds the claim
- 05Measure with an eval set, not with vibes
- 06The failure mode is volume without a point of view
- 07Frequently asked questions
A team gives a model a headline and a tone instruction and receives eleven hundred words back. The grammar is clean, the structure is tidy, the topic is correct, and the piece says nothing. The reflex is to blame the model and go hunting for a better prompt.
The cause is upstream. The agent was asked to produce an argument nobody had made yet, so it did what a model does with no position to defend: it wrote the average of everything already published on the topic. The agent is not a writer you hired. It is a fast, literal production layer that converts a decision into a draft, and where the decision is missing it invents one, and the invented one is always the consensus.
The brief is the product, not the prompt
Almost all of the difference between a usable draft and filler is decided before the agent runs. A brief that produces something publishable contains six things, and the sixth is the one everyone skips.
The audience and their situation. Not "B2B marketers". A demand generation lead at a company of fifty to three hundred people with a CRM, no data engineer, and a board asking why pipeline is flat.
The single claim. One sentence the piece argues that a reasonable person could disagree with. If it cannot be contradicted it is a topic, and you will get a summary.
The evidence. The mechanism, the numbers you may use, the check the reader can run. An agent with no evidence manufactures plausible evidence, the most expensive failure in the pipeline.
The banned claims. Roadmap features described as present, comparisons legal will not sign, any statistic you cannot source, superlatives. A negative list works better than an instruction to be accurate.
The format. Word range, section count, whether lists are allowed, where internal links go, what the opening must not do.
Two examples, one good and one bad. A paragraph from your archive that works and one that is on-topic and lifeless. The contrast teaches more than any adjective.
Such a brief takes twenty to thirty minutes. The draft takes ninety seconds. That ratio feels wrong to anyone expecting agents to save writing time, and it is correct.
Context files beat clever prompts, because prompts do not survive
A clever prompt lives in one person's browser tab. It is not reviewed, not versioned, not shared, and when that person leaves the quality leaves with them. Context files are the alternative: documents in the repository that every content agent reads on every run.
Five are enough. A company file: what the organisation does, for whom, and what it is not. A proof file: every claim you may make, with its evidence and approval status, so the agent has somewhere to check rather than somewhere to guess. A product boundary file: what exists today and what is in development and must never be described as present. A voice codex: not adjectives but rules, with banned constructions. And an inventory of what you have published, each with its claim, so the agent knows what not to retread.
These get reviewed like code, because that is what they are. The marketing as code argument applies directly: a change to the proof file changes what the organisation may assert. The agent skills that do the drafting are thin on top of these files, and the files are where the quality lives.
Agents draft, outline, vary and repurpose. They do not decide
Agents are good at a first draft from a decided brief, at turning messy notes into competing outlines, at generating twenty headline variants once a direction is chosen, at repurposing a finished piece into another format without changing its argument, and at summarising research with sources attached. All of these share one property: the judgement has already happened and only production remains.
Agents are bad at three things no prompt fixes. They cannot have an opinion, having no stake and no experience of being wrong in public. They do not know what is true about your product, and are confidently wrong about it the way a new hire on day two would be, except that they never learn from the correction unless you write it into a file. And they cannot judge what is interesting, because interesting depends on what your audience already knows, while the model's sense of novelty is calibrated to the whole internet.
A check that takes ten minutes: read the last agent draft you published sentence by sentence and ask whether a competitor could publish each one unchanged. If more than half survives, you did not publish content, you published inventory.
Review is editing, and it needs the person who holds the claim
The review step fails when it goes to whoever has capacity, because the defects that matter are invisible to anyone who does not own the argument. Three passes, in order. Is the claim still the claim the brief specified, or did the draft drift to a safer adjacent one? That drift is the commonest failure and it happens quietly. Second, is every factual assertion traceable to the proof file? Third, does it sound like the organisation rather than like a model, judged against the voice codex and not against taste.
Budget fifteen to thirty minutes per piece. If a reviewer spends an hour rewriting, the brief was wrong rather than the draft, and the response is to fix the brief and regenerate.
Measure with an eval set, not with vibes
Every change to a prompt, a model version or a context file moves output quality in ways nobody notices for weeks. The fix is borrowed from engineering practice: a fixed evaluation set. Build twenty briefs covering the range of work you commission and run them whenever anything upstream changes. Score each output on binary checks a second reader can apply consistently: the claim is present and argued, no banned claim appears, every number is traceable, format constraints are met. Twenty briefs on four checks is a two-hour exercise that turns "the new model feels worse" into evidence.
In production, track the share of drafts accepted with light edit and the share deleted outright. A rising delete rate is a context problem, not a prompt problem, every time.
The failure mode is volume without a point of view
The tempting conclusion is that content has become cheap, so publish more of it. That is the failure, and it is punished twice. Search engines increasingly evaluate sites rather than pages: a library of unremarkable pieces does not lift the good ones, it dilutes them, sitewide. Readers punish it faster and more quietly, because someone who finds two of your pieces forgettable will not read a third, and that decision appears in no dashboard.
There is no version of this where the agent supplies the point of view. If the organisation has no argument about its market, agents let you publish the absence of one ten times faster. No tool removes that.
Frequently asked questions
What should a content brief for an AI agent contain?
Six things: the audience described by situation rather than demographics, a single contestable claim the piece argues, the evidence that claim rests on including which numbers may be used, a list of banned claims such as roadmap features and unsourceable statistics, the format constraints, and two example paragraphs, one good and one deliberately mediocre. The examples do more work than any tone instruction.
Why are context files better than prompts for content agents?
A prompt lives in one person's session, is not versioned or reviewed, and does not survive that person leaving. Context files are documents in a repository that every agent reads on every run: a company file, a proof file of approved claims with evidence, a product boundary file stating what does not yet exist, a voice codex with banned constructions, and an inventory of published work. They can be reviewed and audited when an agent gets something wrong.
What is AI content actually good at?
First drafts from a decided brief, competing outlines from messy notes, variant generation once a direction is chosen, repurposing a finished argument into another format, and summarising research with sources attached. These are production tasks where the judgement has already happened. Agents are poor at having an opinion, at knowing what is true about a specific product, and at judging what an audience will find interesting.
How do you measure the quality of agent-produced content?
With a fixed evaluation set rather than by impression. Keep roughly twenty representative briefs, run them whenever a prompt, model version or context file changes, and score each output on binary checks: the claim is argued, no banned claim appears, every number is traceable, format constraints are met. Report a pass rate. In production, track the share of drafts accepted with light edit and the share deleted outright, and treat a rising delete rate as a context failure.
where this lives in the system
shorter reads on this, at aiporate.com
- PlaybooksContent Briefs That Prevent Bad Drafts Instead of Documenting Them Afterward
- PlaybooksAI Content at Scale: The Quality Guardrails That Keep It From Reading Like AI Slop
- RevOpsA Content Production Process That Ships on Schedule
- PlaybooksThe Editorial Calendar as an Operating System, Not a Spreadsheet Graveyard
see where you stand
Twelve questions. Then your build order.
The diagnostic returns your operating stage, the three widest gaps in your motion and what to build first. Two minutes, no sales sequence, one human reply.