Every agentic marketing pilot looks good in the demo. Then it goes live.

The agent personalizes emails on cue, creates a campaign brief in seconds, pulls from your CRM, produces something that would’ve taken a coordinator two hours, and leadership greenlights the pilot. Three months later, the production rollout is in triage.

This triage isn’t new, but it’s entirely preventable. This isn’t about bad technology, but rather what pilots test and what they quietly leave out.

The pilot was never testing what you thought it wasPilot vs. production conditions for agentic marketing: pilots use clean inputs, hand-picked use cases, and human oversight, while production involves messy data, unpredictable use cases, less human oversight, and greater risk.

“Works” in a pilot means something different than “works” in production.

Pilots test whether an agent can do what you ask. Production tests whether it should do what you ask, consistently, in conditions it’s never seen before. In a pilot, a person preps the inputs, the team hand-picks the use cases, and there’s always a human nearby ready to catch anything off base. The whole setup exists so the technology can succeed.

In the real world, production strips all of that away.

What actually breaks

The data isn’t what the demo promised. In a pilot, someone cleans the data. In production, the real data typically includes inconsistencies, missing fields, legacy naming conventions, and cases no one flagged during scoping. The dangerous failure mode isn’t a loud crash, it’s the agent quietly making reasonable-sounding inferences from bad inputs and producing output that looks fine until a human who really knows the account QA’s it.

To create a better process, stress-test against real, uncurated data before you scale, not theoretical edge cases.

Marketing teams never built governance for this. Most marketing governance grew up around human handoffs: someone reviews a brief, someone approves a campaign, and a human reads the output before it goes anywhere. Agentic workflows skip those handoffs, making the process faster, but speed without guardrails creates high-volume, low-visibility errors that compound fast.

The goal isn’t to slow automation down, it’s to make oversight a feature of the workflow, not an afterthought bolted on at the end.

Ownership isn’t clear. In a pilot, ownership is obvious. In production, the agent touches five different teams, so when something goes wrong, it’s unclear whether it’s a data problem, a prompt problem, a logic problem, or a trafficking problem.

To create a clear structure, map accountability before launch to clearly define who owns input quality, output review, and incident response.

The prompt drifts without anyone tracking it. A line gets tweaked to fix a specific output, while the tone gets softened after a client complaint. Without intention, the prompt drifts from what you validated, with no record of why, and the provider has quietly updated the model underneath. You’re no longer running what you tested.

To maintain the prompt, treat it like any other piece of production logic. Version it, document why you made each change, and assign someone to own it.

The agent doesn’t know what it doesn’t know. Campaign logic that’s right in Q3 may be wrong in Q4 when the brand is managing a PR situation. Guidelines for one product segment may not transfer cleanly to another. It has no way to flag any of this on its own. It keeps producing output that looks confident and complete, but is subtly wrong in ways that only someone close to the work would catch.

To maintain integrity, define what you authorize the agent to handle independently, and build escalation logic for everything outside that.

Why this keeps happening

Companies fund pilots to prove feasibility, yet production readiness is a different workstream and rarely gets the same investment. Teams get credit for getting to production, not for asking what production actually requires, and vendors want to move quickly from pilot to contract.

The fix isn’t slower pilots. It’s treating production readiness as its own body of work rather than a checklist you knock out in the final week.

Questions worth asking before you scale

  • When the agent gets input it’s never seen, does it fail gracefully or keep going?
  • Which governance checkpoints are now gone? Was that intentional?
  • Who owns the prompt, and what does the change process look like?
  • If the agent produces bad output at scale, what’s the response path?
  • How will you know six months from now if performance has degraded?
These questions aren’t complicated, they just tend to get skipped in the rush to ship. The pilot proved the concept. Production is where you prove it actually works.

Abby Mates, Principal

Transparent Partners
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.