Lee
Basnight
MF&K
Generative AI quality framework for a creative agency
A creative agency wanted to incorporate generative AI into their creative pipeline. The AI kept letting them down. The outputs were good enough to be tempting and shaky enough, that creative directors quietly went back to doing it by hand. I owned the fix. End to end, across every format the agency worked in, I built the prompt frameworks and the reliability guardrails that made AI output something you could actually ship, then wrote it all down as the standard everyone could work from.
- I created a measurable quality lift in the AI output, driven by prompt engineering and validation guardrails.
- Assessed capability assessments with ROI modeling that told the agency which tools were worth buying.
- Created Best-practice documentation for the company to adopted as its standard operating procedure.
The Problem.
The agency was already running generative AI across all its creative work. The trouble was consistency. One output would land client-ready, the next would need so much cleanup that it had to be remade by hand. You cannot sell a client on AI-augmented creative when the team is quietly redoing the work. So the real question became specific. Where was all that variance coming from and what could I do about it? Nobody doubted the AI could produce good work. The job was getting it to produce good work every single time.
My Role & Scope.
I came in as the AI Creative Strategist, which in practice meant I owned the entire generative AI operation. The systems, the workflow, the quality bar and the standards every format had to clear before it went out the door. I wrote the company-wide best practices, trained the team on them and ran the capability assessments and ROI modeling that decided which tools we actually purchased. I reported into creative leadership and had a real hand in what we spent money on.
Approach.
Diagnose before you build. I pulled a large sample of recent AI outputs from across the agency and traced every failure back to where it went wrong. Once I lined them up, the pattern was obvious. Everything failed in one of three places. The prompt was too vague going in. The output never got checked against the client's brand in the middle. The tone or voice never got a final read before it shipped. From there; I built a three-layer filter. Every output had to clear all three layers before it could reach a client.
The three-layer filter. Every AI output clears prompt spec, brand validation and a tone check before a client ever sees it.
Key Decisions & Rationale.
- I anchored the entire framework to failure modes. Tools come and go every few months; the ways AI output fails stay remarkably constant. Anchor a framework to a specific tool and it dies the day that tool gets deprecated. Anchor it to the failure modes and it keeps working through every tool the agency swaps in. So I built it around the part that stays consistent.
- Prompt templates were organized by use case, versioned like code. A prompt written for one project dies with that project; nobody ever improves it. A versioned template for a use case gets sharpened across hundreds of projects. Every tweak makes the next job a little better. That compounding efficiency is the entire reason to templatize in the first place.
- Validation lives inside the workflow as a real, runnable step. A checklist is the first thing that goes when a deadline hits; a step that sits in the workflow gets done, because at that point skipping it is more work than running it.
- One authoritative document beats a dozen agreements buried in chat threads. A document that names the failure mode and spells out the fix is auditable, teachable and still there after the turnover. Engineering figured this out years ago with runbooks. Creative work needs the exact same thing.
Artifacts Produced.
- The three-layer quality framework itself: prompt template, validation rule, brand and tone check.
- A prompt template library sorted by use case, every template versioned so it only gets better.
- A capability assessment matrix that scored any new tool on workflow impact and ROI before anyone bought it.
- Fine-tuning experiments run across every format the agency worked in, each with a documented baseline.
- The company-wide SOP that became the standard for anything AI touched before it shipped.
What Didn't Work.
The first version of that template library was organized by client. Inside a month, three different clients had three slightly different prompts for the exact same requests. Nobody could say which one was right. That duplication was the tell that the whole structure was wrong. Fixing it meant flipping the library to a use-case-first layout.
What Stuck.
The quality lift was the headline outcome. You could see it in the drop in revision cycles per deliverable. The thing that lasted longer was the SOP. Someone could join the team, read it in under an hour and start working, that same day.