Anatomy of a Four-Agent Cold Outreach Pipeline
One cold outreach prompt demoed beautifully and fell apart on repeat runs. Splitting it into four agents removed 40+ hours of manual work a week for a client.
The first version of this cold outreach system was one prompt. Crawl the website, analyze it, write the email. It demoed beautifully. I could run it in front of someone, watch it produce a genuinely sharp piece of outreach, and feel like the problem was solved.
Then I ran it three times on the same website.
Three runs, three different conclusions. One pass fixated on technical SEO. The next one talked about the design. A third ignored the single most obvious problem on the page and wrote confidently about something trivial. Same input, same prompt, different priorities every time.
That’s when it stopped being a demo problem and became an engineering problem. Not because the output was bad. Because the output wasn’t reproducible, and anything you can’t reproduce, you can’t evaluate. I had no way to answer the only question that matters in production: when this sends a bad email, which part broke?
A single prompt has no seams. It’s one opaque decision. You can’t put a probe inside it, you can’t test a stage in isolation, and you can’t fix one behavior without disturbing the others. It’s a monolith, and it fails like one.
So I broke it into four agents, each with one responsibility and a structured handoff to the next. The final email stopped being one enormous guess and became the result of several small decisions, each of which I could inspect.
The system runs in production for a client today and removes more than 40 hours of manual research and outreach work every week. This is how it’s built, what broke along the way, and the thing I got wrong for months.
Key Takeaways
- A single prompt asked to analyze, prioritize, verify, and write produces different priorities on identical input. Splitting it into four agents creates inspection points between stages, so a bad email can be traced to the stage that caused it.
- The bottleneck was never prompt engineering. It was input quality: JavaScript rendering, cookie banners, page builders, and missing metadata.
- Emails stopped sounding like AI when the writing agent was given less information, not a better prompt.
- Finding nothing worth saying is a valid outcome. The system falls back rather than inventing a problem.
Why couldn’t one prompt run cold outreach on its own?
Because one prompt was doing four jobs at once, and nothing forced it to do them in order. Analyzing a website, deciding what mattered, checking its own reasoning, and writing an email are four different kinds of thinking. Bundled into a single call, they compete. The model would spend its attention on whatever the page surfaced loudest, not on what was actually worth saying.
The failure mode wasn’t a bad email. It was variance.
Repeated runs of the same single-prompt pipeline against an identical website produced different priorities, different conclusions, and completely different outreach. That variance made evaluation impossible: with no seam between analysis and writing, there is no way to attribute a bad email to the stage that caused it.
Splitting the workflow fixed that, and the reason is unglamorous. Four agents give you three places to look. When an email goes out wrong now, I can read the audit output and see whether the observation was wrong or the writing was. That’s the whole benefit. Not intelligence. Attribution.
Each agent has one responsibility, produces structured output, and hands only that output forward. The pipeline is sequential, and each stage receives only the context relevant to its task. That constraint isn’t a limitation I worked around. It’s the thing that makes the output predictable.
What do the four cold outreach agents actually do?
The pipeline runs in order: find the lead, audit the lead’s website, build the personalized presentation, send the email. Each stage narrows the input for the next one.
Agent 1 finds qualified leads against the client’s ICP. Nothing downstream matters if this is wrong. A perfect email to the wrong company is still noise.
Agent 2 audits the lead’s website. This is where most of the real work happens, and it’s covered in detail below.
Agent 3 generates the personalized presentation using the outputs of agents 1 and 2.
Agent 4 sends a conversation-starter email that references one specific finding from the audit by name.
That last sentence is the entire system’s reason for existing. Every other component exists to make one sentence in one email specific and credible. Not the whole email. One sentence. That’s the payload; everything upstream is delivery infrastructure.
The stack is deliberately boring: Next.js with TypeScript, deployed on Vercel, with an orchestration layer coordinating the agents sequentially. For the AI layer I use OpenAI models and pick different models depending on the task. The goal isn’t to maximize intelligence at every step. It’s to balance quality, speed, and cost. Most steps don’t need the smartest available model. They need a predictable one.
The bottleneck was never the prompt
The biggest surprise in production wasn’t the language model. It was websites.
I expected to spend my time tuning prompts. Instead I spent it fighting the input. A lot of websites simply don’t expose clean information. Heavy JavaScript rendering hides the content from a naive fetch. HTML is inconsistent. Cookie banners sit in front of everything. Page builders generate markup that means nothing. Content is duplicated across templates. Metadata is missing entirely.
Feed that to a good model and you get a confident audit of a cookie banner.
The quality ceiling of an agent system is set before the model generates a single token. Most of the engineering effort in this pipeline went into producing better context for the agents, not into making the prompts longer.
This is the part I’d tell anyone building something similar, and it’s the least fashionable advice available. Prompt engineering is where everyone starts because it’s the visible surface. Context engineering is where the results actually come from, and it looks like unglamorous parsing work: extracting real content from rendered pages, normalizing inconsistent markup, deciding what to throw away.
Garbage in, confident garbage out. The model doesn’t tell you the input was bad. It just writes something plausible about whatever it was handed. That’s the dangerous part.
How does the audit agent decide what matters?
It doesn’t try to produce a 40-page SEO report. Its job is narrower and harder: find an observation that could start a real conversation.
The agent looks across messaging clarity, value proposition, CTA visibility, UX friction, trust signals, branding consistency, content quality, basic technical SEO, and conversion opportunities. That’s a wide net. The important part isn’t the net. It’s what gets thrown back.
Every finding passes two filters:
- Is this actually true?
- Would a business owner care?
Only findings that survive both move to the writing stage. Most don’t.
That second filter kills more than you’d expect. Plenty of findings are technically correct and completely uninteresting. Nobody replies to an email because their h1 is missing. A filter that only checks correctness produces accurate, boring outreach. The “would anyone care” test is what turns an audit into a conversation.
If you want to see the scoring half of that logic in a much simpler form, the SEO scorer and lead scorecard on this site apply a comparable idea: score the thing against a fixed set of criteria, decided before any input arrives.
How do you stop an AI email sounding like AI?
You give the model less, not more.
This was the hardest part of the project, and every instinct I had about it was wrong. I kept trying to solve it with better prompts. More instructions, more examples, more careful phrasing about tone. It never worked. The emails stayed recognizably synthetic.
The fix was reducing the model’s freedom.
The writing agent never sees the whole website. It receives one or two validated observations and a strict set of writing constraints. It isn’t allowed to exaggerate. It isn’t allowed to invent compliments. It isn’t allowed to stack multiple insights into one paragraph, which is the single most obvious tell in AI-written outreach.
The objective isn’t to sound impressive. It’s to sound like someone who actually spent five minutes looking at the website.
Restricting the writing agent to one or two validated observations produced more human-sounding emails than any prompt refinement did. Given the whole website, the model performs comprehensiveness. Given one fact, it just says the fact.
That’s the counterintuitive result, and I think it generalizes. A model handed everything it knows will try to demonstrate that it knows everything. A person who noticed one real thing about your site mentions one real thing. The constraint isn’t a workaround for a weak model. The constraint is the product.
What happens when there’s nothing worth saying?
The system sends a less personalized email. That’s a feature, and it took me a while to accept it.
If the audit can’t find a meaningful observation with enough confidence, it falls back to broader business context instead of inventing a problem. It doesn’t manufacture an insight to satisfy the template.
This matters more than it sounds. The whole premise of the pipeline is that the recipient can tell the difference between someone who looked and someone who pretended to look. A fabricated observation doesn’t just fail to land. It actively proves you didn’t look, which is worse than never claiming to have looked at all.
I’d rather send a credible email with less personalization than a highly personalized email built on something that isn’t true. Trust matters more than cleverness. A system that can’t say “I’ve got nothing here” will eventually say something false, at scale, with your client’s name on it.
What I’d build differently
Less prompt engineering. More context engineering.
If I rebuilt this today, the time would go into structured data collection, better website parsing, richer business context, and feedback loops from actual outreach performance. Almost none of it would go into the prompts.
The agents aren’t the competitive advantage. Anyone can write four prompts. The quality of the information those agents receive is the advantage, and that’s a data problem wearing an AI costume.
The full engagement, including what the client owns at the end of it, is written up in the AI agents case study. If you’re building something in this shape, that’s the work I do.
Frequently Asked Questions
Why four agents instead of one prompt? A single prompt asked to analyze, prioritize, verify, and write produces different priorities on identical input. Four agents with structured handoffs create inspection points between stages, so a bad output can be attributed to the stage that caused it. The benefit is debuggability, not intelligence.
What is context engineering? Context engineering is the work of producing high-quality, structured input for a model before it generates anything: parsing, normalizing, extracting, and discarding. In this pipeline it mattered far more than prompt wording, because the quality ceiling is set before the first token.
Why do AI-written cold emails sound fake? Usually because the model was given too much information. Handed an entire website, a model performs comprehensiveness and stacks multiple insights into one paragraph. Restricting it to one or two validated observations produces writing that sounds like a person who noticed one real thing.
What happens if the audit finds nothing useful? The system falls back to broader business context rather than inventing an observation. A fabricated insight proves you didn’t actually look, which is worse than not claiming to have looked. Credibility with less personalization beats personalization built on something untrue.
What does the pipeline actually run on? Next.js with TypeScript, deployed on Vercel, with a sequential orchestration layer passing structured context between agents. The AI layer uses OpenAI models, chosen per task to balance quality, speed, and cost rather than maximizing intelligence at every step.
Final thought
The interesting thing about this system isn’t that it uses AI agents. Four prompts in sequence is not an achievement. The interesting thing is how much of the work turned out to be ordinary engineering: parsing messy HTML, defining structured handoffs, deciding what to discard, and building enough seams into the thing that you can tell which part lied to you.
The AI part took days. The context part took months.
That ratio is the actual lesson, and I don’t think it’s specific to outreach. As models get better, the prompt matters less and the input matters more. The advantage moves upstream, toward whoever can hand the model something clean and true.
Everyone is optimizing the part that’s about to become free.