By clicking “Accept All Cookies”, you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. View our Privacy Policy for more information.
Revenue Operations

When AI Agents Break: How RevOps Teams Build Accuracy Over Time

swirled squiggle accent

Revenue teams should know one thing before deploying AI agents: the work isn’t really in the tooling. Pick the right platform, flip the switch, and let the automation run. The people who aren’t in the know discover — usually after weeks of messy outputs and frustrated stakeholders — that the real work is in the instruction design, the memory architecture, and the iterative systems they either build or don't.

‍

In a recent RevOps Co-op webinar, Matthew Volm sat down with Alex Avila, a solutions engineer at Nooks, and Gerard Martelly, an AI advocate and revenue enablement leader at Vapi, to dig into what actually makes AI agents reliable in RevOps workflows. The conversation moved quickly from theory to practice — covering prompt engineering, agent architecture, deterministic versus non-deterministic workflows, and real production examples from both speakers' teams.

‍

The session opened with an audience poll that revealed a familiar pattern: most practitioners on the call were already experimenting with agents in their workflows, but few felt confident in the accuracy or consistency of what they were building.

‍

The Intern Who Went to Harvard (But Still Needs a Manager)

Before any discussion of tools or architectures, Martelly offered a reframe that anchored the rest of the conversation. The problem most teams experience with AI agents is not the model. It's the expectation.

‍

"Everybody thinking that agents are supposed to give you the perfect answer, tell you everything you need, but that's not really the way they work. An agent really is just an intern who got a degree from Harvard, Yale, Princeton, triple major, got the master's degree. It's actually a postdoc right now, but he's found time to do an internship with you. And an agent really is only as good as you are able to give it direction and like standard operating procedures and instruction." — Gerard Martelly

‍

This framing cuts to the core of what makes agents underperform in practice. The failure is a management failure before it is a technical one. Teams that get frustrated with hallucinations and inconsistent outputs are often teams that gave an under-briefed intern an open-ended task and then blamed the intern.

‍

Avila reinforced this: every AI model, regardless of provider, is generative. Every output is technically brand new, even when the prompt is identical. That inherent variability means the human-designed harness — the instructions, constraints, and environment the agent operates within — determines everything.

‍

What "Harness" Actually Means (And Why It's Not a Buzzword)

The term "agent harness" has become something of a trendy abstraction in RevOps circles. Martelly broke it down to something more immediately usable.

‍

"People will say this harness, and it's like a fancy buzzword for what really just means instructions. The instructions that you give it to operate and the confinement or the environment you allow it to operate within. If I am using an AI agent inside of Notion, Notion therefore is the harness for it." — Gerard Martelly

‍

Whether you're using Clay, n8n, HubSpot workflows, or a custom agent built in Claude's project environment, the harness is whatever structures the agent's behavior and limits its scope. And the quality of that harness depends on one thing that has nothing to do with the tool: clarity about the job to be done.

‍

Martelly outlined a pre-harness exercise to do before any agent build: define the inputs you have, the output you're looking for, and any steps the agent needs to take in between. That clarity, established before touching a prompt, produces dramatically more useful agents than diving into tool selection first.

‍

This connects directly to a pattern RevOps Co-op has explored in the context of AI readiness more broadly — the teams that get useful outputs from AI are the ones that did the definitional work first.

‍

Building a Prompt That Actually Works: RICE and Other Frameworks

Once you have a harness concept in mind, the prompt engineering begins. Avila introduced a practical acronym — RICE — as a baseline for structuring prompts when you're iterating quickly and need a reliable starting point.

‍

The four components: Role (define who the agent is and what capacity it's acting in), Instruction (specific, detailed guidance on what to do), Constraints (what it cannot do, what limits apply), and Example (one or more demonstrations of what a correct output looks like — this is what AI researchers call one-shot or few-shot prompting).

‍

"This is a very easy way to just — when you're trying to rip through iterations and you wanna dial something in, you gotta try a lot of different things — this is a good way to do it." — Alex Avila

‍

Martelly added his own checklist for what a complete prompt needs to specify: the exact output format, the inputs the agent should expect, verification and validation steps, a test at the end, and ideally a command to log every result to a database. That last point — observability — is the one most practitioners skip, and it's the one that makes iterative improvement possible.

‍

On the topic of custom instructions, Martelly shared a technique that most teams aren't using: configuring system-level instructions that act as a standing contrarian voice. His standard setup includes a directive to treat everything it produces as provisional until verified, flag the difference between facts and inferences, and provide a confidence level with every response.

‍

"I want you to act as a contrarian. I want you to help me see my blind spots. If something is a fact, I want you to know it is a fact. If something is an inference, I want you to know it is an inference. I want you to give me your confidence level on the response as well too." — Gerard Martelly

‍

This kind of epistemic discipline — baked into the custom instructions so it travels with every session — is a meaningful upgrade over the default mode, where agents confidently produce whatever output fits the prompt without surfacing uncertainty. For RevOps use cases where outputs feed into CRM data, forecasting logic, or deal routing decisions, that difference matters operationally.

‍

The challenge of getting consistent, trustworthy data out of AI connects to a broader problem the community has been working through — as explored in this discussion of AI readiness and data governance.

‍

Model Temperature, Determinism, and When Not to Use an Agent

One of the most practically useful segments of the session was a discussion of when not to use an agent at all. The concept of model temperature — the parameter that controls how creative or rigid an agent's output is — came up as a tool for managing consistency, but Avila pushed further: some workflows should never route through an agent in the first place.

‍

"If I need something done reliably, and I need the same caliber of excellence and something done exactly the same every single time, I'll use an automation." — Gerard Martelly

‍

The practical distinction Martelly and Avila drew: the quality of your input data largely determines whether an agent is appropriate. Structured, clean, consistently shaped data — form fills, CRM fields, deal stage data — routes naturally to deterministic logic. Conditional filters, HubSpot workflow branches, and n8n automations will outperform an agent every time in these scenarios, because the agent introduces unnecessary variability into a situation that doesn't need it.

‍

Agents become valuable when the input is unstructured and the task requires judgment. Research across heterogeneous data sources, tone evaluation of call transcripts, classification of free-text fields, synthesis across multiple documents — these are the use cases where deterministic logic hits its limits and agents earn their place.

‍

Avila offered a concrete test: "If I could do this with a formula or a filter, should I? Probably, yes. The only reason to bring in an agent is when the formula can't handle the variation in the data." This is a harder constraint to honor than it sounds, because agents feel powerful and versatile in a way that can tempt teams to route everything through them.

‍

The parallel to standard RevOps process design is worth sitting with. Over-engineering workflows is a known failure mode — and the temptation to reach for AI when a simpler conditional would do is the same pattern, one layer up.

‍

From Clay Table to Self-Learning Agent: Avila's Production Case Study

The most instructive part of the session was Avila's detailed account of how a real RevOps-adjacent workflow evolved from a naive first attempt to a production system with meaningful accuracy. The workflow in question: automating the qualification of inbound deals at Nooks to determine whether a prospect was a strong technical fit — without requiring a solutions engineer to manually review every opportunity.

‍

Version one was a single monolithic agent. It ingested Gong call transcripts, HubSpot deal and company data, and a detailed prompt, and produced a classification with a confidence score. Routing logic downstream would either progress the deal or flag it for human review.

‍

"This thing was not accurate, I would say maybe sixty percent of the time. And that wasn't scalable because I had to go clean up the mistakes, or I would have to ask my team to go investigate these deals manually anyways." — Alex Avila

‍

The fix for V2 came from a recommendation to replace the single reviewer with a parallel panel. Five agents, running simultaneously on the same data, each producing independent evaluations. A reconciler agent then reviewed the five outputs and resolved disagreements by majority — a technique borrowed from ensemble machine learning methods applied to LLM evaluation.

‍

"He's like, 'Trust me, you gotta just throw this thing on like Opus five. You have to. Don't — this is not a place to cheap out.'" — Alex Avila

‍

V2 reached approximately 85% accuracy, a significant improvement. But it hit its own limits: Clay table row caps, schema failures when data sent to HubSpot didn't match the expected shape, and the fundamental problem that Nooks's product was evolving so fast that "what constitutes a good fit customer" was changing every few days — requiring constant prompt updates across five parallel agents.

‍

Version three replaced the Clay foundation with Nooks's own self-learning agent, specialized each of the five reviewers with domain-specific prompts rather than identical tasks, replaced the reconciler with a deterministic evaluation tool (Leia), and moved to Temporal for durable workflow orchestration.

‍

The result: approximately 90% accuracy at a fraction of the inference cost, with a learning loop that updated the agent's memory every time a human approved or rejected a routing decision.

‍

The through-line across all three versions is the same: accuracy is a system design problem before it is a model selection problem. Better prompts, better structure, better feedback loops — these matter more than the underlying model.

‍

Gerard's Manager-Worker Architecture: Scoring Account Potential at Vapi

Martelly's example came from the other end of the complexity spectrum: a multi-agent system built to evaluate and score account potential at Vapi, where traditional firmographic signals like headcount are poor predictors of revenue potential in an AI-native customer base.

‍

The architecture: a manager agent receives new accounts from Salesforce and orchestrates a set of specialized sub-agents — a research agent, an estimate agent that models revenue potential, and an independent validation agent that verifies the other agents' work before anything gets recorded. Everything routes back to the manager, which logs every run, every failure, the reason for the failure, and the fix it applied.

‍

"The manager agent not only records every run, records every failure, and it also records its fix for the failure and how it got it to go along the process, which allows us to then look at everything at every week. And we have another — the manager agent produce us a report every week that says, 'This is what happened. These were the failures. This is why it failed. This is how we can update the instructions to ensure that we're not failing anymore.'" — Gerard Martelly

‍

The self-healing loop here is significant. Rather than requiring a human to diagnose failures and rewrite instructions manually, the manager agent is instructed to look for the smallest possible corrective change, apply it, and report on it — a form of automated prompt governance.

‍

When confidence thresholds aren't met, the system routes to a human review loop. The human's decision — and their justification — feeds back into the weekly system update, creating a continuous improvement cycle that gets better with every edge case it encounters.

‍

The architectural choice to separate the research, estimation, and validation tasks across different agents rather than loading all three into a single agent addresses a real technical constraint: token overload. A single agent processing all that context produces more hallucinations and more incomplete outputs than a coordinated set of specialized agents working sequentially.

‍

This mirrors a principle from organizational design that Martelly made explicit: just as you wouldn't ask one person to simultaneously gather research, build a financial model, and quality-check their own work, you shouldn't build agents that way either.

‍

Memory Architecture: How Agents Learn From Their Own Mistakes

The conceptual connective tissue across both case studies is agent memory — the mechanism by which a system can learn from its own history rather than starting fresh with every run.

‍

Avila walked through the spectrum of memory implementations, from the simple (a folder of markdown files that a Claude project reads before responding) to the production-grade (a separate workflow that evaluates every completed run, scores it, and stores the result in a structured database that the agent references on future runs).

‍

"We take all the details from that, and then we ask a separate agent, 'Was this accurate or inaccurate? Was this good or not good? Rate this agent's performance.' And then what we can do is tell it, 'Hey, if it's accurate, I need you to store this away, file it, put it in the filing cabinet, put it in the recipe box for later,' so that it can reference it in the future." — Alex Avila

‍

The filing cabinet metaphor is useful because it distinguishes agent memory from context window — two things that often get conflated. Context window is what the agent can see in a single session. Memory is what persists across sessions and accumulates over time. The latter is what enables genuine improvement rather than just recall.

‍

Martelly connected this back to the management analogy that opened the session: the upfront investment in building a memory system is the equivalent of onboarding a new team member properly. You spend more time at the beginning so you can be more hands-off later.

‍

"You need to give, put more time in in the very beginning, uh, when you're building these systems. You need that observability and memory like we talked about, just so you can check its work, right? I wouldn't give a tool or I wouldn't give a task to Matt and be like, 'Hey, Matt, I need you to run analysis and run all of our Q3 renewals on his first day.'" — Gerard Martelly

‍

This is not a glamorous principle, but it's the one that separates teams building durable AI infrastructure from teams perpetually firefighting broken automations. As the RevOps Co-op community has discussed, the boring foundational work behind great AI is almost always what determines whether the sophisticated layer on top actually delivers value.

‍

Output Schemas, JSON, and Making Sure the Agent Says What You Need

One section of the session that deserves more attention than it typically gets in AI discussions for RevOps teams: output schemas. When an agent's output needs to feed into a CRM field, a routing workflow, or a database record, the text it produces has to match a specific structure. A single-select field can't accept a paragraph. A Boolean field can't accept "probably."

‍

Avila framed the issue directly: most RevOps practitioners are working with typed data in their CRMs and automation tools. That means AI outputs need to conform to specific schemas — and the agent needs explicit instruction on exactly what format to use.

‍

The barrier for many practitioners is JSON syntax, which feels technical enough to be intimidating. Martelly's advice was characteristically practical:

‍

"I'm gonna be real vulnerable. I'm gonna drop a PDF of some of my chats with AI later, and then I'll send them to you, and you guys will see how ridonculous my conversations with AI are. It's mostly me just asking questions." — Gerard Martelly

‍

Asking the agent to suggest its own output format — giving it three options and letting you choose — is a legitimate and effective strategy. The point is not to master JSON syntax but to develop a working relationship with the tool that surfaces options you can evaluate. Resourcefulness, in this context, is the skill that scales.

‍

Key Takeaways

  • Agent accuracy is a management problem, not a model problem. Imprecise instructions produce inconsistent outputs regardless of which underlying model you use. Treat agent briefing with the same rigor you'd apply to onboarding a new team member.‍
  • RICE is a reliable prompting baseline. Role, Instruction, Constraints, and Example is a repeatable structure for building prompts when you're iterating fast and need to dial in consistency.‍
  • Audit logs are non-negotiable. Every production agent workflow should log its outputs, flag failures, and record the reason for each outcome.‍
  • Use deterministic logic when the data allows. Filters, conditional branches, and automations outperform agents on clean, consistent inputs. Reserve AI for unstructured and ambiguous data.‍
  • Parallel agents outperform monolithic agents on complex evaluations. Running five specialized reviewers and reconciling their outputs is more accurate than asking one agent to handle everything — and more robust to individual model failures.‍
  • Agent memory is what turns a one-time workflow into a learning system. Without a mechanism to store and reference historical runs, agents reset with every execution. A simple feedback database — even a folder of markdown files — changes this fundamentally.

‍

The underlying current across every section of this session is the same: AI agents for RevOps are an instruction design challenge before they are a technology challenge. The teams making real progress are not the ones who found the best model or the most sophisticated platform — they're the teams that invested in clarity, observability, and structured feedback loops from the beginning, and built from there.

‍

Learn more about how Nooks helps revenue teams accelerate pipeline through AI-powered sales engagement and agent-native workflows built for modern go-to-market (GTM) teams.

‍

Looking for more great content?

‍

Check out our blog, join our community and subscribe to our YouTube Channel for more insights.

Related posts

Membership Options

Pick your path

Join 21,000+ RevOps pros. Start free, or unlock the full community and everything that comes with it.

Free 👋
Free forever

A newsletter full of best practices and lessons learned, plus a heads-up on upcoming in-person and digital events.

Get started
Starter 🕺 Popular
The full community

Our community of thousands of RevOps pros, plus a peer-to-peer matching program.

Join Now
Rise 🚀
Everything, leveled up

Everything in Starter plus course and conference discounts, our knowledge hub, career coaching, and more.

Join Now

Reach your RevOps goals

Our average member has 5+ years of RevOps experience — so you'll have real-time access to seasoned pros. All we ask is that you're generous with your knowledge in return.

Become a Member →