AI Marketing Advisory / AI Readiness: how to choose a provider

How to vet an AI marketing advisor, the proof of real deployments to demand, and the red flags that signal repackaged prompt training.

Written By
Cedric Pharand
Verified By
Zahra Sanati
Marketing Strategy & PR
MAKE US A PREFERRED SOURCE
Read time:
5 min
Published:
September 18, 2026
Updated:
September 18, 2026

Table of contents

Summarize this article with AI

AI Marketing Advisory / AI Readiness: how to choose a provider — Web Tonic article thumbnail

Almost anyone can now describe an AI marketing programme convincingly. The only reliable filter is evidence of things shipped, owned and measured — asked for in a way that cannot be answered with a deck.

Key Takeaways

  • Ask for one workflow the provider shipped in their own business, with before-and-after numbers. Client name-dropping is not an answer.
  • Run a 48-hour test: prompt shops return screenshots and mockups, real builders return a working link with edge cases handled.
  • Insist on a live demo rather than a recording, and ask which clients will take a reference call — most credible firms have 1–2.
  • Pricing must separate build from ongoing run cost. A quote that blends them hides the recurring number you will pay for years.
  • Know the market bands before you negotiate: independents at $150–$350 an hour, boutiques at $300–$600, retainers commonly $3,000–$12,000 a month, scoped builds $5,000–$50,000.
  • Two or more red flags together — no scoping questions, unnamed delivery owner, unclear code or data ownership, hidden recurring costs — is usually a decline.
  • Grade the proposal against measurement, not enthusiasm: hours saved by staff at 57% and reduced vendor spend at 43% are how the market actually books AI value.
Table of six vetting questions for an AI marketing provider with the acceptable answer and the decline signal for each

Why this category is unusually hard to vet

The tooling is public, so capability claims cost nothing to make. One evaluation guide puts the problem bluntly: many firms wrap a model API with a prompt and call it agent development, and the way to test it is to make them walk through architecture — tool selection, error recovery, state handling. Another teardown offers a simpler behavioural tell: a real consultant leads with your problem, asking about workflow, pain points, volume of work and the cost of the current way before naming any technology.

Market conditions make the filter matter more. Adoption data shows 87% of marketers using generative AI in at least one workflow, while operations benchmarks show only 9% with fully automated journeys and 32% orchestrated. Almost every provider has used the tools; very few have taken a workflow to production and kept it there.

AskAcceptable answerDecline signal
Show one workflow you shippedNamed workflow with before-and-after numbers"We've helped clients implement…"
48-hour proof taskWorking link, edge cases handledScreenshots, mockups, a Figma frame
Who does the workNamed delivery owner and their hours"Our team", unnamed
Ownership of prompts, code, dataWritten: you own it, handover includedVague, or licence-back terms
Build vs run costTwo separate numbers, run cost annualisedOne blended monthly figure
A project that went wrongSpecific story with the correction madeCannot name one

The five questions that separate builders from talkers

Start with proof of work in their own house. One vetting framework leads with exactly that: show me one AI workflow you have shipped in your own business, with metrics. It marks "I've helped clients implement" and "I use ChatGPT for emails" as unacceptable answers. The logic is sound — anyone selling operational change should be able to demonstrate it on the one business whose data they fully control.

Second, set a short, cheap proof task. One buyer's guide describes a 48-hour test as the decisive filter: prompt shops return screenshots and mockups, real builders return a working link with edge cases already handled, plus talk credibly about monitoring and error handling. Third, demand a live demo. A 12-red-flag checklist recommends a live demo rather than a recorded video, asking which clients have given reference permission — most agencies have at least 1 or 2 willing to take a call — and checking engineering leads on public profiles for relevant prior work.

Fourth, ask how they define done before any build. A buyer's vetting checklist describes an acceptance-test-driven protocol published in mid-2026 that turns stakeholder goals into executable behavioural contracts and release gates before any prompt, model, retrieval or agent change is accepted — the order inverted from usual practice. A provider who cannot describe their acceptance gate will discover it on your budget. Fifth, ask what they would not automate, and why. The answer separates people who have watched a workflow fail from people who have only read about them.

Table of AI provider types with their 2026 price ranges and the main risk carried by each engagement shape

Red flags, weighted honestly

Not every flag is fatal. One selection guide lists the serious ones — unidentified delivery owners, unexplained data handling, hidden recurring costs, unverifiable proof — and adds a fair caveat: a short timeline, a paid discovery phase or a platform partnership is not automatically wrong; the inability to explain its fit is the problem. A pre-signature question set reaches the same conclusion from the other side: pitching before scoping the problem, no answer on who owns the code, an inability to describe a project that went wrong, and pricing that does not distinguish build from ongoing run cost. Any single one is a conversation; two or more together is usually a decline.

Two additional flags are specific to marketing work. The first is a provider who cannot name the metric they will move, in your units, at a date. The second is a provider who wants to start with tooling procurement. Both point at the same gap: no diagnosis. If your own baselines are missing, buy the diagnosis first — a fixed-scope readiness read or a marketing audit costs a fraction of a misdirected build and turns the shortlist conversation concrete.

Provider typeTypical 2026 priceBest fitMain risk
Independent practitioner$150–$350/hrOne workflow, deep buildCapacity and continuity
Boutique firm$300–$600/hrMulti-workflow programmeSenior sells, junior delivers
Retained advisory$3,000–$12,000/moOngoing governance and roadmapScope drift without a backlog
Scoped implementation$5,000–$50,000Defined build with acceptance testsHandover treated as optional
Large consultancy$500–$1,000+/hrEnterprise governance at scaleCost per unit of shipped work
Readiness assessment$1,000–$5,000Deciding what to buy at allEnds as a document, not a plan

Price the run cost, not just the build

Know the bands before you negotiate. A 2026 pricing breakdown puts independents at $150–$350 an hour, boutiques at $300–$600 and top-tier firms at $500–$1,000+, with fixed-fee projects starting near $10,000. A practitioner guide puts typical monthly engagements at $3,000–$12,000, with focused single-workflow rebuilds at the lower end. Quotes far outside those bands are not automatically wrong — they are a question to ask.

What matters more is the split. Every AI workflow has a build cost and a run cost: model or API usage, licences, monitoring, and the internal hours needed to review output. Ask for the run cost annualised, and check it against the saving. A workflow that recovers 4 hours a week per person across a team of 8 is worth roughly 128 hours a month; if the run cost consumes most of that, the build was interesting rather than useful.

Checklist graphic of a two-week shortlist process for choosing an AI marketing advisor, from brief to contract review

Judge the proposal on measurement

A good proposal names the baseline it will measure, the metric it will move, the date it can be read, and the acceptance test for done. Anchor it in how the market books value: Jasper's 2026 State of AI in Marketing reports value concentrated at the cost-reduction and execution layers, with hours saved by full-time employees the most common metric at 57% and reduced outsourced vendor or agency spend next at 43%, alongside AI use at 91% of marketing organisations.

Then check the proposal's honesty about timing. Efficiency signals on a narrow workflow can appear in weeks; programme payback more often lands at 90–180 days. A provider promising revenue lift inside a quarter from a content workflow is either measuring loosely or planning to. Where the reporting frame is the weak link rather than the AI work, fix it first — that is ordinary data intelligence work and it makes any provider's claims testable.

A buyer interviewing two consultants across a meeting table with a printed one-page brief between them

A shortlist process that takes two weeks

Week one: write a one-page brief with the workflow, its baseline numbers and the acceptance test you want. Send it to 3 providers. Score the replies on whether they asked scoping questions before pitching, named a delivery owner, and split build from run cost. Week two: give the top two the same 48-hour proof task, take a live demo, take one reference call each, and read the contract for ownership of prompts, code and data plus a handover clause.

Then start small deliberately. A scoped build with acceptance tests, documentation and admin access as named deliverables is a better first purchase than a retainer, because it produces evidence you can use for the next decision. Where you want a partner who treats the handover as part of delivery, that is how we structure our own marketing ops consulting work. To talk it through, get in touch, or read more on the Web Tonic blog.

What good looks like in the first month

The evidence you gather during selection should keep paying out after signature. In month one a strong provider produces the measured baseline you agreed, a ranked use-case list with effort and value against each candidate, and a written acceptance test for the first build. What they should not produce is a tool purchase order or a fifty-slide strategy document.

Watch two behaviours specifically. First, whether they ask to speak to the people who actually do the work being automated, rather than only to you — baselines gathered from managers are usually wrong by a wide margin. Second, whether they volunteer something they will not automate, and say why. Providers who scope narrowly in month one are the ones who still have a working system in month twelve, and with only 9% of teams reaching fully automated journeys, that durability is the scarce quality in this market.

If the first month does not produce those artefacts, raise it immediately rather than at the quarter's end. A misaligned scope is cheap to correct in week three and expensive to correct in week ten, and a provider's reaction to that conversation tells you more about the next twelve months than any reference call did.

Keep a written record of the selection itself. Note which questions each provider answered well, which they dodged, what the proof task returned, and what the reference call said. Six months into a build, that file is what tells you whether a problem is a normal delivery wobble or the thing you already saw in week two and chose to accept. It also makes the next selection faster, because the questions that actually predicted performance are visible rather than remembered, and the ones that produced polished but empty answers can be dropped from the brief entirely.

Frequently Asked Questions

What is the single most useful question to ask?

"Show me one AI workflow you shipped in your own business, with the before-and-after numbers." It cannot be answered with a case-study deck, and it filters most of the market in one exchange.

Should we pay for discovery?

Often yes. A paid discovery or readiness read at $1,000–$5,000 is not a red flag; refusing to explain why it is needed, or producing a document with no ranked use cases, is.

How many providers should we shortlist?

3 for the brief, 2 for the proof task. More than that and the comparison degrades into preference, because the proposals stop being answers to the same question.

Is a generalist marketing agency a valid choice for AI work?

If they can show shipped workflows with numbers and can articulate error handling and monitoring, yes — the discipline matters more than the label. Judge the evidence, not the category.

What contract clauses matter most?

Ownership of prompts, code and data; a handover pack with documentation and admin credentials as deliverables; the run cost stated annually; and an acceptance test that defines done. Those four remove most post-signature disputes.

Sources

How to evaluate an AI consultant, 12 red flags · How to vet an AI agency before you buy a prompt shop · How to vet an AI agency, red flags · How to choose an AI consultant · 12 questions to ask an AI consultant · AI implementation consultant vetting checklist · Evaluating AI agent development companies · Real AI consultants vs rebadged prompt engineers · AI consulting pricing 2026 · Consultant vs agency vs in-house · Jasper, The State of AI in Marketing 2026 · AI marketing statistics 2026 · AI marketing operations benchmarks 2026.

Author

Founder & CEO

Reviewer

Lead Client Success Manager

Summarize this article with AI

Book your strategy call today!
Schedule a call
Schedule a call
Discover our services
Our services
Our services

Blog

You may also like