Skip to main content
DoneThat

AI Adoption GuideMarketingMeasure

Natural-language analytics agent

An LLM answers ad-hoc performance questions over the warehouse without SQL, using tools like Tableau Pulse or Cortex Analyst.

Marketing processResearchPlanCreateLaunchMeasureReport

By Don, DoneThat’s AI coach · updated

What a natural-language analytics agent does

A natural-language analytics agent lets a marketing analyst ask performance questions in plain language and receive drafted answers grounded in governed metrics and approved campaign tables. Instead of writing SQL or rebuilding a dashboard for every ad-hoc question, the analyst describes the question, the time window, and the segments that matter. The agent proposes a query plan, retrieves the allowed data, and returns a draft answer with the filters, joins, and metric definitions it used.

The outcome is speed in the measure stage: faster turnaround on recurring and exploratory questions without inventing new KPI definitions. The agent does not replace the measurement stack. It sits on top of it, constrained to metrics, dimensions, and tables that the analytics team has already certified.

Typical questions include spend and return by channel for a campaign flight, conversion rate shifts after a creative swap, or whether a segment’s CPA moved relative to the prior period under the same attribution window. The value shows up when those questions would otherwise wait in a ticket queue or force an analyst to recreate logic that already exists in a semantic layer.

When this approach fits measurement work

This pattern fits teams that already have a stable metric layer and reliable campaign fact tables, but still spend too much analyst time on one-off pulls. If performance questions arrive constantly from brand, growth, and media partners, and most of them map to the same definitions of spend, reach, clicks, conversions, and revenue, a natural-language agent can shorten the path from question to draft answer.

It is a poor fit when metric definitions are still contested, when campaign identifiers are inconsistent across platforms, or when the warehouse tables the agent can see are incomplete. In those cases the agent will either refuse to answer or, if poorly constrained, invent a plausible story from partial data. Prefer fixing the data contract first.

The agent also fits better when questions are diagnostic rather than causal. Natural-language retrieval can surface that CPA rose in a region after a bid change. It cannot, by itself, prove that the bid change caused the rise. Causal claims still need experiment design or mix modeling, covered in related measure work such as Automated incrementality experiments and Channel-level marketing mix modeling.

How the agent works with governed metrics

The agent should operate against a curated semantic layer, not against raw platform exports. Governed metrics carry explicit definitions: numerator, denominator, attribution window, currency, and exclusion rules. Approved tables expose campaign, channel, creative, geo, and audience dimensions with documented grain. When a user asks for “ROAS last week,” the agent must resolve ROAS to the certified formula and apply the same week boundary the rest of reporting uses.

A reliable flow looks like this:

  1. The analyst asks a question in natural language, including the campaign, period, and comparison if needed.
  2. The agent maps terms to certified metrics and dimensions, or stops if a term is ambiguous or undefined.
  3. It drafts a query against accessible tables only, with explicit filters and joins.
  4. It returns a draft answer: numbers, breakdowns, and the definition and query assumptions behind them.
  5. The analyst reviews the draft, checks edge cases, and only then shares or acts on the result.

Empty output is required when metric definitions are missing or when the needed data is not in the accessible set. Silence is safer than a confident guess. If “brand lift” is not defined in the metric catalog, or if creative-attribute tables are not readable by the agent, the response should state that the question cannot be answered under current governance rather than approximating from another KPI.

For creative-level questions, the same rule applies: only use attribute fields that are populated and governed. Deeper creative-attribute work is covered in Creative-attribute performance attribution.

Human review before decisions

The agent drafts answers and queries. Analysts verify them before any budget, creative, or channel decision. Human-in-the-loop review is not optional overhead; it is the control that keeps speed from becoming silent error.

Review should cover four checks. First, definition match: did the agent use the intended metric, window, and currency? Second, population match: are test spend, brand campaigns, or excluded geos correctly in or out? Third, grain match: are rows at campaign, ad set, or creative level, and did rollups double-count? Fourth, comparison fairness: is the baseline period seasonally or mix-comparable?

Only after those checks should numbers move into a read-out, experiment brief, or budget reallocation. The agent can accelerate drafting of charts and narrative for stakeholders, but the analyst owns the decision-grade claim. If the draft conflicts with a certified dashboard for the same filters, stop and reconcile before publishing either number.

Access controls matter as much as metric governance. The agent should inherit the analyst’s permissions. It must not surface restricted brand or partner tables simply because a question was phrased cleverly. When permissions block a join, empty or partial output with an explicit access reason is the correct behavior.

What to put in place first

Start with a short, certified metric dictionary for the questions the team asks weekly: spend, impressions, clicks, conversions, CPA, ROAS, and any revenue definition used in marketing reporting. Document grain and keys for campaign and creative entities so joins do not silently fan out. Expose only the tables and views the agent is allowed to query, and log every drafted query with the user, timestamp, and metric IDs used.

Train analysts on how to ask well-scoped questions: named campaigns, explicit dates, and named metrics. Ambiguous prompts (“how did we do?”) should trigger clarifying questions from the agent, not a default national ROAS pull. Build a small library of verified question templates for recurring asks so the agent’s drafts stay aligned with how the team already thinks about performance.

Measure success as time-to-verified-answer and rate of empty or clarifying responses when definitions are missing, not as volume of unchecked answers. If empty responses are common, that is a signal to extend the metric catalog or fix table coverage, not to loosen constraints. Keep causal and mix-model workflows separate so the agent remains a fast, governed readout layer rather than an informal substitute for experiment design.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first