Direct answer: Design an incrementality test for intent-based advertising by pre-registering the business decision, eligible population, experimental unit, treatment, counterfactual, primary outcome, conversion lag, sample and power assumptions, contamination controls, analysis and stopping rule. Intent data may define eligibility or planned strata; it is not evidence that the ads caused pipeline.

Who is this for? Marketing analytics leaders, RevOps teams, paid-media directors and agencies deciding whether to launch, retain or expand intent-driven ad spend. The guide also shows when a valid test is too small to answer the decision and should be reported as inconclusive.

Start with the business decision and counterfactual

Write the decision before the method. Examples are: expand the intent-based campaign to all eligible accounts; renew the data and media program; shift budget from ICP-only targeting; or stop a creative and audience combination. Each decision needs a threshold for action, the cost of a wrong decision, and the period over which the effect matters.

Then define the estimand – the exact causal quantity. A useful form is: ‘Among eligible high-fit accounts with recent research in approved topics, what is the incremental probability of a sales-accepted opportunity within 120 days from offering the declared advertising treatment, compared with the existing non-ad treatment?’ This states population, treatment, comparison, outcome and window. ‘Did intent ads work?’ does not.

The counterfactual is what would have happened to the same eligible population without the additional treatment. Attribution reports do not observe that state. A randomized holdout approximates it by making treatment assignment independent of the potential outcome. A matched comparison tries to approximate it with assumptions. A pre/post comparison assumes other changes did not matter. Make the assumption visible.

Choose one primary outcome close enough to the revenue decision and frequent enough to measure. Won revenue may be economically ideal but too rare or delayed; qualified meetings or sales-accepted opportunities may be usable intermediates when their definition is stable. Add guardrails such as unsubscribes, complaint rate, cost, sales capacity and low-quality conversions.

Define data, owners, windows, and workflow before launch

Freeze the protocol before launch. Record the hypothesis, eligible and excluded units, assignment and stratification, treatment, control experience, budget, start and end, lag, primary and secondary outcomes, minimum detectable effect, power target, data sources, missing-data handling, analysis, multiplicity rules, stopping conditions, owners and approval. Store a versioned copy that cannot be silently rewritten after results appear.

Build a unit ledger from raw intent and fit evidence. Preserve account, domain and person candidates; source; topic; observed and received time; freshness; identity confidence; customer and opportunity status; suppression; geography; and planned stratum. Remove ineligible units before randomization. If eligibility is updated after treatment based on engagement, the analysis becomes selected on a post-treatment variable.

Name owners for protocol, media delivery, audience construction, CRM outcome, data engineering, privacy or legal escalation, analysis and client communication. Separate the analyst who monitors data integrity from the executive who wants an early win. Establish a deviation process: a broken feed or policy issue can justify a change, but the change must be logged and its effect assessed.

Run preflight checks: expected eligible units, baseline outcome, intraclass correlation if randomizing groups, likely match and delivery, conversion lag, required duration, sales capacity, contamination paths and feasibility. Google’s experiments guidance notes that interference, power and inconclusive results matter. A calendar deadline is not a substitute for statistical information.

Compare five experiment and measurement methods

The five methods below are ordered from stronger direct control of the counterfactual to a diagnostic that does not establish incrementality. The best choice depends on the available assignment mechanism, unit independence, sample, data rights and business decision. No method earns causal language merely because software labels a report ‘lift.’

1. Randomized person or account holdout

Best fit: Use when the platform and data architecture can assign eligible people or accounts to treatment and control before exposure, the unit can be kept reasonably independent, and both groups can be measured through the declared outcome window. Account randomization is often more defensible than person randomization for B2B programs because several people at one company may influence the same opportunity.

Experimental unit: Choose person, cookie, account, buying group or another unit based on how treatment is delivered and how the outcome occurs. Randomize at the highest level needed to limit spillover. If one account’s employees can enter both groups, person-level assignment may contaminate the result.

Required data: A frozen eligible population, stable intent and fit strata, assignment record, exposure and campaign receipts, suppression list, cost, primary and guardrail outcomes, conversion lag, CRM reconciliation, and a rule for accounts entering or leaving during the test.

Implementation: Pre-register the hypothesis, minimum detectable effect, allocation, exclusions, duration and stopping rule; stratify or block on a few important pre-treatment factors; randomize; verify balance; prevent cross-group activation; monitor data quality without peeking for a winner; and analyze all assigned units under the declared rule.

Primary analysis: Estimate the difference in outcome rate or value between assigned treatment and control, with uncertainty and cost. Use intent strata to evaluate planned heterogeneity only when the sample can support it. Report intent-to-treat as the default so selective exposure does not silently redefine the groups.

Proof to request: Preserve the randomization code or platform receipt, unit list, balance checks, treatment integrity, excluded records, outcome definitions, missing-data handling, analysis code, confidence interval and deviation log. A dashboard screenshot is insufficient.

Meaningful limitation: Small B2B populations, account overlap, sales contact outside the ad program, shared devices and long conversion lag can reduce power or contaminate groups. A valid design can still return an inconclusive result; randomization does not guarantee a decisive answer.

2. Randomized geo experiment

Best fit: Use when individual or account assignment is unavailable but media can be varied across geographic units with limited spillover, adequate historical stability and enough comparable regions. This is common when the channel or privacy design supports geography more readily than user-level holdout.

Experimental unit: The unit is a market, state, city group, designated market area or another non-overlapping geography. Pair or stratify geographies using pre-treatment outcome and spend patterns, size, vertical mix and other declared factors before random assignment.

Required data: Consistent geo-level media cost and outcomes, sufficient pre-period history, the treatment schedule, boundaries, account or revenue geography rules, conversion lag, cross-border or travel indicators, and a definition for national accounts whose outcomes span several regions.

Implementation: Select eligible geographies, test pre-period comparability, form pairs or blocks, randomize treatment, hold other material activity stable where practical, run for the declared duration, monitor delivery and spillover, and use a model specified before results are inspected.

Primary analysis: Estimate incremental outcome and incremental return on ad spend across randomized geo pairs or strata. Google Research has published geo-experiment methods and work on paired-geo incremental return; the appropriate model still depends on the buyer’s data.

Proof to request: Request the geo eligibility rules, pairing variables, assignment record, pre-period fit, delivery and spillover checks, outcome mapping, model specification, uncertainty, sensitivity analysis and a complete list of concurrent geo-specific campaigns or sales changes.

Meaningful limitation: B2B markets may be too heterogeneous or too few, and company headquarters may not represent where buying activity occurs. National campaigns, sales territories and cross-geo exposure can weaken separation. A poorly matched geo test can be more misleading than a clearly labeled observational analysis.

3. Matched-market quasi-experiment

Best fit: Use only when randomization is genuinely unavailable and the team has several untreated markets or account cohorts with credible pre-treatment similarity. It can inform a decision, but it is a quasi-experiment and should not be described as equivalent to random assignment.

Experimental unit: Use geographies, accounts, industries or time-series panels that can be matched using pre-treatment variables unaffected by the campaign. Do not match on post-treatment engagement, conversions or intent changes because that conditions on consequences of treatment.

Required data: Long enough pre-period history, treatment timing, outcomes, costs, fit and baseline intent, concurrent campaigns, sales coverage, market changes, and a documented reason each comparison remained untreated. Preserve the candidate pool before selecting the best-looking matches.

Implementation: Predeclare matching variables and tolerances; build the comparison without viewing post-period results; inspect pre-trends; set the analysis and placebo tests; track concurrent changes; and run sensitivity analyses with alternate credible matches. Treat unexplained pre-trend divergence as a stop signal.

Primary analysis: Estimate the difference in change between treated and matched comparison units, with uncertainty and sensitivity to the match. Synthetic-control or regression adjustments may help in appropriate settings, but model complexity cannot remove unobserved confounding by assertion.

Proof to request: Ask for the full candidate pool, matching code, balance and pre-trend plots, excluded units and reasons, concurrent-event log, placebo and sensitivity tests, model outputs, confidence intervals and the exact claim language proposed for stakeholders.

Meaningful limitation: Selection bias remains: treated markets or accounts may differ in unobserved ways, and choosing comparisons after seeing results can manufacture a lift. Use causal language cautiously and explain why randomization was not feasible.

4. Platform conversion-lift study

Best fit: Use when an ad platform offers a conversion-lift study for the account, campaign, region and outcome, and its feasibility check indicates enough scale. Platform studies can provide a randomized holdout within the platform’s delivery system with less custom engineering.

Experimental unit: The platform determines eligible users or another supported unit and assigns a control that is withheld from the experimental campaign. Confirm how identity, cross-device exposure, other campaigns and conversions are handled before accepting the result.

Required data: Eligible campaigns, sufficient expected conversions, reliable conversion tracking or qualified offline events, study dates, budget, audience and geographic scope, conversion lag, and an account-level record of other activity that can reach the same users.

Implementation: Use the platform’s feasibility process, freeze the primary outcome and study scope, keep instrumentation stable, avoid conflicting experiments, run for the recommended duration, and export all available study definitions and uncertainty. Google’s Conversion Lift overview distinguishes attributed from incremental conversions, while its setup guidance describes feasibility and study considerations.

Primary analysis: Report the platform’s incremental conversions or conversion lift, uncertainty and study population, then connect the conversion definition to qualified pipeline. Do not automatically generalize a platform-specific result to every channel, audience or future period.

Proof to request: Retain the official study setup, feasibility result, eligibility and control definitions, campaign scope, outcome, dates, budget, other-campaign conflicts, result, confidence and platform documentation. Reconcile the measured event with CRM quality.

Meaningful limitation: Eligibility and minimum scale can exclude many B2B programs. The platform controls parts of assignment and reporting, cross-channel exposure may remain, and the measured conversion may occur well before pipeline. The study answers a bounded question, not the entire go-to-market impact.

5. Observational attribution diagnostic

Best fit: Use as an operational diagnostic when experimentation is unavailable, when a team needs journey context, or when deciding whether a future experiment is worth funding. Attribution can describe observed paths and allocation under a model; it should not be called incrementality.

Experimental unit: The records are observed people, accounts, sessions, opportunities or touches rather than randomized units. Define identity resolution and attribution rules explicitly, because changing either can change the result without changing buyer behavior.

Required data: Timestamped exposures and interactions, identity and account confidence, campaign cost, CRM stages, outcome values, attribution window, model definition, missing-channel assessment, topic and fit cohorts, and sales activity. Preserve untracked exposure and unknown identity as limitations.

Implementation: Reconcile sources, deduplicate events, freeze attribution definitions, analyze cohorts and lag, compare with historical or matched baselines, and label the result observational. Use the findings to identify plausible mechanisms, data gaps and candidate units for a later experiment.

Primary analysis: Report observed conversion and pipeline rates, attributed credit under the named model, account journeys and sensitivity to alternate windows or rules. Avoid converting attributed revenue into incremental revenue. Correlation can prioritize action without proving causation.

Proof to request: Request event lineage, identity logic, missing-channel analysis, attribution rules, alternate-model sensitivity, outcome reconciliation, cohort definitions and a visible caveat explaining why the comparison does not establish a counterfactual.

Meaningful limitation: Selection, unobserved demand, sales effort, seasonality and identity gaps can explain apparent lift. A sophisticated dashboard does not solve the missing counterfactual and must not be used to overstate causality in a client renewal.

Separate causal experiments from attribution and observation

Randomized experiments estimate an effect under the declared assignment and implementation. Quasi-experiments estimate an effect under additional assumptions about comparability and trends. Attribution distributes observed outcome credit among recorded touches under a rule or model. Observational cohort analysis describes association. All can be useful; they answer different questions.

Use attribution to debug paths and reconcile data, not to replace the counterfactual. Use observational differences to size a hypothesis or find segments, not to claim lift. Use matched markets when randomization is impossible and the assumptions survive pre-trend and sensitivity checks. Use randomized account or geo designs when exposure can be separated and the sample supports the decision.

A common failure is method switching after the result. A randomized test is underpowered, so the report substitutes an attributed pipeline number and calls the program successful. Keep the planned primary result, including an inconclusive interval, and present secondary diagnostic evidence as secondary. Decision-makers deserve to know what was and was not learned.

Budget sample, media, tooling, analyst time, and opportunity cost

Budget five resources. Sample is the number of independent units and expected events. Media is the spend needed to deliver meaningful treatment. Data and tooling covers eligible audiences, assignment, exposure, outcome and analysis. Analyst and engineering time covers protocol, QA, reconciliation and sensitivity. Opportunity cost includes withheld advertising, delayed rollout and sales capacity.

Power planning begins with the baseline outcome rate or value, smallest effect worth acting on, variation, clustering, allocation, expected delivery and attrition. A sample-size calculator can produce a number only after those inputs are defensible. For account-randomized B2B tests, several people share one outcome, so counting cookies as independent units exaggerates information.

Do not increase spend solely to obtain significance if the incremental opportunity could not repay the test. Conversely, a cheap underpowered test can waste all of its budget by producing an interval too wide for the decision. Report the range of plausible effects and whether each side would change the decision.

Agency pricing should state the protocol, data preparation, test administration, media operations, analysis, revision and client meeting separately. BrandWell agency plans range from $2,500 to $5,000 per month, depending on topic count, term, and available contractually scoped topic exclusivity. The current written quote and Order Form control. Pricing and product details need current approval.

Choose metrics, segments, confidence checks, and reporting rules

The primary metric should match the decision. Common choices are incremental qualified conversion rate, incremental opportunities per eligible account, incremental pipeline value or incremental return on ad spend. Secondary metrics explain mechanism: reach, exposure, clicks, engaged visits, meetings and stage movement. Guardrails capture cost, complaints, low-quality volume, sales workload and customer suppression failures.

Plan a small number of segments before launch: fit tier, intent recency, topic cluster, existing engagement or account size. Randomize within important strata when practical. Do not search dozens of slices for one favorable result. Adjust for multiple comparisons or label segment results exploratory and use them to design the next test.

Confidence reporting should include the point estimate, interval, sample and event counts, assignment balance, treatment delivery, missing data, contamination and practical decision threshold. Statistical significance is not the same as commercial importance. An estimate can be statistically detectable but too small to fund, or commercially promising but too uncertain to expand.

Use the predeclared duration and stopping rule. Stop early for safety, policy or severe implementation failure; otherwise avoid weekly winner checks. If a sequential design is used, specify its monitoring boundaries in advance. Add conversion lag after the final exposure before closing the outcome window.

Know when the available scale cannot answer the question

The test is decision-useful when it has enough independent units and expected outcomes, assignment can be maintained, treatment differs meaningfully, instrumentation is stable, outcome lag fits the planning horizon, and a plausible result would change the budget decision. Run the power analysis before launch, not after the test misses significance.

It is too small when the minimum detectable effect is larger than any believable business benefit, when only a handful of accounts or opportunities determine the result, or when geographies cannot be separated. It is also infeasible when nearly every eligible account is already exposed by other campaigns, when sales changes treatment based on group status, or when the qualified outcome arrives after the renewal decision.

Options for insufficient scale are to aggregate carefully across time or clients only when definitions and rights permit; use a more frequent, still meaningful intermediate outcome; test a larger eligible universe; evaluate creative or delivery mechanics instead of pipeline; or document that the question cannot currently be answered. Do not weaken the outcome to a click merely to obtain a favorable result.

Use intent, identity, and activation data without treating them as proof

Intent can support the design in four legitimate ways: define an eligible research cohort before assignment; stratify randomization by topic or recency; create treatment messages aligned with declared needs; and measure whether effects differ across a small number of planned cohorts. Fit can restrict the population to accounts able to buy. Identity data can connect outcomes when permitted. Activation records show whether assigned treatment was deliverable.

None of those inputs proves intent or exposure. Offsite topic activity may be account-level and probabilistic. Person resolution may be wrong. An uploaded audience may have a low match rate. An assigned unit may never receive an impression. Preserve the full chain: raw evidence, eligibility decision, assignment, audience receipt, exposure where available, qualified outcome and analysis inclusion.

Analyze intent-to-treat as the main randomized estimate: compare units as assigned. A treatment-on-the-treated estimate asks a different question and can be biased if exposure is selective. If shown, use an appropriate method and explain its assumptions. Do not simply discard treatment accounts that received no impression and compare only the easiest-to-reach users.

BrandWell can support agencies with topic reports, cohort preparation and agent-ready operating instructions, subject to product review. It is not an experimentation platform and does not prove incremental impact. Claude or ChatGPT can execute approved analytical steps; Moxby is a separate optional browser path. Human analysts must validate data and claims.

Control selection, contamination, privacy, and overclaiming risk

Selection bias appears when high-intent or high-fit accounts receive treatment and are compared with an unmatched general population. Avoid it by randomizing within the eligible universe or using a defensible pre-treatment comparison. Do not condition on post-treatment clicks, visits, seller actions or opportunity creation.

Contamination occurs when control accounts see other campaigns, treatment and control employees share an account, sales contacts the groups differently, geographies overlap, or audiences are reused elsewhere. Map every path before launch, measure known spillover and use cluster assignment when it reduces interference. Some contamination attenuates a true effect; selective contamination can bias in either direction.

Privacy controls include data minimization, source and rights review, purpose limitation, client separation, access, retention, deletion, suppression and clear treatment of identity uncertainty. The experiment protocol should never encourage an upload that violates platform rules. Avoid sensitive topics and do not reveal inferred behavior in creative or outreach.

Overclaiming includes calling attributed revenue incremental, generalizing beyond the eligible population, hiding an inconclusive interval, omitting concurrent sales changes, and reporting subgroup winners discovered after the fact. Use an evidence ladder: randomized causal estimate; quasi-experimental estimate with assumptions; observational association; operational metric. Label each accurately.

Turn experiment evidence into agency decisions and renewal reporting

An agency should incorporate experiment design as a decision service: quarterly question selection, feasibility and power review, protocol and approvals, audience and data QA, test administration, deviation logging, qualified-outcome reconciliation, analysis, and an executive decision memo. It should not promise a positive lift; it should promise a transparent process within the written scope.

The report opens with the decision and primary result, including uncertainty. It then shows implementation integrity, cohort and exposure, guardrails, secondary diagnostics, limitations, financial scenarios and the recommended action. Renewal can be supported by credible evidence of incremental value, but an inconclusive test may recommend improving scale or instrumentation rather than renewing unchanged.

For an agency using BrandWell, the intent-data offer remains distinct from the legacy SEO writer. Agencies can purchase BrandWell’s $70 seven-day reseller pilot. It includes agency-branded topic reports and the complete sales playbook under the current written pilot terms. Other product capabilities and any topic exclusivity remain subject to their separate current written scope. Those product statements require current verification and do not alter the experiment’s causal standard.

The most valuable experiment is not the one with the largest reported lift. It is the one that gives the client a defensible next decision, records uncertainty and prevents an attractive correlation from becoming an expensive recurring assumption.

Use the $70 pilot to test client demand

BrandWell’s agency entry point is a $70 reseller pilot that lasts seven days. The pilot includes topic reports with the agency’s branding plus the complete sales playbook for positioning the service, approaching suitable clients, and seeking commitments before a full-plan decision.

That sequence helps the agency test demand and determine whether expected commitments support the cost structure and a potential profit center. BrandWell does not guarantee commitments, cost coverage, or profit. Review the $70 seven-day reseller pilot.