Simulated decision case · 06 / 07

When an AI Pilot Cannot Prove Whether It Worked

This simulated case follows an AI ad-optimisation pilot that looked positive early and negative later. Intake can show that the team cannot explain the reversal and that success criteria were never agreed. That does not prove the model failed. Stage 2 would inspect the test design, control group, configuration changes, market movement and guardrail metrics. The directional decision is FIX: pause, establish experiment governance and run one controlled re-test with pre-registered thresholds and a stop rule. If a clean re-run cannot be secured on acceptable terms, the default becomes STOP as “unproven”, not “failed”.

What does this case show? Without a protected control, pre-agreed thresholds and a single owner, neither positive nor negative pilot results support a defensible AI decision.

01 · Organisation and value

Business context and management question

Organisation and value proposition

A mid-sized free-to-play mobile game studio earns most of its revenue from advertising. Ad frequency is a sensitive lever. More ads can lift short-term revenue but may damage retention, ratings and lifetime value. A vendor offers machine learning to set individualised ad-frequency caps.

Management question

A paid pilot began with an apparent revenue uplift and ended below baseline. Should the company continue, restructure or terminate it?

Relevant value-chain context

The model changes the ad-configuration layer, but performance signals originate upstream in ad demand and infrastructure, and downstream in player behaviour. The full chain must be instrumented to interpret the result.

02 · Approaching the real problem

From symptom to a testable problem hypothesis

Early signals visible through intake

The following items may be reported during the initial intake or interviews. They are not treated as measured performance values.

  • Early pilot dashboards show an apparent uplift, while later results show a decline.
  • The team cannot explain the reversal.
  • Ad configuration changed during the pilot.
  • Market-wide ad prices moved during the same period.
  • No locked holdout or written stop rule was agreed before launch.
  • No single owner is accountable for revenue attribution across teams.

Priority problem hypothesis

Initial intake, interviews and limited document signals may make the following hypothesis a priority. It is not presented as a confirmed root cause until it is tested against operational evidence.

The working hypothesis is that the organisation cannot yet run a trustworthy experiment on this revenue lever. The visible symptom is a pilot that appears to have lost money. The deeper issue may be a contaminated test, no pre-committed success criteria and fragmented ownership. On that evidence, the result is not proof that the AI failed. It is proof that the organisation cannot yet tell whether it worked.

Customer and internal value to be validated

The following are value hypotheses, not achieved or validated outcomes.

Customer value: Players experience ad frequency directly. A governed system should protect retention, stability and perceived fairness while testing revenue value.

Internal value: A clean decision avoids scaling an unproven claim and avoids killing a potentially useful lever. It also strengthens vendor accountability and future experiment quality.

03 · AI alignment

From the first AI reflex to a value-aligned design

01 · First reflex

Proposed technology

Scale from the early uplift or stop from the final decline.

02 · Alignment gap

Why should it be reconsidered?

  • No protected holdout or change freeze.
  • Success and guardrail thresholds were not registered before the test.
  • Configuration changes and market movement contaminate attribution.
  • No single source joins revenue, configuration history and player cohorts.
  • No named owner has authority over experiment integrity.
03 · Priority direction

Controlled AI Ad-Frequency Re-Test

The AI role, human decision boundary, data readiness and measurement logic are designed together. No scale decision is made without pilot evidence.

Public boundary: This page shows the decision logic. It does not publish MoreSight’s detailed question architecture, scoring rules or project analysis templates.
04 · Opportunity map

AI options derived from the real problem

This table is not an investment decision or a definitive ranking. Priority and evidence readiness are reassessed during Stage 2.

Opportunity hypothesisAI roleValue connectionEvidence readinessDirectional priority
Experiment-governance standardExplain + GovernMake future AI results decidableLeadership ownership and process change requiredHighest; prerequisite
Causal pilot monitoring and change logDetect + ExplainSeparate model effect from other changesJoined data and frozen control requiredHigh; re-test foundation
Segment-level ad-cap optimisationPredict + RecommendRevenue with player-experience guardrailsClean experiment and raw-data access requiredConditional; re-test only
Portfolio-wide automatic rolloutExecuteScale revenue optimisationCurrent evidence is contaminatedNot recommended

Priority initiative to test if evidence supports it

The design below is not a solution commitment. It is a pilot hypothesis to be validated.

Controlled AI Ad-Frequency Re-Test

  • Business decision: Does the AI-driven ad cap improve revenue against a locked control without harming retention, stability or ratings?
  • Primary users: Monetisation, analytics, product, engineering and the executive experiment owner.
  • Inputs: Treatment and holdout cohorts, revenue and ad metrics, configuration-change log, eCPM and fill-rate context, retention, stability and rating data.
  • AI functions: Monitor experiment integrity, attribute differences, flag contamination and evaluate pre-registered decision thresholds.
  • Outputs: Clean treatment-versus-control result, guardrail status, contamination log and one-page scale-or-stop memo.
  • Human decision boundary: The AI or vendor does not define success after seeing the result. Management agrees the thresholds in advance and the experiment owner controls changes during the test.
  • Measurement logic: Evaluate revenue uplift, day-7 retention, rating, stability, configuration compliance and statistical uncertainty together.
05 · Evidence and claim boundary

What can intake show, and which claims require data?

Areas the intake can direct

  • Reports of conflicting pilot results and unexplained reversal.
  • A signal that the experiment changed while it was running.
  • A hypothesis that the decision problem is experiment governance before model performance.
  • An early warning that management may be treating contaminated evidence as a scale-or-stop result.

Evidence required for Stage 2

  • Pilot design, cohort definition and vendor dashboard exports.
  • Raw or independently reconcilable revenue and ad reports.
  • Configuration and code-change history during the test.
  • Market eCPM, fill-rate and infrastructure context.
  • Retention, stability, rating and other guardrail metrics.
  • Commercial terms and raw-data access rights.

Analyses possible when evidence is available

  • Reconstruct the experiment timeline and identify contamination events.
  • Compare treatment and control in the same markets and time window.
  • Test whether the reversal coincides with configuration or market changes.
  • Evaluate guardrail metrics and confidence intervals.
  • Assess whether vendor claims can be reconciled independently.

Claims we will not make without evidence

  • That the model failed or succeeded from the existing contaminated result.
  • That early uplift should be annualised.
  • That player harm did or did not occur without guardrail data.
  • That the initiative deserves scale before a controlled re-test.

Readiness gaps and risks

The following are readiness hypotheses that require document and process review.

  • No mandatory experiment template with pre-registered thresholds.
  • No change-freeze protocol during live tests.
  • Fragmented ownership across ad operations, analytics and engineering.
  • Limited raw-data access or vendor transparency.
  • Infrastructure issues can be misattributed to the AI layer.

Decision language by evidence level

  • Stage 1: Provisional FIX. Treat the pilot as unproven and pause any scale decision.
  • Stage 2: If the original test is materially contaminated, issue an evidence-supported FIX and define one controlled re-test.
  • After the re-test: SCALE if the pre-registered threshold is met and guardrails hold. STOP if it is not, or if a clean re-test cannot be secured.
06 · Validation plan

Move to the next defensible decision in 90 days

This plan is adapted to data access and client scope. Stage 1 alone does not include implementation or outcome validation.

PeriodPurposeMain actions
Days 1–30Governance and baselineName one experiment owner, write the holdout and change-freeze standard, define success and stop thresholds, and complete infrastructure hygiene checks.
Days 31–60Controlled re-testNegotiate the re-run, freeze unrelated changes, run treatment versus a locked holdout in the same markets and window.
Days 61–90Decision gateEvaluate only against the pre-registered thresholds and produce a one-page Scale or Stop memo for management.

Recommended validation measures

  • Revenue uplift versus holdout over the agreed test window.
  • Day-7 retention and store-rating guardrails.
  • Crash and stability metrics.
  • Percentage of configuration changes logged and authorised.
  • Ability to reconcile vendor and internal reporting.
07 · Decision gate

When should management Scale, Iterate or Stop?

Scale

The treatment meets the pre-registered revenue threshold and no player-experience or stability guardrail is breached.

Iterate

The direction is positive, but uncertainty, sample size or implementation quality prevents a scale decision.

Stop

The threshold is missed in a clean test, or the organisation cannot secure an interpretable re-test on acceptable terms.

Decisions management must make

  1. Who has authority to freeze changes during the test?
  2. What revenue threshold and guardrails define success before launch?
  3. What is the maximum additional spend for a re-test?
  4. Will the experiment standard become mandatory for all revenue-affecting AI pilots?

Transferable lesson: Before asking whether an AI initiative works, ask whether the organisation would be able to tell.

Disclosure: This is a simulated composite case. Organisation characteristics, circumstances, data and outcomes have been altered, combined or synthetically generated. It does not represent a specific organisation or client engagement. Numerical examples are not achieved client results.
Your initiative

Which decision does your AI project deserve?

Assess the available evidence, business value and implementation conditions together.

Start a Business Value Review