marketing​context​graph
Causal AI × Agentic AI

Causal evidence,
stored properly.

Marketing measurement produces causal effects. Almost nothing preserves them — their uncertainty, their expiry, the population they hold for, the method that established them, the decision they justified. The Marketing Context Graph is a governed substrate that does, so causal evidence can be composed, audited, transported, and refused.


Two problems, one cause

Measurement produces evidence. Then the evidence goes nowhere.

01 — Causal AI · identification

Organisations centralise petabytes of signal into flat, columnar warehouses on the assumption that machine learning will surface the causal relationships inside. It won't. Many of the hardest measurement problems are structural before they are statistical — no amount of estimation care recovers an effect the representation cannot express.

A channel's worth for a budget decision is its total effect: its direct path plus every indirect path through the channels it influences. Recovering that requires composing paths through causal structure. A flat table has none. And a model that conditions on a mediator — as media-mix models routinely do by including branded search — reports only the direct path and makes the indirect one disappear.

Not just MMM The same failure afflicts touch-based attribution, causal attribution, and observational measurement generally. Anything that must recover a causal effect from marketing data and carry it forward to a decision.
FLAT REPRESENTATION TV SOCIAL DISPLAY SEARCH 41,20028,900 12,40019,750 38,70031,050 11,90021,300 44,10027,600 13,05018,900 39,80030,400 12,75020,600 42,35029,200 12,10022,150 four independent columns no relationship is represented, so none can be composed conditioning on SEARCH deletes the indirect paths entirely CAUSAL STRUCTURE TV SOCIAL DISPLAY BRANDED SEARCH SALES halo solid = mediated path  ·  dashed = direct path total effect = direct + indirect, recovered by traversal
Fig 1 The same marketing system under two representations. On the left, spend is four independent columns — the relationships that carry TV's value through branded search are simply not in the data model, so no estimator can recover them. On the right the structure is explicit and each channel's total effect is a composition over paths.
02 — Agentic AI · context

Agents are now asked which campaigns should I scale? and optimise my mix. These are not language problems. Answering them requires incremental effect and marginal return, saturation and interaction, mediation and halo, short- and long-run response, the uncertainty and strength of evidence behind each effect, the population it applies to, the constraints that bind, and how all of it moves over time.

Conventional agent memory — vector stores, memory banks, entity caches — stores flat facts or loosely connected entity graphs. It helps an agent remember. It cannot help it explain.

The retrieval gap An agent that retrieves the sentence "TikTok ROAS is 3.4×" has retrieved a number, not a basis for action. It does not know whether that was measured experimentally or observationally, whether it still holds, whether it holds here, or at what spend it saturates.

What the substrate stores

An effect is not a number. It is a number with conditions attached.

The MCG does not solve causal identification. It preserves, governs and operationalises the evidence that identification produced.

Formally the graph is a tuple G = (E, R, γE, γR) — typed entities, typed directed relations, and a context map for each. Everything that makes an effect safe to reuse lives in that context: not a separate node type, not a side table, not tribal knowledge in an analyst's head.

That is the whole idea. The diagram below is one edge.

Position A governance argument, not a new estimator. The measurement methods are classical — mediation and path analysis, informative structural priors, transport. The substrate's job is to carry what those methods produce, not to replace them.
Why it has to be attached A stored effect detached from its method, window and population is indistinguishable from a guess with good posture. The conditions are not metadata — they are what makes the number actionable or not.
TV BRANDED SEARCH HALO_EFFECT_ON γ_R — RELATION CONTEXT EFFECT φ = 0.34 of branded-search response is caused by TV UNCERTAINTY 95% CI [0.28, 0.40] carried, not discarded at write time METHOD randomised geo experiment the best single predictor of trust VALID 2025-04-01 → 2026-03-31 expires; supersession is explicit POPULATION US · 42 DMAs · Q2 mix bounds where it may be reused
Fig 2 One governed edge. The relation carries the effect, but the relation context carries everything that determines whether the effect may be used — how confident to be, what established it, when it stops being true, and for whom it was ever true. A conventional warehouse stores the first row and throws away the other four.

Two clocks

What was true, and what you knew.

Every fact carries a valid time — when it held in the world — and a transaction time — when the system came to believe it. With both, the graph can be replayed exactly as it was known on any past date. The snapshot operator Σ(tv, tr) is what makes that mechanical.

This sounds like bookkeeping. It is the difference between a backtest that means something and one that quietly grades itself with tomorrow's answer key.

Why a single clock fails A store with only valid time overwrites the past when a correction arrives. Ask it how good your Q1 estimate was and it reports today's corrected number as though you had always known it.
TRUTH 2.0 Q1Q2 Q3Q4 Q5 corrective geo-experiment BELIEVED at the time 3.43.4 3.43.4 2.0 FLAT STORE one clock 2.02.0 2.02.0 2.0 reports 0% error MCG Σ(tv,tr) two clocks 3.43.4 3.43.4 2.0 reports 44% error
Fig 3 The same history, queried two ways. A flat store returns today's corrected value for every past quarter, so its backtest reports that the model was always right. The bitemporal snapshot returns what was actually believed — exposing both the real historical error and the 70% over-investment made while the belief was wrong.

Four studies

Every number comes with the condition under which it stops being true.

Four decision-centric studies across 8,000 simulated budget decisions, all scored against known oracles. The scope conditions are in the margin, and they are the most useful part.

Study 1 — path composition

Composing total effects instead of conditioning on mediators

43% less budget-decision regret

At full cross-channel mediation. The advantage grows monotonically with how mediated the system is — and is exactly zero when channels are independent.

Study 2 — bitemporal governance
0% vs 44% historical error

What a flat store reports, against what actually happened. The gap concealed a 70% over-investment across five quarters.

Study 3 — decision provenance
6 of 7 query classes flat recall cannot answer

Multi-hop evidence chains, policy compliance, rejected alternatives, as-of temporal lookups, policy-drift audits, trace audits. Not answered badly — not representable.

Study 4 — transport
≤ 0.05 vs 2.0 RMSE as populations diverge

Reweighting a source effect by the target's own composition, against applying the source headline number directly.

Standing caveat Every result here is simulated against a known oracle. These are existence proofs that isolate a mechanism. They are not estimates of what you would see on your own data.
Limit · Study 1 The gain comes from encoding causal structure, not from storing it in a graph. A hand-specified structural equation model recovers the same total effects with no graph involved. What the substrate adds is that the structure is versioned and reusable rather than re-specified per analysis.
Limit · Study 2 Regime-change dates and validity windows are treated as known. A deployment would have to detect them, and detection error would erode the margin.
Limit · Study 3 The baseline is a recall-oriented store — which is what most agent memory actually is. A well-modelled bitemporal relational schema would answer most of these too. The finding is about flat recall, not about graphs beating databases.
Limit · Study 4 Transport's advantage holds only when the target has poor spend overlap. Give the target rich variation of its own and plain back-door adjustment competes with transport and then beats it.
4%6% 8%10% 0%25%50% 75%100% CROSS-CHANNEL MEDIATION → BUDGET-DECISION REGRET flat model composed paths identical here — nothing to compose
Fig 4 The left edge is the point. With no cross-channel mediation there are no paths to compose and the two systems are exactly the same. The advantage is created by causal structure in the world, not assumed by the method — which also means a business whose channels genuinely don't interact gains nothing here.
DECISION D1 shift 20% ($50K) Meta → TikTok 2025-10-06 APPLIED_POLICY BudgetPolicy v2.1 — as believed APPROVED_BY VP Marketing — required >$25K PRECEDENT_FROM D0 — Q3 TikTok test SUPPORTED_BY geo-lift 1.8× · p = 0.02 SUPPORTED_BY MMM 3.4× ROAS · CPM −15% CONSIDERED_REJECTED 30% ($75K) — creative coverage one traversal · why, under what policy, on what evidence, against what alternative
Fig 5 The reconstructed answer to why. Because the policy reference resolves through the snapshot operator, the decision is bound to the policy version believed on the day it was made — not the version current when someone asks. A flat store can return any one of these boxes. It cannot return the structure that connects them.

Across populations

An effect measured there is not automatically an effect here.

A geo-lift experiment identifies an effect in the population it ran in — a set of regions, a season, an audience mix. Budget decisions are routinely made for a different one. Moving the number across is not a copy; it is the transportability problem.

The graph carries a selection diagram: an annotation marking exactly what differs between the two populations. When the annotation licenses the transfer, traversal evaluates the transport formula. When it doesn't, there is no answer to give — and the correct behaviour is to say so.

The one place representation is load-bearing Within a single population, identifiability doesn't care whether structure lives in a graph or a relational schema — the substrate adds governance, not identifying power. Across populations, identifiability depends on the selection diagram, which is a graph. This is the regime where the representation genuinely changes what can be known.
SELECTION DIAGRAM XYZ S spend outcome audience mix marks the only difference source: randomised experiment, E[Z] = 0 target: observational, E[Z] = d TWO OUTCOMES S MARKS EVERY DIFFERENCE P*(Y | do(X)) = E [θ(Z)] source strata effect × target's own Z-distribution AN UNMARKED DIFFERENCE EXISTS NOT IDENTIFIED no derivation exists — the traversal returns nothing
Fig 6 Transport is licensed by an annotation, not by convenience. The selection node S states what differs between the populations; the transport formula combines the source's stratum-conditional effect with the target's own composition. If something differs that S does not mark, the formula is invalid and no amount of computation rescues it.
The part most systems don't have
NOT IDENTIFIED

Sometimes the correct output is no output.

No formal result can rescue an unmarked difference, because the diagram is the only statement of what differs. So the graph runs an admissibility probe against a small randomised target sample and abstains when the residual says the structure has moved.

Abstention rises with the size of the violation — 4% on a genuinely transportable control, 61% at a moderate unmarked difference, 100% once it is large. A system that always answers will answer these cases too, confidently, and with roughly twenty times the error.

What refusal costs The probe needs a small randomised sample in the target — far cheaper than a full experiment, but not free. And it declines about 4% of transports that would have been perfectly valid. That is the price, and it is worth naming rather than burying.
Read this before believing any of it

Where this fails.

Three things a fair reviewer would raise. We would rather raise them first.

All of it is synthetic.

Every study simulates a world, then demonstrates a mechanism inside the world it simulated. That isolates the mechanism cleanly and proves nothing about magnitudes in the field. Real-data application with experimental holdouts is the necessary next step and has not been done.

Within one population, you probably don't need a graph.

Identifiability is a property of causal structure and available interventions, and it does not care whether you store that structure as a graph or a relational schema. The representation only becomes causally load-bearing across populations. One of the four studies lives in that regime. Three do not, and the paper says so rather than blurring it.

Governance can manufacture false authority.

A wrong effect that arrives with a validity window, a provenance chain and a credible interval is worse than a wrong number on a spreadsheet, because every signal a reader uses to gauge trustworthiness reports healthy. Better governance of bad evidence produces a more convincing artefact, not a more correct one. An ungoverned spreadsheet at least advertises its own unreliability.

Why publish the failures Because the alternative is worse. Anything a competent reviewer would find in an afternoon is better said first — and the scope conditions are the most useful part of the work for anyone deciding whether it applies to them.

The paper

The Marketing Context Graph

A Governed Substrate for Causal Marketing Measurement and Auditable Decision Provenance. Full treatment — set-theoretic definitions, the snapshot operator and its look-ahead guarantee, three propositions, four studies, and a limitations section longer than most papers' results.

Affiliation and conflict of interest

This work originates at Lifesight. Anil Kumar Singh is its CTO and co-founder; Rajeev Nair is also affiliated with Lifesight. The company develops marketing-measurement software including media-mix modelling and incrementality tooling, and the Marketing Context Graph reflects an architecture in commercial development.

We state this at the top rather than the bottom. No proprietary dataset, customer data or commercial product is evaluated or benchmarked anywhere in the paper; all results derive from synthetic data with known ground truth, and every comparison is against a standard public method — a flat observational MMM, back-door-adjusted estimation, and a flat fact/vector store.

Corrections welcome If you find an error — in the mathematics, in a claim that overreaches, or in a scope condition we have stated too generously — we would rather hear it than not.