Ask a quantitative model why it expects a position to work and the honest answer, most of the time, is: because things that looked like this have worked before. That is not a criticism. Correlation-based methods are the foundation of quantitative finance for good reasons — they are tractable, they are testable, and within a stable regime they are often right.

But "within a stable regime" is doing a great deal of work in that sentence, and the moments when it fails are precisely the moments that matter most. This article is about a different foundation: building research around directed relationships — what acts on what — rather than around co-movement. It is written for two readers at once: the principal who needs to understand what the difference buys, and the researcher who will want to know where the claims stop.

What correlation captures, and what it does not

Two series that move together are correlated. That observation is useful — it lets you predict one from the other, hedge one with the other, or combine them into a portfolio whose variance you can estimate. Almost all of quantitative infrastructure is built on refinements of that idea: covariance matrices, factor models, regression-based signals, statistical arbitrage.

What correlation does not tell you is why the two series move together. There are three possibilities, and they have very different consequences:

  • A causes B. Real yields rise; equity multiples compress. Intervene on the first and the second follows.
  • B causes A. The direction is the other way round, and intervening on A does nothing to B.
  • Something else causes both. A confounder — growth expectations, say — drives real yields and multiples together, and the relationship between them is real in the data but not a lever.

A correlation matrix treats all three identically. It records that the series co-move and how strongly. It cannot distinguish the cases, because distinguishing them requires information that is not in the joint distribution of the two variables.

Figure 1. The same three variables, two representations. The correlation view records that everything moves together. The causal view asserts that growth expectations drive both yields and multiples, and that the yields–multiples correlation is real in the data but is not a lever.

Why the distinction matters when it matters

In a stable regime, you can often ignore the distinction. If A and B have moved together for a decade, predicting B from A works whether A causes B, B causes A, or growth causes both — as long as growth keeps behaving as it has.

The cases where it fails share a shape: something changes upstream. A confounder that was stable begins to move. A relationship that ran one way starts running the other. A structural link — a central bank rule, a market convention, a regulatory constraint — is removed, and the co-movement it produced disappears with it.

A correlational model has no representation of "upstream." It encoded the association and nothing about the mechanism that produced it. So it gives no warning: the correlation simply decays, or reverses, and the model keeps predicting from a relationship that no longer exists. The post-mortem usually concludes that the regime changed — which is true, and is also a description of exactly what the model was unable to see.

The deeper problem is that post-hoc attribution does not fix this. Factor loadings tell you what the model did. Shapley values tell you which inputs mattered to the output. Neither tells you whether the relationship the model learned was causal, spurious, or the artefact of a confounder nobody named. Explanation after the fact is commentary on a decision already made; it is not the same as having encoded the mechanism.

What a causal structure is

A causal model represents relationships as a directed graph: variables are nodes, and an edge from A to B asserts that A acts on B — that if you could intervene on A, B would change, and not merely that they have been observed to move together.

Building such a graph is called causal discovery, and it is where the real work is. Some structure is known from domain knowledge: policy rates act on short-term yields, not the reverse. Some can be inferred from data using methods that exploit conditional independence — if A and B are correlated but become independent once you condition on C, then C sits between them, and the algorithm can often orient the edges. Some remains genuinely ambiguous, and an honest system says so.

The output is a graph in which every edge carries three things: a direction, an estimate of strength, and a statement of the evidence for it. That last part is what distinguishes a research tool from an oracle.

What direction lets you ask

The practical difference between a correlational model and a causal one is the kind of question you can put to it.

A correlational model answers conditioning questions: given that real yields are at this level, what has typically happened to multiples? It looks up periods that resembled the present and reports what followed. That is reasoning by analogy, and its reliability depends entirely on the present resembling the past in the ways that matter.

A causal model additionally answers intervention questions: if real yields moved fifty basis points and credit spreads were held constant, what would the structure imply for these instruments? It does not look for similar periods. It propagates the change through the graph, along the directed edges, holding fixed what you have asked it to hold fixed, and shows you the path the effect took.

The distinction is the one between observing and acting. Conditioning tells you what the world looked like when a variable happened to take a value. Intervening tells you what the model expects if you set the variable to that value — which is the question a scenario is actually asking.

Figure 2. A conditioning question looks backward for resemblance. An intervention question propagates a change forward through the graph, holding fixed what you have chosen to hold. Only the second is what a scenario is actually asking.

What a causal model cannot promise

This is the part a researcher will read most carefully, and it should be stated plainly.

A causal graph is a model, not the world. Every edge is an assertion backed by evidence of some strength. Some edges are well-supported by domain knowledge and data; some are inferred and could be wrong; some plausible edges are missing because the data never revealed them. A system that presents a graph without exposing the confidence in each edge is presenting a more elaborate kind of guess.

Financial data is hostile to causal discovery. Most discovery algorithms assume the data is drawn from a stable process — that the relationships being estimated hold across the sample. Financial time series are non-stationary: relationships drift, regimes shift, and an edge estimated over one decade may be weaker, absent, or reversed in the next. This does not make causal methods useless in finance; it makes honesty about the estimation window and its limits essential.

Unobserved confounders remain the central risk. If a variable that drives two others is not in the data, the algorithm may draw an edge between them that does not exist. Domain knowledge is the primary defence — knowing what to include — and the secondary defence is treating any surprising edge as a question rather than a finding.

An intervention result is a modelled path, not a forecast. It tells you what the structure implies under stated assumptions. The assumptions may be wrong. The structure may be incomplete. The value is in making the reasoning inspectable — you can see which edges the result depends on and decide whether you believe them — not in the number at the end.

What good practice looks like

Given those limits, a causal research process has a recognisable shape.

Assumptions are stated, not embedded. Every scenario names what is intervened on and what is held fixed. Every edge shows its evidence.

Domain knowledge and data discovery are combined, not opposed. Known structure constrains the search; the search surfaces relationships the researcher had not considered; the researcher decides which to keep.

Regime awareness is built in. Edges are estimated over windows the researcher chooses and can compare. A relationship that holds in one window and not another is a finding, not a bug.

A human decides. The model produces structure, strengths, and modelled paths. It does not produce recommendations, and it does not act. The researcher reads the graph, questions the edges, runs the scenarios that matter, and reaches their own conclusion — for which they, not the model, are accountable.

Everything is auditable. Six months later, when someone asks why a decision was made, the graph, the assumptions and the scenario that informed it can be reproduced. That is what "explainable" has to mean in a research context — not that an explanation was generated afterwards, but that the reasoning was on the table at the time.

The practical difference

Correlation tells you what has moved together. That is valuable, and nothing here argues for discarding it. Causal structure tells you what the model believes acts on what, lets you ask what would follow from a change rather than what followed a resemblance, and exposes the assumptions the answer depends on.

For a research team, the difference shows up in the questions that can be asked, the failures that can be anticipated, and the conversations that can be had with a risk committee — where "the model says" is a weaker position than "here is the structure, here are the assumptions, here is where I think it might be wrong."