Every analysis in CausoAI comes back with a Causal Readiness Score — a number between 0 and 100. It is not a measure of whether the result is good news. It is a measure of how much weight the result can bear.
This matters because the alternative is worse. Most analytics tools hand you a number with no indication of whether the data behind it could support the conclusion. The score exists so that a weak analysis announces itself instead of quietly looking like a strong one.
The four bands
- —80–100: High confidence. Fine for budget decisions, strategy and reporting upward
- —60–79: Directionally reliable. Use it to form a view and prioritise, not to move money today
- —40–59: Treat with caution. Real assumptions are doing work here — say so out loud when you share it
- —Below 40: Do not act on it. Something important is missing and needs fixing first
Check 1: Do you have enough data? (25%)
The first check looks at whether the file can support any reliable answer at all: how many rows there are, how much is missing, and whether there is enough variety in the numbers to compare anything.
Typical problems: fewer than a hundred rows, a key column empty for more than one row in five, or an outcome with almost no variation — everybody converted, or nobody did.
These gaps are the most fixable of the five. They usually mean widening the date range, filling in a column, or narrowing the question to something the data can actually answer.
Check 2: Can the question be answered at all? (25%)
This is the one that catches genuine dead ends. It asks whether, given the diagram and the columns available, it is even possible to separate the effect you want from everything else going on.
It looks for three things in particular: a factor that influences both sides and is missing from the file, so its effect gets attributed to your treatment; something being held constant that should not be, which distorts the comparison rather than cleaning it up; and timing that does not make sense, such as a treatment recorded after the outcome it supposedly caused.
A low score here is the most serious of the five, because more data will not rescue it. If the effect cannot be separated out, the fix is to change the diagram or add the missing column.
Check 3: Can the method actually run on this? (20%)
A question can be answerable in principle and still be impractical with the data in front of you. This check looks at whether the chosen method has enough to work with.
The usual culprit is a lopsided split. If only two percent of your customers received the thing you are studying, or only twenty of them did, there is very little to compare against, and any answer will be unstable no matter how large the file is overall.
Check 4: Are the groups comparable enough? (20%)
The whole approach depends on comparing like with like. This check tests whether that is actually possible — whether, for each kind of customer who got the treatment, there are similar customers who did not.
Where there are not, the analysis is extrapolating rather than measuring. That happens most often at the edges: your very largest accounts may all have received the treatment, leaving nothing to compare them against. This check also flags when the answer will come back with a range so wide that it cannot inform a decision.
Check 5: How easily could something you missed explain this? (10%)
The last check stress-tests the result. It asks how strong an unmeasured factor would have to be to wipe out the effect you found. If it would take something enormous — a factor bigger than anything already in your data — the finding is sturdy. If a modest one would do it, the finding is fragile.
It carries the smallest weight because it is diagnostic rather than disqualifying. It tells you how nervous to be, not whether the analysis is valid.
Read the gap list, not just the number
The score is a summary, and summaries hide things. More useful is the gap list: each specific problem, which check it came from, and whether it is blocking or a warning. Blocking gaps should stop you acting. Warnings lower your confidence without invalidating the result.
A pattern worth recognising: an analysis scores 71, which sounds workable, but the shortfall is one blocking gap in check 2. The right response is to fix that gap, not to proceed on the grounds that 71 seems decent. A single serious hole is not offset by four areas that are fine.
What to do when the score is low
- —Not enough data: widen the date range, fill the gaps, or ask a narrower question
- —Question not answerable: revisit the diagram, and add the factor that is influencing both sides but missing from the file
- —Method cannot run: the split is too lopsided — gather more cases on the thin side, or compare fewer things at once
- —Groups not comparable: restrict the analysis to the part of your base where you have both kinds of customer
- —Fragile result: get data on the factor you suspect is missing, or present the finding with the caveat attached