Field Notes

When AI Gives You a Story the Data Hasnt Earned Yet

When AI Gives You a Story the Data Hasnt Earned Yet

Don't let AI explain your data—make it earn the explanation.

Don't let AI explain your data—make it earn the explanation.

Share:

Copy link

Copied!

One rabbit hole I’ve been wandering down is using Claude as a first-pass analyst for trend analyses. The workflow is simple enough: give it a model or dashboard, ask it what it thinks is happening, and then use the response as something to push against.

I’m not trying to outsource the work (😏). The real goal is to see what hypotheses it generates before I get too attached to my own.

This has worked better than I expected, but it has also exposed a failure I didn’t appreciate at first.

The most obvious failure mode is hallucination (ex: Claude invents a number). Annoying, but fairly easy to catch if you’re paying attention.

The more interesting failure mode is when the model uses real numbers to build a story the data hasn’t actually earned yet.

Here’s the example.

I was reviewing Net Revenue Retention (NRR) for a Series A SaaS company. At the blended level, NRR was sitting around 95%.

Then I split the customer base.

Enterprise was holding at 99%, while SMB had slipped to 85% (down from 95% two years prior). Given this dichotomy, I asked Claude why SMB was deteriorating relative to Enterprise.

Claude cited expansion dynamics.

The model’s interpretation was that SMB NRR had weakened because expansion revenue from customers who stayed was no longer sufficient to offset losses from downgrades and churn.

It wasn’t a far-fetched explanation. If expansion slowed while churn remained relatively stable, SMB NRR would naturally deteriorate.

This is where things got slippery.

The explanation fit the headline metric, but I wasn’t sure it had actually explained the dynamic. Finance is full of reasonable-sounding narratives. Some of them are useful. Some of them are pure optics.

So I dug in further.

Instead of asking whether Claude’s explanation made sense, I asked: if this explanation were true, what else would I expect to see?

If SMB NRR was primarily deteriorating because expansion had weakened, I would expect the spread between NRR and Gross Revenue Retention (GRR) to compress, since expansion would no longer be offsetting churn to the same degree. I would not necessarily expect logo churn to accelerate materially, nor would I expect ARR per churned customer to increase.

The first prediction held. The second and third did not.

The NRR-GRR spread was compressing, which suggested expansion was contributing to the deterioration. At the same time, logo churn was accelerating and ARR per churned customer was increasing, which suggested that larger SMB customers were also leaving the platform.

At that point, the question was no longer whether Claude was wrong. That framing is too crude, and it makes the whole exercise less useful than it could be.

A better question was whether the evidence supported Claude’s explanation more than the alternatives.

As far as I could tell, it didn’t.

Expansion was almost certainly part of the story, but it wasn’t sufficient to explain the full deterioration in SMB NRR. The underlying retention metrics pointed toward a more complicated set of dynamics than a single weak expansion narrative could capture.

Claude had not hallucinated. It had done something more subtle and probably more common: it connected a few true observations into a causal explanation before checking whether that explanation had survived the right tests.

That is a very human mistake, which is probably why it’s easy to miss.

Finance teams do this all the time. A metric moves, someone finds two supporting data points, and before long, a narrative hardens into the company’s explanation for what happened. Then it travels from the analysis to the memo to the board deck, picking up confidence along the way.

This is where AI can either help or make the problem worse.

If you treat the first answer as the answer, it can help you ship a cleaner version of a weak story. If you treat the first answer as a hypothesis, it can help you get to better questions faster.

That distinction has become the main change in my workflow.

Whenever Claude explains why a metric moved, I now ask what evidence would make that explanation less likely. Not what evidence supports it. What evidence would weaken it.

In practical terms, that means asking:

  1. What other metric should move if this explanation is true?

  2. What would I expect to see if the explanation is false?

  3. Which competing explanation would produce the same surface-level pattern?

Those questions are usually where the useful analysis begins.

I suspect this applies well beyond NRR. Forecast vs Actuals, cohort performance, sales productivity, and GTM efficiency metrics all have the same basic problem. Several explanations can fit the visible facts, and the cleanest explanation is not always the one best supported by the data.

For now, I’m less interested in whether AI can explain financial metrics on the first pass. I’m more interested in whether it can help generate hypotheses that are specific enough to test.

That feels like a more honest role for the tool, and probably a more useful one for Strategic Finance.

Published:

More to explore