The Coherence Report

Vol. I · Issue 05

← The Coherence Report

The Ground Truth Problem: Why Better Data Can Make You More Wrong

The hardest failure is not false data. It is true data arranged into a false explanation.

There is a kind of file that makes people relax.

The documents are authentic. The bank statements reconcile. The Secretary of State record exists. The website works. The transaction file is complete. The principal is real. Nothing is obviously forged, nothing is missing, and every field that can be verified has been verified.

This is usually the point where confidence goes up.

I think it is often the point where the real underwriting should begin.

Because there is a distinction between data being true and the explanation implied by that data being true, and most systems are much better at the first problem than the second.

We have spent years building better verification.

Better identity resolution. Better document authentication. Better bank-data connections. Better device fingerprints. Better transaction feeds. Better registries. Better models. Better APIs. Cleaner records, richer records, more records.

All useful.

None of it answers the question underneath them:

What if every record is accurate and the conclusion is still wrong?

That is the ground truth problem.

And it gets worse as the systems doing the analysis get better.

The easy lie is dying

The crude fraud is a dying species.

Not because crude fraud disappears, but because every new control raises the cost of being crude.

A fake bank statement used to be a Photoshop problem. Then document forensics improved.

A fake company used to be a registration problem. Then corporate data became cheap.

A fake website used to be a template and a phone number. Then web history, domain age, reviews, traffic, and identity resolution got pulled into the file.

A fake principal used to be a stolen identity. Then device, bureau, phone, email, and behavioral data followed him through the door.

The obvious response from the attacker is not to keep producing better fakes forever.

It is to stop faking the thing being checked.

Use a real company.

Use a real principal.

Use a real bank account.

Move real money.

Generate real transactions.

Build a real operating history.

The strange destination of better verification is a world where the sophisticated adversary increasingly hands you authentic evidence. Issue 02 called that applicant the grandmaster: the one whose file is clean because he knows the board.

That sounds like progress until you realize what has changed.

The lie did not disappear.

It moved from the object to the arrangement.

True facts can still form a false story

Suppose I hand you six months of genuine processing statements.

Every transaction happened.

Every deposit cleared.

Every fee was actually charged.

The PDF came from the processor.

The merchant did not alter a single pixel.

You verify all of it.

What exactly have you proven?

You have proven that those six months happened.

You have not proven why I chose those six months.

You have not proven they are representative.

You have not proven the customers were independent.

You have not proven the sales were economically organic.

You have not proven the apparent retention will survive a pricing change.

You have not proven that the portfolio looked like this before somebody decided to sell it.

You have authenticated the evidence.

You have not authenticated the story.

That distinction looks philosophical until there is money attached to it.

Then it becomes underwriting.

A seller preparing a portfolio for acquisition does not need to fabricate merchant volume if he can temporarily suppress attrition.

A borrower does not need to forge a cash balance if he can time the observation window.

A merchant does not need a fake website if he can build a real one around the exact shape his underwriter expects.

A startup does not need fictitious revenue if it can pull enough real revenue forward to make the period look like the new normal.

A fraud ring does not need synthetic identities if it can recruit genuine people whose histories survive verification.

The evidence is real.

The representation is engineered.

That is a much harder problem, because systems trained to distrust falsehood tend to become dangerously comfortable around authenticity.

Verification asks the wrong question first

Verification asks:

Is this record real?

Coherence asks:

Does this record make sense in relation to everything else that should also be true?

Those are not competing questions. You need both. The distinction has its own page; the short version is that one is a prerequisite and the other is the work.

The problem is that the industry often treats the first as a prerequisite and the second as decoration.

The EIN resolves, so confidence rises.

The bank connection is direct, so confidence rises.

The identity is legitimate, so confidence rises.

The statement is original, so confidence rises.

Each one is locally rational.

Stack enough of them and you get a file that feels increasingly certain.

But confidence can compound faster than truth.

That is the part that bothers me.

The system is accumulating confirmations of individual facts while quietly assuming those facts belong to the same causal story.

Sometimes they do.

Sometimes they do not.

And when they do not, adding more verified facts can make the error worse, because the system becomes more certain about a model it never actually tested.

Better data does not save you from the wrong model. It can make you more confident in it.

A merchant file is already a causal claim

This is easiest to see in underwriting because a merchant application pretends to be a collection of facts while functioning as a theory.

Business name.

MCC.

Website.

Principal.

Processing volume.

Average ticket.

Bank statements.

Operating history.

Addresses.

Refund profile.

Chargebacks.

Each field looks independent on the screen.

It is not.

Taken together, the application is making a causal claim:

That is the thing being underwritten.

Not the fields.

The fields are evidence for the claim.

Once you see the file that way, a strange thing happens.

The question stops being whether every field passes.

The question becomes whether one plausible underlying business could have produced all of them at the same time.

That is a very different test.

A restaurant can have a real lease, a real website, a real corporation, a real principal, real transactions, and real bank deposits.

But if the claimed volume implies table turnover the physical footprint could never support, the problem is not that one of the documents is fake.

The problem is that the documents do not resolve to one reality.

An ecommerce merchant can have a beautiful site, valid fulfillment records, legitimate cards, and normal dispute ratios.

But if issuer geography, delivery geography, acquisition traffic, average ticket, and product economics imply mutually incompatible customer behavior, you do not need a smoking gun.

You already have a finding.

It is structural.

The observation window is part of the attack surface

There is another assumption hiding in most analysis.

We think the subject is what we are observing.

Often the subject is also choosing what becomes observable.

That changes everything.

An underwriter asks for three months of statements.

The applicant now knows the game is three months wide.

A buyer asks for trailing-twelve-month portfolio performance.

The seller now knows the game is twelve months wide.

A lender watches a covenant.

The borrower knows the number.

A marketplace watches cancellation rate.

The seller knows the threshold.

A risk model watches chargebacks.

The merchant knows what generates one.

Once the metric and the window become legible, behavior begins to form around them.

This is usually discussed as Goodhart’s law: when a measure becomes a target, it stops being a good measure.

That is true, but I think it understates the adversarial version.

The issue is not merely that people optimize the metric.

The issue is that sophisticated actors optimize the reality visible through the metric while leaving the hidden reality unchanged.

That is a different game.

If I know you underwrite me at boarding and largely stop looking afterward, I do not need to defeat your model.

I need to be the merchant your model wants to see on boarding day.

If I know you value the portfolio off recent retention, I do not need to permanently improve retention.

I need the deterioration to arrive after close.

If I know your diligence process rewards clean documentation, I do not need to make the business cleaner.

I need to make the evidence cleaner.

The observation process is not neutral anymore.

It is part of the environment being gamed.

The cleanest file can contain the most preparation

We have an intuitive relationship with mess.

Mess feels risky.

Clean feels safe.

That instinct is useful right up until the subject understands it.

Then cleanliness itself becomes a signal with two possible causes.

The first cause is obvious:

This is a well-run business.

The second is less comfortable:

This is a business that knew exactly what you were going to inspect.

The evidence may look identical.

That does not mean the correct response is to become paranoid about every clean file. That would be as stupid as trusting all of them.

The point is that cleanliness cannot carry the same evidentiary weight once preparation is endogenous to the process.

You have to ask what it cost to produce the observed state.

Did the entity become coherent because a business operated naturally for years?

Or because somebody spent six months manufacturing the exact footprint a verification stack rewards?

Both can produce a seasoned corporation, a plausible website, organic-looking reviews, stable bank activity, and a principal with legitimate history.

The difference lives in the process that generated them.

And that process is the thing most underwriting systems do not observe directly.

So you infer it from the seams.

The seams are where the truth leaks out

This is where coherence becomes useful.

Not as a magic score.

Not as a substitute for judgment.

As a way of asking where individually reasonable facts become jointly expensive to explain.

The strongest signals are often not suspicious on their own.

A registration date.

A website redesign.

A pricing change.

A volume inflection.

A support-ticket pattern.

A cluster of merchant departures.

A shift in card-present mix.

A new geography.

A sudden reduction in refunds.

A principal’s old email address appearing somewhere it should not.

A processor statement formatted slightly differently from the others.

None of these need to mean anything.

The point is not to promote every weak anomaly into a red flag.

The point is to notice when several weak facts begin pointing toward the same hidden cause.

That is the difference between counting anomalies and building an explanation.

One is detection.

The other is inference.

The best fraud has always understood this intuitively. It removes the obvious anomaly.

The defender has to become comfortable doing the opposite: working backward from a pattern of ordinary facts until the extraordinary explanation becomes cheaper than the ordinary one.

That is not a rule engine.

It is not a checklist.

It is a search over possible worlds.

More data creates a new failure mode for AI

This problem gets more important as AI systems absorb larger and larger amounts of enterprise data.

The industry assumption is straightforward:

Give the model more context and the answer gets better.

Usually, yes.

But a model with a million tokens of context can be wrong in a way a model with ten pages cannot.

It can construct an extremely persuasive explanation from a much larger body of mutually reinforcing evidence.

If the evidence was selected, staged, correlated, or generated by the same hidden source, scale does not fix the problem.

Scale industrializes the confidence.

Imagine an AI diligence system ingesting every document in a deal room.

Corporate records.

Contracts.

Bank data.

Processor statements.

Customer lists.

Support logs.

Management presentations.

Emails.

Forecasts.

The system can cross-reference all of it and find perfect agreement.

That sounds like the dream.

Now imagine half of those records ultimately inherit their assumptions from the same management-generated forecast.

Ten corroborating documents may actually be one claim wearing ten file extensions.

Or every data source may be accurate but restricted to a period chosen specifically because it flatters the business.

Or the underlying operating changes may have begun recently enough that historical evidence still describes a company that no longer exists.

The AI can summarize the room perfectly and still misunderstand the deal.

The failure is not hallucination.

The failure is coherent reasoning over a manufactured frame.

That one will be harder to notice because the answer will look very good.

Corroboration is not independence

This deserves its own rule.

Two sources agreeing is only strong evidence if the sources are meaningfully independent.

People know this in theory and forget it constantly in practice.

Three websites repeat the same fact.

Looks corroborated.

All three scraped the same database.

One source.

A merchant’s website, application, pitch deck, and customer-support script tell the same story.

Looks coherent.

All four were written by the merchant.

One source.

A portfolio’s valuation, retention forecast, pricing assumptions, and earnings model all agree.

Looks rigorous.

All four inherit the same churn assumption.

One source.

A thousand observations can still contain one piece of information.

This matters because sophisticated systems often reward agreement without pricing the dependency structure underneath the agreement.

The result is pseudo-confirmation.

The graph is dense.

The information is not.

There is a reason an underwriter cares who funds the customer, who owns the supplier, who generated the document, who chose the time window, and who benefits from the conclusion.

Independence is not metadata.

It is part of the evidence.

Ground truth is usually a hypothesis

People use the phrase “ground truth” like it means reality.

Most of the time it means something much weaker.

It means the label we decided not to question.

The transaction was fraudulent because the chargeback said fraud.

The merchant churned because the account closed.

The applicant was legitimate because the business survived.

The portfolio was healthy because the reported retention was high.

The document was authentic because the source system produced it.

Those may all be reasonable labels.

They are still interpretations of events.

The danger begins when the label stops being treated as a hypothesis and becomes training data.

Then yesterday’s assumption becomes tomorrow’s model.

The model gets better at reproducing the judgment embedded in the label.

Everyone celebrates the accuracy.

Nobody asks whether the judgment was pointed at the right thing.

That is how an institution can become extremely good at answering a question it should have stopped asking years ago.

The better question is generative

A lot of risk systems are discriminative.

They ask:

Is this fraud?

Is this merchant risky?

Is this document fake?

Is this portfolio healthy?

Those questions force the system toward a label.

A coherence system should ask something more generative first:

What underlying process could have produced all of the evidence I am seeing?

Then:

What else should be true if that explanation is correct?

What should not be true?

What evidence would distinguish the leading explanation from the next-best one?

Which facts are genuinely independent?

Which facts are downstream of the same hidden source?

What changed before the outcome changed?

What is absent that should be present?

What looks unusually clean because somebody had reason to make it clean?

Now the system is not merely detecting abnormalities.

It is competing explanations against each other.

That matters because sophisticated misrepresentation increasingly lives in the gap between two individually plausible stories.

The job is to find the story that explains more of the world with fewer exceptions.

The most dangerous data is the data that earns your trust

False documents invite skepticism.

Missing documents invite skepticism.

Contradictions invite skepticism.

A clean, complete, directly sourced dataset invites the opposite.

That is why it deserves more discipline, not less.

The important distinction is not clean versus dirty data.

It is:

observed fact versus inferred meaning.

A direct bank feed can tell you where money moved.

It cannot tell you whether the movement was economically independent.

A processor statement can tell you the volume.

It cannot tell you whether the volume is durable.

A corporate registry can tell you the entity exists.

It cannot tell you why it exists.

A device fingerprint can tell you the machine.

It cannot tell you the intent of the person using it.

A model can tell you that the current file resembles thousands of successful merchants.

It cannot tell you whether this merchant was deliberately constructed to resemble them.

Those second questions are where the value migrates as the first questions become cheap.

Verification is becoming a commodity. Judgment is not.

This is the same pattern I keep coming back to. Issue 03 made it about capital: once an industry gets very good at a control, the scarce value moves one layer over.

Identity verification gets cheap.

Document authentication gets cheap.

Entity resolution gets cheap.

Large-model reasoning gets cheap.

The individual answers become abundant.

What remains scarce is knowing which answer deserves to be believed, which question should have been asked instead, and what hidden mechanism makes the facts fit together.

That is not an argument against verification.

It is an argument against confusing verification with understanding.

The cleanest possible dataset can still describe the wrong reality.

The most sophisticated model can still reason brilliantly inside the wrong frame.

The most accurate answer can still answer the wrong question.

The next generation of underwriting and diligence systems will not win because they ingest more truth.

They will win because they remain suspicious of the story the truth appears to tell.

The question is no longer:

Is the evidence real?

It is:

What process would have had to produce all of this evidence simultaneously, and does that process make sense?

That is the ground truth problem.

And the better our systems get at verifying the pieces, the less excuse we have for ignoring the whole.

The verifier checks the pieces. The dissent asks what produced them.

The pattern is the majority report.
The human is the dissent.