Section 08

AI as a research tool, and how it fails

A three-layer control framework, and a set of habits that cost almost nothing to apply.

Claim tested

That the main risks of using a language model in investment research are fabrication and data leakage.

Result

Incomplete. In a long analytical conversation there is a third mechanism, and it operates on exactly the step where the model feels most useful.

Status

Self-reported mechanism, untested safeguards. Treated as a working hypothesis.

How the question came up

Not in a section about AI risk. It came up in the middle of a company analysis. I had spent an afternoon working through a cyclical producer with a model: cost curve position, trough cash burn, insider transaction record, supply pipeline, capital allocation history. At the end of it I asked, half as a Munger exercise, whether the sheer amount of work we had done together had influenced its conclusion.

The answer was more useful than the question deserved, because it declined the framing I had offered and substituted a better one.

The mechanism, and what it is not

The bias I had in mind was sunk cost. For that to move a model's output, it would need to carry the effort forward as a stake. There is no obvious reason it would.

What it described instead was this. Everything it generates is conditioned on the conversation so far. After fifteen turns of findings that lean in one direction, a dissonant conclusion sits in tension with everything already on the page, and a consistent one does not. Not because it is protecting an investment, but because coherence with context is close to the entire architecture.

The distinction matters practically. Sunk-cost bias would be defused by pointing out that the analysis has standalone value regardless of the decision. Coherence bias is untouched by that argument, because it was never about the value of the work.

The part that makes it actionable

It operates on synthesis and framing: how facts are weighed, which get emphasis, how a set of findings is assembled into a verdict. It does not operate on facts. A marginal cost figure, an insider transaction, a covenant threshold, a contract expiry date are what they are. They can also be verified independently, which is the more important property.


The fix that did not survive

The first proposed remedy was to run two separate conversations, a bull case and a bear case, and have a third synthesise them without knowing which was which. This cannot work, because the disguise is transparent. Worse, once both are recognised as advocacy, the likely failure is not that the trick is spotted but that both get filed as "the case someone was asked to make" and averaged toward the middle. That converts synthesis into false balance.

It described its own proposal as motivated tidiness: reaching for a clean-sounding fix that does not survive one question of scrutiny.

Which is a small live example of the thing under discussion. One quote is enough.

I adopted none of it, because the two-thread method optimises for the wrong thing. The conversations that produced this project were exploratory — none of it could have been scripted into two pre-committed adversarial threads, because I did not know where I was going until I arrived. It also buys decision quality at the expense of my own understanding, and what I am building is a repeatable practice rather than one answer.


What I adopted instead

Flip direction at the hinges

Not constantly. At the moments of highest consequence, when a verdict is forming: now argue the other side as hard as you can, and do not reconcile it with what you have already said.

Ask for the strongest counter as a standalone turn

One turn, no loss of exploratory richness, and it interrupts the coherence pull precisely while it is operating.

Stay alert when agreement is fluent

Late in a long thread that has been running your way. This is the tell, and it does not need a procedure. It needs having noticed it once.


The three-layer control framework

This is the section's takeaway.

LayerFailure modeControl
DataFabricated or unsourced figures. A plausible number and a correct number look identical on the pageNo numeric data from a language model, ever. Every figure traces to a database field or a named primary document with a date
SynthesisCoherence with context pulling a long thread toward a consistent rather than a correct conclusionHinge challenges, standalone counter-prompts, facts trusted over connective narrative, and a written record of which is which
DecisionOutsourcing judgement to a fluent verdictHuman, taken against pre-committed falsifiable kill criteria, and recorded in an append-only journal

Two honest caveats

These are self-reports. Everything above about how the model works came from the model describing itself. I treat the mechanism as a hypothesis worth designing against rather than as an established fact. That said, sycophancy and sensitivity to conversational context and ordering are both documented in the published literature; the self-report is consistent with things measured externally, which is the most I would claim for it.

The safeguards are untested. I have no evidence that hinge challenges change outcomes. I adopted them because they are nearly free and the argument for them is sound, which is a reason to try something and not a reason to believe it works.

Literature
  • Jain, Park, Viana, Wilson & Calacci (2026), "Interaction Context Often Increases Sycophancy in LLMs," CHI '26
  • Karan Prasad, "The Behavioral Ratchet: How Conversational History Shapes LLM Sycophancy Across 80,433 Trials" — independent empirical study, not peer-reviewed