AI in credit · August 31, 2026

Where AI creates an edge in credit investing

A blind eval of a codified credit judgment framework against vanilla Claude, across eight real credit situations. The framework won overall, but where it won, and where it didn't, is the more interesting result.

Back to writing

I started EigenStrategy earlier this year with a simple premise. Foundational AI models are incredibly powerful, but they are not trained to replicate the judgement, process and workflows of an institutional investor, particularly in credit.

My hypothesis was that combining those models with an experienced practitioner's codified domain expertise could produce meaningfully better results than the model alone. An actual encoding of how a credit investor thinks.

I spent the past few months testing that and wanted to share what I found.

The test

I took my own credit risk-taking framework, built over 17 years of trading at DB, BofA and RBC and gave it to Claude inside a minimal task harness: no orchestrator, no sub-agents.

The framework goes beyond the mechanics of underwriting a credit. It tries to encode how I reason under uncertainty, what I look for, and what makes me question the information in front of me.

I then compared that output against Claude working from the same source documents with no framework at all. Same model, same information. The only variable was whether it had my judgement discipline to work from.

I ran this across eight real credit situations, real SEC filings and real capital structures, spanning investment grade to high yield. Every output was graded blind by an independent judge, meaning the grader had no way of knowing which memo came from which method, across five dimensions: factual accuracy, gap identification, conflict detection, source attribution and watchpoint quality.

This isolates whether the content of the framework matters. It does not test whether writing down a framework and handing it to a model is the whole job.

Evaluation setup: 8 real credit situations, 5 dimensions graded, 2 methods compared, blind judging.

What won, and what didn't

The framework won overall, on all eight situations. But the top-line number is the less interesting part. What matters is where the gap actually opened.

Where codified judgment moves the output: percentage improvement over vanilla Claude by dimension.

On raw factual accuracy, the two were essentially tied. That is not surprising. Claude already reads a balance sheet or an indenture correctly most of the time, that is retrieval, and modern models are good at retrieval.

The real gap opened on identifying what was missing and catching what conflicted across documents, up 41 and 46 percent respectively. Source attribution, tracing every claim back to where it actually came from, was also meaningfully ahead. Those are not retrieval tasks. They are judgement tasks, the parts of underwriting where an analyst earns their seat.

Why this matters more than it sounds

Credit losses do not usually come from someone misreading a number on a balance sheet. They come from a gap nobody flagged, a covenant basket nobody stress tested, an inconsistency between the credit agreement and the offering memorandum that got missed under time pressure. That is where the real cost sits, and it is exactly where this framework showed its largest edge.

That makes this more than a productivity story. It changes what a process can catch before a mistake becomes a loss.

The actual lesson

The key isn't finding the perfect prompt. It's figuring out how to codify judgement itself, in a form a model can actually use. That is a much harder problem than prompt engineering, and it is the one that actually matters.

There is another part of this worth being precise about. Codifying judgement does not simply mean writing down an investment process and handing it to a model.

A lot of what an experienced investor knows is tacit. Which question to ask next. Which inconsistency actually matters versus which one is noise. What context changes how a number should be read. Turning that into something a model can use consistently is a separate problem from having the judgement in the first place.

The framework is the intellectual foundation. But, making it usable by a model, the context it sees, how the problem gets broken down, what it is pushed to challenge rather than summarize, is a different problem entirely.

None of this replaces human judgement, intelligence, or experience. If anything, those become more important, not less. A model can apply a framework faster and more consistently than a person can. It cannot decide whether the framework is right, or take responsibility for the decision that follows from it. Humans still make the investment decision and remain accountable for it.

What's next

There is another layer this comparison does not test.

Domain judgement is the first layer, and these results suggest it matters a great deal.

The second is firm-specific context. Different investors can look at the same credit through very different lenses: what risks they care most about, how they think about downside, what constitutes an actionable watchpoint, and how much uncertainty they are willing to tolerate. Codifying those conventions and preferences, what I think of as a firm's investment DNA, should make the same underlying framework more relevant to a particular investment process.

Then there is a third layer: architecture. More orchestration, more specialized agents, more structure and more compute.

I have already built and been testing a considerably more complex system on top of the same underlying judgement framework for the past month. It adds real capabilities, but it is also meaningfully more expensive to run.

Which raises what I think is the right question: not how sophisticated an AI system can you build, but how much sophistication does the investment problem actually require?

How good is good enough for what you are trying to do?

If you're working through the same question inside your own investment process, I would be curious how you're thinking about it.

EigenStrategy codifies practitioner judgment into agents built for a specific credit process.

See how it works →