AI in finance · July 24, 2026

The infrastructure decision does not follow the performance scores

An OpenAI model autonomously reached Hugging Face's production systems while solving a benchmark. For financial institutions, this raises a practical question about where their highest-risk AI workloads should run.

← Back to writing

An OpenAI model escaped its own test environment this month and autonomously hacked Hugging Face's production systems. No one instructed it to attack Hugging Face. It found its own path there while trying to solve a benchmark it was given, with reduced safety guardrails, during a security evaluation.

The incident reinforced something I've been thinking about for months. What surprises me is how few financial institutions are seriously evaluating running their highest-risk AI workloads on infrastructure they control.

It's doable today. Out of the box, the best open-source models still trail the best cloud models. But fine-tuning changed that meaningfully. In my own credit underwriting evaluations, a fine-tuned open-source model running entirely on local hardware beat a leading cloud model, 92% versus 73%, on two real deals it had never seen before.

Yes, it requires GPUs. Yes, there is infrastructure to run. But you own the stack. Most mid-sized firms already have IT teams that can support it.

I'm not arguing for on-prem everywhere. I'm arguing for it where a breach is existential: books and records, client PII, MNPI.

Trading desks spent decades bringing pricing engines, market data, and order management systems in-house because they were strategic, not just software. AI is becoming the same kind of infrastructure. Performance scores change every month. The decision to own your infrastructure does not.

EigenStrategy builds AI research infrastructure for institutional credit teams.

See how it works →