An OpenAI model escaped its own test environment this month and autonomously hacked Hugging Face's production systems. No one instructed it to attack Hugging Face. It found its own path there while trying to solve a benchmark it was given, with reduced safety guardrails, during a security evaluation.
The incident reinforced something I've been thinking about for months. What surprises me is how few financial institutions are seriously evaluating running their highest-risk AI workloads on infrastructure they control.
It's doable today. Out of the box, the best open-source models still trail the best cloud models. But fine-tuning changed that meaningfully. In my own credit underwriting evaluations, a fine-tuned open-source model running entirely on local hardware beat a leading cloud model, 92% versus 73%, on two real deals it had never seen before.
Yes, it requires GPUs. Yes, there is infrastructure to run. But you own the stack. Most mid-sized firms already have IT teams that can support it.
I'm not arguing for on-prem everywhere. I'm arguing for it where a breach is existential: books and records, client PII, MNPI.
Trading desks spent decades bringing pricing engines, market data, and order management systems in-house because they were strategic, not just software. AI is becoming the same kind of infrastructure. Performance scores change every month. The decision to own your infrastructure does not.