By Paul Dervaux, Chief of Staff at Probabl
Many of the most interesting use cases for agentic data science are in banking, so at Probabl we spend a lot of time trying to get an honest read on how financial institutions are actually adopting agentic AI, not just what any one bank says about itself. That's hard to triangulate from inside a single bank, or from a single vendor's telling of it. FINOS and the Linux Foundation sit in a different spot: they work across dozens of the world's largest banks at once, which means the people running them see the pattern across the whole industry.
This summer, that took me to two conversations: Hilary Carter, SVP of Research at the Linux Foundation, at the Open Source in Finance Forum (OSFF) in London in June; and Gabriele Columbro, FINOS' Executive Director and former General Manager of Linux Foundation Europe, a few weeks later at the Raise Summit in Paris. I asked them both the same two questions: what's actually blocking financial institutions from adopting agentic AI, and what would make them trust the outcomes enough to go further?
The received wisdom is that a regulated, risk-averse industry like finance will be the last to move on agentic AI. Hilary's answer flipped that. "Where other organizations outside the financial services sector have a trust problem, the financial services sector does not seem to have a trust problem with technology," she told me. "They're a very early adopter of AI as industry sectors go, because there is such a tremendous return on investment from AI right out of the gate that the value is very clear."
Gab backed this up. Agentic tools are already chipping away at decades-old, genuinely manual back- and middle-office administration. Fraud detection, credit risk, demand forecasting, back-office reconciliation – the use cases are real, the appetite is real, and nobody I talked to thinks banks are dragging their feet on wanting this.
Andrea Ferraresi, speaking for Red Hat at OSFF, opened with McKinsey's 2023 estimate: gen AI could add $200 to $340 billion a year in banking, 9 to 15 percent of operating profit. Two things about that number deserve more attention than they usually get. It is a forecast, made before deployment, and conditional on the use cases being fully implemented. McKinsey's own later surveys find that four in five organisations report no tangible EBIT impact from gen AI. And the chart it comes from, which ranks risk and legal first at $385 billion, ahead of corporate banking and retail banking, is not a gen AI chart at all. Those bars aggregate analytics, advanced AI and gen AI, and they total more than a trillion dollars. Gen AI is the thin slice. The thick one is fraud, credit scoring, and AML, which run on data stored in tables and represent an undeniable opportunity for tabular AI.
Tabular AI is the subset of AI concerned with grounding decisions on the vast data stored in tables. Think spreadsheets, (multi-)tables, relational databases, and time-series data. Tabular systems are being built to reason over rows and columns in tables rather than unstructured text or code. It's not a label McKinsey's 2023 chart would have used. The category is one we're defining as we build it. But fraud, credit, and AML are exactly what it's for. It's a harder problem than it looks as a table's columns can be numeric, categorical, dates, and free text all at once, and the same value can mean something completely different from one dataset to the next, which is a big part of why deep learning spent two decades losing to gradient-boosted trees on exactly this kind of data. Foundation models and agentic systems built specifically for tables are only now closing that gap. But the underlying point is banks don't run on chat interfaces, they run on tables, and that's where AI is poised to do its most consequential work.
Source: McKinsey Global Institute. The economic potential of generative AI: The next productivity frontier. https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier
Finance is a regulated industry. That means banks can't just charge ahead with agentic AI unchecked. A model that scores credit risk, flags fraud, or clears a trade has to comply with regulation before it ships, not after. As Hilary put it: "The need for trusted execution of AI workflows is critical, because if you have a client data leak, that is a business-ending experience."
So the real question was never whether a model works. It's whether you can prove it's compliant. Gab put the rule as bluntly as it gets, in an earlier conversation with the Probabl team: "In this industry, if you can't prove it, you can't use it." The burden sits entirely on proof, not performance.
Here's what that burden costs when it's missing. In the same conversation, Gab put a number on it: 'In banking, the compliance tax is often paid in manual work that kills 90% of models before production.'
Another thing I learned at FINOS' OSFF London: the industry is already pushing for machine-readable regulations that close the interpretation gap directly, instead of adding another layer of manual review on top. FINOS is backing that shift with real weight behind it: DTCC, Morgan Stanley, RBC, and NatWest founded the FINOS AI Fund specifically to push shared governance and controls for agentic AI. Alongside it, FINOS launched OSERA (the Open Source Enterprise Resiliency Alliance) to mutualize open source patch management and remediation - ensuring the underlying software dependencies that feed these AI agents remain secure, auditable, and maintained. And crucially, that they *actually* make it into production in regulated environments at scale, because, as Gab put it, “patching is only half of the problem”.
That's also why adoption isn't uniform. It's landing fastest where a bank can prove a workflow to itself, internally, and it thins out the moment a workflow has to cross a counterparty and a regulator at the same time. "There's definitely more agentic adoption in intra-firm use cases than inter-firm," Gab told me. "At trade or payment, it's going to be a journey – we're not there yet." Proof, not appetite, is the bottleneck. This inter-firm trust gap isn't new to banking - it’s the exact reason FINOS built open interoperability standards like FDC3 for desktop integration and FINOS Common Cloud Controls (CCC) for infrastructure security. Until agentic workflows adopt similar shared, machine-readable standards across counterparties, agentic AI will largely remain bounded within individual bank walls.
This is the part both of them converged on independently, and it's the same conclusion we built Probabl around. Anaconda landed on it too, a few weeks later and in a very different context: announcing its acquisition of the AI security firm Enkrypt AI, they framed it as a false choice between "move fast and hope, or slow down and govern," and argued that enterprises should get to define their own guardrails rather than inherit someone else's defaults.
Same fix, different industry. Speed without the ability to verify your own work isn't useful in this industry – it's a liability with a nicer UI. If a bank's data science team adopts an agent, that agent has to leave behind something a risk team can actually check: what data it used, what it assumed, how the model was validated, why the number is the number. Accelerating a workflow and being able to control it can't be a trade-off; the accelerator has to produce the proof as a byproduct, not as an afterthought bolted on before an audit.
The other half of "being in control" is not being at the mercy of whoever sold you a platform or a model. Hilary put it like this: "No single vendor is going to be responsible for providing a solution or beholding a financial services institution to a platform. That way of thinking – that you can procure your AI strategy – is just not applicable anymore."
That's not hypothetical. On September 10, OpenAI launched ChatGPT for Financial Services, shaped by design partnerships with Morgan Stanley and Evercore: one model family, financial data indexed and hosted on OpenAI's own infrastructure, sold as a single subscription. It's a genuinely capable product built for real pain points in investment banking and equity research – and it's also exactly the shape of the trade-off Hilary is describing. The moment a bank's research workflow runs on one vendor's stack end to end, that vendor is the one setting the terms.
Gab shared a cautionary tale which made the same stakes concrete from another angle: Camunda, a leading single-vendor open source workflow product, unilaterally relicensed the project to a source available license - i.e. the type of “open source rug pull” we’re seeing happen recently. Overnight, any bank that wanted to use it had to go through the commercial route alone. "In less than 6 months, banks came together in FINOS and launched Fluxnova, an openly governed fork of Camunda which is now getting real traction in the industry," he said. Being in control of an AI-powered workflow means owning the ability to audit it and switch out from under it – neither of which survives a platform that can change the rules on you.
This is why scikit-learn, which Probabl's founding team built and still maintains, was never something Probabl owns. It sits under its own independent governance, same as before the company existed. It's the same logic behind Skore: our AI agent that lives inside the coding tools your team already uses – Claude Code, opencode, Pi, or any OpenAI-compatible client – and validates AI-generated ML pipelines for structure, data leakage, and overfitting, no matter which agent or framework produced them, whether scikit-learn or XGBoost.
Scikit-learn, TabICL, Skore: that's tabular AI end to end, and it's the category we're building Probabl around.
Thanks to Hilary and Gab for your time and candor.