top of page

Stop Playing with AI, It's Time to Put an Agent to Work

  • Writer: Alan Thomas
    Alan Thomas
  • 10 hours ago
  • 10 min read

CTRL AI DEL is Dailoqa's weekly, no-nonsense look at what's actually happening with AI in financial services. This issue looks at why so many AI investments underwhelm, and the practical framework that separates a genuine agentic AI deployment from another expensive pilot.


this is the featured image of CTRL AI DEL post topic Stop Playing with AI, It's Time to Put an Agent to Work

Most enterprise AI investments have failed to live up to expectations since individual solutions don't communicate with one another. Agentic AI puts an end to this by using reasoning, planning and carrying out multi-step work throughout a process, as opposed to merely improving a single task in isolation. To achieve this safely it is necessary to have a proper framework, beginning with defining the business problem, securing early support, and establishing governance along with setting up a centre of excellence before scaling, not after.


Where the Enthusiasm Meets Reality


There is a sense in the market that AI has not quite reached the expectations that were set for it. Fraud detection has improved. Document search and synthesis have improved. Customer communication is more personal, and trading algorithms can analyse market data and sentiment with far less effort than before. But is that it. With all the money that has gone into this, the clients we talk to are somewhat underwhelmed by progress so far, and the more interesting question is why these functions still are not talking to each other.


That disappointment shows up in the data too. MIT's Project NANDA found that roughly ninety five percent of enterprise generative AI pilots failed to produce any measurable financial return, and the researchers behind the study pointed to poor integration rather than weak models as the actual cause¹. The technology mostly works. What is missing is the connective tissue between one AI function and the next.


Fraud detection improves in isolation. Document search improves in isolation. Each function gets marginally better at its own job, and none of them talk to the others, which means a flagged transaction still has to be picked up, interpreted and acted on by a person before anything actually happens. That handoff is where most of the promised value quietly leaks away.


This pattern has a name now. Without a central function coordinating AI investment, most large organisations end up with what practitioners call agent sprawl, a scatter of overlapping tools bought by different teams, each with its own governance, its own vendor relationship and its own definition of success⁴. None of those tools are individually bad. The problem is that nobody owns the space between them, so the organisation ends up with a dozen small improvements and no compounding effect.


What Agentic AI Actually Changes


Agentic AI is often described as the more disciplined answer to this problem, the adult in the room. It is not a cleverer chatbot or a better looking dashboard. It is designed to carry out complex, multi step work that spans process, workflow and function, rather than improving one task in isolation. It can reason through a problem, plan a sequence of actions, and execute a full task with limited supervision.


Imagine a series of agents acting in unison when dealing with a suspicious transaction. The first of them freezes the account, another prepares the compliance report that is needed in this case, and a third sends the necessary information to a human analyst having already put together the relevant context, rather than sending a basic alert that the analyst would have to look into from the beginning. This kind of coordinated response is no longer just a matter of theory. Fiserv's newly launched agentOS platform is currently implementing this exact pattern in production, and one credit union that is using it has already reduced a daily reporting task that previously took ten minutes down to just a few seconds². The change is not due to a single ingenious model carrying out all the tasks; it is the result of several agents, each carrying out one specific job well, passing the work on to one another without any person having to manually fill in the gaps between them.


AI is still an emerging technology, and progress is moving quickly enough that the gap between what was promised and what is actually being delivered is finally starting to close.


We touched on this distinction in an earlier issue, the difference between a system that executes a single instruction and one that pursues an intention across several steps. That distinction is exactly what separates an isolated point tool from an agentic one. A document summarisation tool executes an instruction. An agentic KYC workflow pursues an intention, gathering the documents, checking them against policy, flagging exceptions and routing anything unusual to the right person, without waiting for a new instruction at every step. The value was never in making any single step faster. It was always in removing the gaps between steps, which is exactly what a collection of disconnected point solutions cannot do, however good each one is individually.


The Two Things Firms Skip


Start with the business problem, not the technology


It is far more fun to buy the newest tool than to sit with a genuinely painful process. But before getting drawn in by what an agent could do, it is worth pinning down what it should do. Is the loan application process a mess. Does the compliance team spend most of its week buried in paperwork. Do the systems handling those two problems even exchange information with each other. That shift in emphasis, starting from the problem rather than the product, tends to separate a genuinely useful investment from an expensive experiment that never quite lands.


Bring the whole business along early


An autonomous agent cannot simply be dropped into a live environment and left to find its feet. These systems touch legal, operations and frontline teams, and their actions carry real consequences. Agentic AI can also unsettle how a firm is organised functionally, and it is not unusual to meet resistance from people who are wary of losing responsibilities or visibility they have spent years building. We have seen use cases submitted by stakeholder teams that were so minimal it was obvious they had little real interest in taking part. Getting genuine buy-in early is difficult, but it is what separates a system that is technically compliant from one that actually works the way the business works.


The kind of teams most likely to resist are those which have the most to lose should a process that they control first become visible, measurable and comparable throughout the business. That resistance usually doesn't make itself known directly; instead it appears as a use case that is submitted half-heartedly, as a workshop that is attended but not really engaged with, or as a pilot that is quietly denied the access it needs in order to succeed. By recognising that dynamic from the start, rather than viewing each objection as entirely technical, one can usually save several months on the project later.


Five Steps to Get This Right


None of what follows is complicated. Most firms already know this. Very few actually do it.


Identify and prioritise the use cases that matter


Instead of looking at the vendor's brochure, focus on the business processes where actual friction occurs. The natural tendency is to go for the easiest victory, but a better aim should be the process that is nearly breaking down—the one causing the most harm the longer it is left unattended. When setting priorities, do so based on where true value lies, not on which demonstration looks most impressive to a boardroom audience, and be clear from the outset about how value is to be defined, not only after the project has started.


In reality this generally involves selecting from a limited list which is used by most banks, in the areas of KYC and onboarding, loan origination, and complaint resolution. All of these processes have the same structure: they are multi-step, involve a number of handoffs between teams, and leave a paper trail that regulators will one day want to examine. It is exactly because of this common structure that they respond so well to an agentic approach and so badly to a separate, standalone solution. When value is defined at the outset it could be something like cycle time, error rate, or the number of cases that have to be escalated to a senior reviewer, but it must be a figure that the business is already monitoring at the moment, not a new measure created specifically to make the pilot look good later on.


Put a governance framework in place before you need one


This is not paperwork for its own sake. Singapore's Infocomm Media Development Authority published one of the first governance frameworks built specifically for agentic systems in January 2026, organised around four principles, bounding risk upfront, making a human meaningfully accountable, building in technical controls, and giving end users a genuine role in the process³. A company needn't take up that specific framework, but the general structure is helpful. It is necessary to decide beforehand which agents and tools are really required, how they are developed, deployed and monitored, and also who is to be held responsible if anything goes wrong, since something eventually will. An agent which cannot be audited is not suitable for use in a live environment even if the demonstration had suggested otherwise.


Run a real pilot before a full rollout


No one will buy a car without testing it by driving it first, and this reasoning holds true in this case as well. Therefore, launch a small and carefully controlled pilot scheme, demonstrate its value, work out the various difficulties that are certain to arise, and then see how the system performs when actually used under real conditions before rolling it out across the whole company. Approach it as a real test and not merely as a routine step leading up to a decision which has already been made.


Why the retrofit always costs more


Firms that skip governance early tend to assume they can add it later once the pilot proves itself. In practice, retrofitting audit trails, permission scoping and decision logging onto a system that is already live and already trusted by the business is a far harder job than building those things in from day one. The system has to keep running while the change happens, the people using it have already formed habits around how it behaves, and any gap discovered during the retrofit becomes a live incident rather than a design decision made on paper. Governance built in from the start is cheaper in nearly every case, it is simply less visible on the timeline that gets presented to the board.


Train people to work with the agent, not just around it


Dropping new technology on a team and expecting them to figure it out rarely goes well. People need to understand what the agent can do, where its limits are, and how to work alongside it, partly so they grasp its capabilities properly, and partly so they stop assuming it is quietly coming for their job. A team that understands a system tends to trust it more, and a team that trusts it tends to actually use it.


Build a centre of excellence rather than letting every team improvise


Without a central function, most organisations end up with agent sprawl, a scatter of overlapping tools, inconsistent governance, and teams solving the same problem in different, incompatible ways⁴. A dedicated centre of excellence gives a firm a single source of truth for best practice and strategy, so the approach to agentic AI becomes a coordinated effort rather than a collection of disconnected side projects competing for the same budget.


The centre of excellence is also where the dial we have described elsewhere in this series actually gets set and adjusted. A loan origination agent handling a straightforward, low-value application can run with a light touch. The same agent handling a complex commercial facility needs the dial turned down, with a named person checking its work at defined points. Without a central team tracking which use cases sit where on that spectrum, every department ends up guessing at the right level of autonomy on its own, and guesses drift toward whichever setting is most convenient that week rather than whichever is actually appropriate for the risk involved.


AI that stays confined to isolated pilots was always going to underwhelm. Agentic AI, built on a real framework rather than enthusiasm alone, is the version of this technology genuinely worth betting the business on. The firms that treat it that way now will have a real head start once everyone else catches up.


FAQ on Agentic AI


A few questions we hear most often when firms are working through this in practice.


What is agentic AI, in practical terms?


It is AI that can reason through a problem, plan a sequence of steps, and carry out a multi-step task across systems, rather than answering a single question or completing one isolated task.


Why do so many AI pilots fail to show a return?


Most failures come down to integration, not model quality. A tool that improves one task in isolation, without connecting to the processes and systems around it, rarely produces a measurable business result.


What should a firm do before investing in agentic AI?


Identify a real, painful business problem first. Choosing the technology before the problem is one of the most common reasons agentic AI projects stall.


What does a governance framework for agentic AI actually need to cover?


At minimum, it should bound the risk of each use case upfront, name a human who is meaningfully accountable, build in technical controls, and give end users a genuine role in how the system is used.


Why does agentic AI need a pilot before a full rollout?


A pilot exposes the integration problems, edge cases and process gaps that never show up in a vendor demo, and it does so at a scale where mistakes are cheap to fix.


What is an agentic AI centre of excellence, and why does it matter?


It is a centralised team that owns strategy, governance and best practice for agentic AI across the business, preventing every department from running its own uncoordinated pilot.


How much human oversight does an agentic AI system need?


Enough that every agent action can be traced back to a named owner and an audit trail, with a human able to intervene at defined points rather than only after something has already gone wrong.


7. Glossary


  • Agentic AI, AI that can reason, plan and execute multi-step tasks across systems with limited human input at each step.

  • Agent Orchestration, the coordination layer that lets multiple specialised agents work together on a single workflow.

  • Centre of Excellence (CoE), a centralised team responsible for an organisation's AI strategy, governance and best practice.

  • Agent Sprawl, the accumulation of disconnected, inconsistently governed AI tools and agents across an organisation.

  • Pilot Program, a small, controlled deployment used to test an AI system's value and behaviour before a full rollout.

  • Human-in-the-Loop, an oversight model where a person reviews or approves an agent's action at a defined point in a workflow.

  • Audit Trail, a recorded, traceable history of what an AI agent did, why, and under whose authorisation.


8. References


1. Forbes, MIT Finds 95% Of GenAI Pilots Fail Because Companies Avoid Friction, August 2025. https://www.forbes.com/sites/jasonsnyder/2025/08/26/mit-finds-95-of-genai-pilots-fail-because-companies-avoid-friction/

3. Singapore IMDA Model AI Governance Framework for Agentic AI, summarised via WithTAI, Agentic AI Governance Framework 2026, 2026. https://withtai.com/knowledge/what_is_the_agentic_ai_governance_framework_2026_and_why_should_leaders_care_now.php

4. AI Agent Square, AI Center of Excellence Guide 2026. https://aiagentsquare.com/blog/ai-center-of-excellence-guide

Comments


bottom of page