Your Company Is Throwing Away Its Thinking
Dispatches from the Agentic Frontier is a regular intelligence briefing for leaders in knowledge-intensive sectors. Each dispatch translates evidence from the frontier of Agentic AI – practitioner experience, investor signals, strategy research, market events – into what it means for enterprise leaders building for competitive advantage. All are filtered through the Intelligence Capital framework developed in The AI Your Competitors Can't Buy.
Summary
Your business runs on thousands of judgement calls a week, made well below the Executive Team. Nobody records the reasoning behind any of them. When two people are handed the same case, they reach different answers, and none of your systems can tell you which one was right.
Insurance companies measured this. Two underwriters priced the same risk 55% apart. Their executives had guessed 10%.
AI is about to make this worse or better for your company, depending on a key design choice. The efficiency everyone is buying from vendors will be bought by your competitors too, and it will leave you level with them. The reasoning behind the decisions made within your company is the one thing a competitor cannot get hold of at any price, because it is made up of the cases only your firm has decided.
This article is for those who are approving AI spend and want to know what they can get from it that their competitors can’t.
Managing judgement within a firm
Most of the judgement in a business is exercised many levels below the board, by people whose names leaders mostly don’t know. The underwriter pricing a risk that sits at the edge of what the firm normally accepts. The credit officer setting terms for a customer whose payments have just started slipping. The hiring manager choosing between two candidates, neither of whom matches the specification. The project manager deciding whether a supplier's extra charge is fair.
There are hundreds of them in a mid-sized company and thousands in a large one, and between them they settle tens of thousands of cases a year that the standing rules can’t reach.
Standing rules are the policies, thresholds, guidelines and precedents a company has already written down. They answer most cases. The ones they do not answer are the cases that stop with a person, and those are the decisions this piece is about.
The book, by a Nobel Laureate and colleagues, documents how much professional judgement varies from case to case, independent of any directional bias.
Hand the same case to two of those people and they will not reach the same answer. Most firms have no idea how far apart.
Insurers measured it. Putting the same risk in front of two underwriters and the prices came out 55% apart at the median; the same claim in front of two adjusters, 43%. Executives at the firms audited had predicted differences of around 10%. The figures come from the professional noise audits reported by Kahneman (Nobel laureate in Economics), Sibony and Sunstein in Noise (2021), which found the same pattern wherever they looked, from courtrooms to consulting rooms.
Nobody audited was incompetent. Competence and consistency are separate properties, and most firms hire for the first while leaving the second unexamined.
In theory, firms hold the answer to every decision. The problem is that the reasoning behind nearly every one of them is never written down anywhere it can be found again, which is why the variance persists and why no corporate system can say which is right.
In this article we use insurance as a worked example, because it illustrates nicely the range of decision making that goes on within one of the sectors most exposed to AI. But my argument is relevant to any company that relies on cognitive labour.
The impact of Agentic AI
Strip the technology out of any AI business case and what is left is a promise to improve some set of decisions: to make them faster, cheaper, more consistent, or simply better. Two questions test that promise. Which part of deciding does it improve? And how long does the advantage last once competitors buy the same thing?
Agentic AI is different from what came before. Earlier tech systems optimised paperwork and left the deciding to a person: they filed the documents, ran the numbers, and put the file in front of someone who then made the call. An Agentic system, however, reads the file itself, checks it against the firm's rules, decides the cases it is authorised to decide, and passes the rest to a person with its working shown. Advanced AI systems that do this are already in production in some sectors and will rapidly become pervasive over the next few years.
Two things are emerging from this development, one of which is worth far more than the other.
1.) The work of deciding gets cheaper. The reading, the checking, the cross-referencing and the drafting are done by the system rather than by a person, and the routine cases are settled outright without one. What reaches an experienced person is the part that actually needed them. Fewer expensive hours per case, and more cases per person.
But this new level of efficiency is available to any company which writes a cheque. Your competitors can buy comparable systems from the same vendors and can reach a similar cost base; how fast and how well you do it is an execution question.
When every firm in a market can serve a customer more cheaply, the savings tend not to stay in their margins for long. In most markets they go fairly quickly into lower prices and better service, because that’s what firms do with a cost advantage when they are competing for the same business. (Insurance takes longer, because expense savings sit in the ‘combined ratio’ before they ever reach pricing. But the destination is the same.)
Jamie Dimon, CEO at JPMorgan, put it to analysts on 14th July this year, on the bank's second-quarter earnings call: "you don't uniquely benefit from AI". Customers get the benefit in the end. So does every competitor. The margin is real while it lasts, which creates the case for capturing it early rather than the case for expecting it to hold.
2.) The reasons behind each decision can now be kept. An Agentic system has to set out what it thought and why in order to act, and when an experienced person overrules it, the correction is their judgement stated. Both get written down automatically as the work happens, rather than by asking anyone to stop and write a report.
This works in the opposite direction to the first, because the gain is never competed away. Over time, more of the difficult decisions get handled to the standard of the firm's best people without its best people having to be in the room. This improves the decisions that set margins, and it frees scarce experts to focus on the cases that still need them.
Rivals can hire good people, to match the output. What hiring cannot reach is the record: a decade of reasoning on the firm's own cases, which stays behind when the person who wrote it leaves. The rival pays permanently for what the record partly banks.
This assumes, of course, that the reasoning going in is worth keeping. A record built indiscriminately gives a firm consistent versions of its own worst decisions, at speed. So, it’s important that what goes into it is a choice about whose judgement the firm wants to propagate. The underwriting audits described in Kahneman’s Noise are the reason that choice cannot be left to chance.
A firm already investing in Agentic AI is almost certainly buying process automation: submission triage, document handling, claims processing in insurance, for example. This is the right first purchase. It delivers the cheaper processing, and the savings pay for what comes next. The same programme can capture the second thing or throw it away, and most firms make that choice without knowing they made it.
What actually goes into a decision?
Let’s take, as an example, a risk that sits at the edge of the appetite for an insurer to insure. This is a good case to walk through because the accountable person is named and the outcome eventually gets recorded against them. The opportunity to underwrite a risk comes to an underwriting team and stops with whoever is accountable for the decision, often a team leader or a senior underwriter three levels below the executive floor. What matters is that the decision stops with them.
The file arrives as raw material: documents, figures, dates, correspondence. Organised, much of it now by machine, it tells the decision-maker what is true of this particular client, exposure and history. The firm then brings what it knows: the guidelines, precedents, playbooks and pricing rules that answer most cases most of the time. On this case they aren’t enough, which is why it stopped where it did.
What takes over is judgement: the ability to decide well when the standing rules are not clear. What the accountable person does with that ability is reasoning: she weighs the risk exposure against two past cases she remembers, considers declining, prices the uncertainty, drops an exclusion as unworkable, hears out a colleague who disagrees, maybe asks her personal AI assistant for input, then arrives at her view. The decision ends the deliberation. The risk is accepted, on adjusted terms.
The case then leaves the firm's hands. Whether it turns out well depends partly on the decision and partly on what happens next in the world: whether the client's business hits trouble, whether the market moves, whether the risk that was priced (say a weather event) ever materialises. Months or years later a result arrives, in the form of a loss, a renewal, a profit or a complaint.
Substitute other trades, in other sectors, and the anatomy is the same. The credit officer, the hiring manager and the project manager we referred to at the start go through the same sequence: raw material, encoded knowledge that runs out, judgement, reasoning, decision, and a result that arrives later.
Now track what the company keeps from any of them. The documents sit in the document store. The decision is logged in whichever system of record owns it, whether that is the policy admin or claims platform, the credit system, the applicant tracking system or the ERP: what was decided, by whom, on what date, at what price.
The result eventually reaches the ledger, which is to say it lands in the loss figures and the profit and loss account. The reasoning – the weighing, the two remembered cases, the colleague's disagreement and why it lost – is held in exactly one place, which is the head of the person who did it. When she retires, moves to a competitor, goes on holiday, is in a meeting when the next similar case arrives, or simply forgets, the firm keeps the answer but loses the reason.
Judgement transfers through many examples
Judgement is an ability. It lives in a person the way fluency in a language or skill at negotiation does, and it cannot be copied out of anyone's head. What can be written down is what happens when the ability is used: the reasoning on one particular case, which has a beginning, an end, and content another person can read.
This is how expertise has always been handed on. A junior sits alongside a senior for years, watching case after case, and builds the same ability from enough examples of it in use, which is why apprenticeships typically take years. The law does it at scale: a judge's approach to interpretation is written down nowhere, but read a few hundred of that judge's reasoned decisions and a competent lawyer can predict how the next case will go.
Some people may object that expert judgement is tacit: held as instinct, impossible to write down. That is true of the ability itself, but it does not block what follows, because what gets written down is never the ability, just each occasion it’s used. One underwriter's reasoning on one risk is an event. It happened, it had content, and it can be set down in a few hundred words. Enough of those events, held together and retrievable, is what we call a reasoning record.
Every firm already has a ‘judgement frontier’
Decisions divide into two kinds. The difference is when the thinking behind them was done, rather than how important they are.
In a routine decision, the thinking was done in advance. A straightforward insurance claim is paid in seconds with nobody weighing anything at the moment of payment; the weighing happened months earlier, when somebody ruled what counts as straightforward, set the thresholds, and corrected the system's early mistakes. In a difficult decision, the thinking happens fresh on the case, because the standing rules don’t reach far enough to resolve it.
The line between the two merits a name: the judgement frontier. It’s the line between the cases a firm's systems can already handle as well as an expert would, with nobody watching, and the cases that still need a person.
Every firm draws that line already, in its authority limits and escalation rules: what a person can sign off alone, and what has to go to someone more senior. The line moves outward one recorded case at a time, and only where there is volume. A decision taken thousands of times a year shifts it within quarters; one taken three times a decade doesn’t shift it at all.
The kinds of things firms can retain
An objection at this point might be that firms already record enormous amounts. They do. Sorting what they capture helps to clarify the argument:
Traces record what was done. System logs, decision logs, audit trails: the action, the timestamp, the actor, the amount. Every modern platform produces them, and regulation obliges firms to keep them. They hold the answer with none of the weighing.
Transcripts record talk in which reasoning occurred, in a form no future decision can use. Most firms hold years of email threads, chat channels and meeting recordings containing genuine deliberation: positions weighed, objections raised, alternatives discarded. A recording of a difficult meeting can hold the entire reasoning behind a major decision while contributing nothing to the next one, because nothing structures it for the case where it would be useful and nothing brings it back at the moment of deciding.
Reasoning records are the kind almost no firm holds. They carry the reasoning attached to the case: the question, the options weighed, the evidence relied on, the disagreement, the decision, the reason. And they are read back when the next similar case is being decided.
In our client work and on the public record, firms hold traces and transcripts by the terabyte, but reasoning records almost nowhere.
Read-back is the test: it’s the question to put to any supplier selling memory, context or knowledge capture. When the next similar case arrives, does the system put the reasoning from the last one in front of whoever is deciding, without anyone asking it to? The same content is just an audit trail when it sits in the archive, but it becomes an asset the moment the next case draws on it.
One element of the entry carries more information than any other: the override, where an experienced person overrules the system, the guideline or a colleague. It contains an answer plus the reason the standing rule was wrong, which is a correction to the firm's encoded judgement itself.
Today it is the entry least likely to be written down with its reason.
How Agentic AI enables a new organisational capability
Anyone who has held an operating role for long enough may have bought something like this before, and watched it fail.
The lessons-learned database, the post-mortem template, the knowledge management programme, the wiki that three people maintained: most of them asked experts to write down what they knew, and most of them decayed. Anyone who has bought this before has good reason to be sceptical.
Where the pattern did hold, in case law and in aviation incident reporting, two things were true. Writing up the reasoning was somebody's job rather than an extra. And the write-up was read by people whose next decision depended on it.
Both failed everywhere else. The writing cost fell on the busiest people in the building, so the case got decided, the write-up got deferred and the reason evaporated. And the reading depended on somebody knowing an entry existed and going to look for it, which nobody does under time pressure.
Agentic AI changes both conditions, for three reasons.
AI agents show their working as a by-product. An Agentic system reasons in text, because it has to in order to act. When it checks a submission against underwriting appetite, weighs an ambiguity, and refers the case, the account of why exists already, and nobody has to stop and produce it.
People's reasoning gets captured at the point of challenge. When an experienced person overrules the system's recommendation, that disagreement is the moment their judgement is most visible, and capturing it costs a sentence rather than a report, because the system has already set out what it thought and why. The same applies when experts work through their own personal AI agents day to day: the corrections they make are their judgement, stated.
The record can be read back without anyone remembering to look. This is the key characteristic of a well designed agentic AI architecture. A person handling the next similar case would have to know the earlier case existed and go searching for it. An Agentic system handling that case retrieves the earlier reasoning as a matter of routine, which is what turns a written record into a working asset.
However, there’s a caution that cuts across all three: making it easy is not the same as making people willing to participate.
Experienced people consenting to have their reasoning recorded, read back and argued with is the most demanding condition in a programme of this kind. It is hardest exactly where the volume is. A call centre supervisor or a claims team leader measured on cases closed per week has every reason to decide the case and move on, and none to spend three minutes explaining why. This is a management problem, best answered by giving the people whose judgement fills the record both status and a stake in what it becomes.
None of this happens by default. AI agents deployed one process at a time keep whatever they learn inside whichever vendor's system produced it, in formats nobody else reads, and it is gone if the platform is changed. For the reasoning to accumulate as something the firm owns, some key design choices have to be made at the start:
One record per case that every AI agent/agentic system and every person writes into;
Encoded rules setting what each AI agent/agentic system may decide alone;
The record must be held separately from any single vendor's software, so it survives a change of supplier.
Those choices sit inside the layer of the enterprise stack most firms do not yet have, which we call the Coordination Layer: the AI workforce that does the work, and the record, decision rights and governed system access that make its work accumulate. It sits on top of core systems rather than replacing them.
We recommend renting the models and the platforms, and owning the record, its schema and the decision rights.
Example of the Coordination Layer applied to an insurance carrier
The three choices should be seen as design constraints on programmes most firms are already funding or will soon fund, which is what makes them cheap to get right now and expensive later, even though they are not free: the record has to be specified and somebody senior has to own it.
Timing decides how much of it a firm gets. The rest of the AI programme can be built at whatever pace suits appetite and budget, and not too much is lost by waiting. The reasoning record is different, because it only ever holds cases decided after it exists. Every month it runs late is a month of reasoning the firm has already paid for and can never get back.
It reports early, too, which matters when the market-facing results are going to be some distance into the future. Within a quarter a leader can see what share of the difficult cases left a readable record behind them, and what share of new decisions drew on one. These are two important numbers to put in front of a board while it waits for the rest.
Results alone cannot grade decisions
The instinctive alternative to recording reasoning is measuring results instead.
Every result — the loss, the renewal, the profit, the complaint, whatever eventually lands in the numbers — is part decision, part luck. A carefully underwritten risk can turn bad in month three because the client's largest customer went under. A lazy one can run clean for five years because nothing happened to it. Neither tells you anything about the quality of the decision behind it until enough cases pile up to separate the two.
Some decisions grade themselves quickly. A decision is graded when enough has happened to show whether it was right. Flag a claim as suspicious and the investigation tells you within weeks whether you were right. Reprice a product and the conversion rate moves the same month. Where the verdict arrives that fast, a firm can learn from results alone, and systems that do this are on sale to every competitor.
The decisions that set margins are not like that. An insurer's casualty book can take a decade to show what was written into it. A credit limit, a supplier appointment or a senior hire all run to their own slow schedules.
This shapes what firms can currently prove about Agentic AI. Our review of the best-evidenced deployments worldwide, across all sectors, finds every qualifying case reporting from the fast-verdict territory: fraud hit rates, processing times, error counts.
The two that come closest to decision quality are both fraud programmes. Ping An's micro-agent claims and fraud network is credited by the company with RMB 6.44bn (around $900m) of fraud savings in the first half of 2025. Commonwealth Bank reports fraud losses down more than 20% in the first half of its 2026 financial year, with roughly three-quarters of its card-fraud rules now generated by an AI agent.
Neither has published the working, so in neither case can the contribution of the AI agents be separated from the fraud controls around them. And no case yet shows recorded reasoning as the mechanism behind better decisions. That is exactly what our argument predicts: nothing can learn from a verdict that arrives years late and is mostly luck, so the (current) published wins cluster where the verdict is fast.
A reasoning record changes what a firm can learn, in two ways that results cannot reach:
Before the verdict arrives, it allows the only fair question about a long-cycle decision in its first year: was the reasoning sound, given what was knowable at the time? A firm with the record can answer that in month three. A firm without it waits years to find out whether it has been making the same mistake all along.
After the verdict arrives, it turns the result into a lesson. The score records that the match was lost. The recording of the match shows why.
And it reaches the half of the business that results never touch. The submission turned away, the deal passed on, the applicant refused: a decline produces no result anyone will ever observe. Refusals shape an insurance company's book as much as acceptances, and on its declines a firm learns through the record or it learns nothing.
The first firm to build this owns the asset and the evidence together, and what it risks is small: a specified record, and the discipline to write into it. Even on the most pessimistic case, what it ends up holding is a complete account of how its hardest decisions were made, which pays for itself many times over in regulatory explanation, succession and training.
Good decisions come in two parts
Whether each decision was right, given what was knowable, is decision quality, and that is the part every firm believes it manages. Whether the decisions fit together is ‘decision coherence’, and almost nobody measures it.
A firm can score well on every decision individually while the set loses money. The underwriting audits from Kahneman’s analysis are coherence failing at its simplest: same firm, same standing rules, same risk, two prices 55% apart, with neither underwriter doing anything wrong by the standards they were held to. That is what makes the loss invisible in every report the firm produces.
Decisions fail to fit together in two further ways any executive will recognise. Across functions: one team resolves a contested question one way while another prices the opposite reading, each defensibly, leaving the firm with a position neither would have taken alone. And against the strategy, which is realised in the thousands of operational decisions taken beneath it, so decisions that inherit none of its choices leave it as a document.
A shared record fixes all three, and what decides which one it fixes is who reads it. Read back by the same desk, it makes each decision-maker consistent with their own past calls. Read across desks and functions, it lets the firm decide as one. A contested question gets settled once and applied everywhere. And the choices made in the strategy show up in the everyday decisions that are supposed to deliver it, rather than being restated in a document nobody consults.
Consistency on its own is available to anyone - any firm running one encoded set of rules gets it. What lasts is consistency anchored to a firm's own precedent, with the disagreements kept.
The danger named earlier returns here in a sharper form. Consistency built on a wrong premise makes every decision wrong in the same way, at speed, with the human variety that used to catch it removed. The protection is that the record holds disagreement as content: dissent and overrides are captured with their reasons, so a premise stays visible and challengeable in a way a silently shared assumption never was.
Three things a firm holds, but only one of them lasts
Everything above sorts what a company holds into three kinds of capital. What separates them is how long an advantage built on each one survives.
What the firm knows is Information Capital: its data, its information, its accumulated know-how. Any advantage here is temporary. The same vendors sell the same feeds and models to every buyer, and even a firm's own transaction and loss history is a head start with a horizon, because a rival writing similar business accumulates its own.
Who the firm employs is Human Capital: the judgement held in its people. The advantage is real but it walks. A poached expert takes her future judgement to the rival; the recorded reasoning of her past decade stays where it was written.
What the firm has recorded of its own reasoning is Intelligence Capital, and this is the only one that produces durable competitive advantage. The record is made of decisions already taken, on that firm's cases, in time that has already passed. A rival with more money cannot get to it, and neither can one that hires the same people.
Each wave of enterprise AI has made one of these cheap. Analytics and machine learning made information cheap to produce, which is why the economists Agrawal, Gans and Goldfarb argued in Prediction Machines (2018) that cheap prediction leaves judgement as the scarce and expensive part.
Generative AI made recorded knowledge cheap to reach. Agentic AI now reaches judgement itself, the part those economists expected to stay scarce. If implemented effectively it is the first wave that produces a written account of judgement being used while the work is done. None of this depends on human judgement staying scarce. It depends on the record of how judgement was used on a firm's own cases staying out of a rival's reach, which holds whoever, or whatever, is using it.
Any capability sold to every buyer produces an advantage that lasts until competitors finish installing it, and then drains away, quickly into lower prices in most markets and more slowly through the combined ratio in insurance. The record of how a firm's own people reason produces an advantage that does not run out, because it is built from cases only that firm has decided.
So the question to ask of the next AI investment that comes up for approval is simple. In three years, what advantage will it have built that a competitor with the same budget will not also have? If the answer is none, there is one change to make, and it does not need a new budget: start capturing the reasoning now, on the programme already approved. The record only holds cases decided after it exists.
Simon Torrance is CEO of AI Risk, an Agentic AI strategy and implementation consultancy. For the foundations of the Intelligence Capital thesis, see The AI Your Competitors Can't Buy.