Mind the Gap: AI Efficiency, Human Oversight and the Risks in Between
July 2026As artificial intelligence (AI) becomes increasingly integrated into professional services, the benefits and the risks are revealing themselves. Initial assumptions of boundless productivity gains are being met with real-world challenges around the quality of output and the reliability of human oversight.
AI appears to create visible savings at the point of production; however, the potential risks it generates are less visible and more difficult to quantify but are of equal importance. Whether the professionals reviewing AI-generated work can sustain meaningful oversight, and whether insurers will continue to underwrite the associated risks, are questions that demand immediate attention.
The legal position
The legal framework is reassuringly familiar. Courts in England and Wales, and in other common law jurisdictions, have not needed new principles to deal with AI. In a professional services context, a professional who relies on AI-generated output owes the same duty of care as in any other context. The professional who delegates a task to a tool remains responsible for the output that tool produces. The act of delegation does not discharge the duty.
This consistency provides reassurance to clients, insurers and regulators, who know where responsibility sits. Professional services firms are responding through the implementation of robust AI policies and guardrails which ensure that output is signed off by a suitably qualified professional. This is referred to as the ‘Human-in-the-Loop’ (HitL). In a HitL system, the human is primarily responsible for validating AI outputs, but may also be required to ensure the accuracy and appropriateness of inputs where these materially affect the reliability of those outputs. It should be noted, however, that not all AI applications operate on this basis; some are designed to function autonomously, without human review of individual outputs, which raises distinct questions about accountability and risk allocation.
How professional services are changing
Whilst the extent of change is uncertain, there is a consensus that AI will reshape professional services firms significantly. Predictions from legal technology commentators including Richard Susskind and consultancy firm, McKinsey, suggest that junior and mid-level headcount in professional services firms will fall significantly as AI absorbs routine generative work. Estimates suggest a 30–50 percent reduction in overall headcount over ten years, concentrated in the lower tiers. Responsibility for supervision will sit with a smaller senior layer, expected to review substantially more AI-generated output without a corresponding increase in capacity.
This creates an obvious pressure on traditional development pathways, as when junior professionals are no longer doing the routine supervised work that builds expertise, the pipeline producing future senior professionals becomes thinner and shallower. The people expected to check AI output in 2035 will have spent less of their careers doing the underlying work themselves. A senior engineer reviewing AI-generated calculations in ten years’ time will be doing so across a higher volume of output, with a smaller team, having spent less of their career doing those calculations personally. This is a challenge that has not yet been addressed.
Can humans hold the line?
Concerns over the HitL oversight assumption are not limited to whether we can sustain a pipeline of future humans with sufficient expertise. There are also questions over whether humans will be capable, from a psychological and motivational perspective, of holding AI to account as well as the extent to which AI errors can be identified in the first instance.
- Automation bias and error identification
When AI produces output that is fluent, well-structured and formatted like expert work, reviewers are inclined to treat it as accurate. This is a well-documented tendency in how human attention interacts with authoritative-looking material. A well-presented output is not the same as a correct one, but once that distinction is lost in practice, human review becomes endorsement rather than an independent check. For professionals accustomed to reviewing and verifying information, the challenge extends beyond a tendency to trust well-presented material. AI-generated output may contain assertions that are compelling but fundamentally flawed, where the defect (often referred to as an AI hallucination) originates deep within the system eg a defective algorithm or the model’s misinterpretation of source material, with no outward indication of where the error may lie. Identifying the source of such misinformation can be akin to ’finding a needle in a haystack’, and the verification process may in some cases take longer than performing the underlying task independently. If something subsequently goes wrong, the record will show that a human reviewed and approved the output. Whether that reviewer had a realistic opportunity to identify the error is the harder question — and precisely the one that will matter in a claim.
An additional risk arises from litigants-in-person using AI to formulate claims. Without legal expertise to assess outputs, there is a heightened risk of inaccurate, misconceived or poorly evidenced claims being advanced. In turn, this may increase claims frequency and defence costs, as parties and courts expend additional time identifying and addressing errors.
- Scaling volume without scaling oversight
AI potentially allows teams to produce substantially more output which must be checked. Businesses pushing for an edge, whether competitive or financial, may not recognise the need to increase their capacity to check AI output as it is produced. Validating output can be a tiring and stressful exercise, which requires utmost attention to detail. Businesses will need to carefully consider the extent to which humans can, realistically, engage in that activity for long periods of time without becoming blind to the importance of their exercise. In the medium/long term, human capacity to validate AI output may be seen as an obstruction to progress. Some industries are already looking to move to ‘human-on-the-loop’, where only anomalies and exceptions are sent to humans to review. This requires a level of comfort in the automated output which is not scrutinised by a human.
- The problem compounds over time
As AI-assisted output becomes normalised, the baseline expectation for speed of delivery, and the required volume of output rises. What begins as a productivity improvement becomes minimum expectation, and returning to a more deliberate process feels like a step backwards, even if it produces more accurate outcomes.
- The incentive structure makes this worse
AI adoption is typically described in terms of what it saves at the point of generation. The pressure to deliver falls on those closest to the work, whilst the consequences of any resulting errors are pushed into the future. With an increased volume of work comes the increased pressure to get it done quickly, and with no incentive to slow down and take time to verify, this may well perpetuate a rise in mistakes. If there is not a culture of ensuring that the output is genuinely reliable, then there is a risk that the assumed productivity gains at the generation phase simply shifts the burden, and the associated costs, to a different area of the business.
- The latency problem
The risk increases the longer it takes for errors to surface. One error can create a serious issue, but repeated, automated errors can be catastrophic, especially if originating from a systemic defect. A gap between work production and error discovery creates a window during which the error can be repeated. This means that the risk profile of a firm using AI which is tested immediately, is materially different from one using AI for outputs which will be relied on in years to come.
The insurance position
As explored further in other articles in this series, the US insurance market has shown initial signs of caution over AI risk, and similar concerns around exclusions and alternative cover are beginning to emerge in the UK market. Whilst it is far from clear how insurers in all jurisdictions will react, the breadth of what is being excluded in the US is wider than most policyholders anticipate.
The underlying concerns, including the difficulty of observing the risk, inability to reconstruct the chain of responsibility and doubt about sustainable pricing are not jurisdiction-specific and so we are keeping a close eye on developments.
For risk managers, the immediate priority is to understand how existing policies would respond to an AI-related claim. It is not safe to assume that current coverage is adequate simply because it has not yet been tested. Engaging with insurers now, whilst AI exclusions are not yet standard in the UK, creates an opportunity to establish coverage terms and document AI governance practices.
It is also important to note that many indemnity insurance policies are on a ‘claims made’ basis, where the claim attaches to the policy in force at the time that the claim is made. By way of example, an error in a will drafted in 2026 may not come to light until many years into the future, by which time insurers may have deemed such reliance on AI to be an uninsurable risk, such that any liabilities for the 2026 errors become excluded from cover in subsequent policies. The latency risk discussed above, therefore, has direct implications for the availability and adequacy of insurance cover.
Conclusion
The challenges outlined above are interconnected. The effectiveness of HitL oversight depends on the expertise, capacity and incentives of the individuals performing it, and each of these is under pressure as AI adoption accelerates. The latency risk compounds these concerns: errors that remain hidden create exposure that may only materialise years after the work is produced, by which time the insurance landscape may have shifted materially.
Notwithstanding the clamour for AI adoption, any changes to processes should be made with the court’s approach to liability kept in mind. The human remains responsible for the output and so must be empowered to ensure that the quality of that output is appropriate.
Underwriters need to consider whether the organisations being assessed have genuinely thought through where verification responsibility sits, and how long it will take for errors to surface. HitL oversight ranges in practice from a rigorous, documented review process to a cursory sign-off by someone without sufficient expertise or capacity to scrutinise the output to the requisite standard. Emerging regulatory and professional obligations, including under the SRA Standards and Regulations, FCA requirements and the EU AI Act, increasingly emphasise that human oversight must be real, evidenced and auditable, reinforcing the need for firms to maintain a demonstrable record of their review processes.
The businesses that most successfully deploy AI are likely to be those that recognise how efficiency and quality are not automatically aligned. The HitL is only as effective as the loop allows them to be. The guardrails and checks need to be designed honestly, identifying where in the process the verification burden falls hardest, where errors can remain hidden and the consequences if those emerge in the future.
Ultimately, AI adoption by professionals for the provision of their services comes with equal opportunity and risk. Any business engaging in AI should be looking to manage the risk by scrutinising their internal safeguarding processes and liaising closely with their insurers at each step.
Download PDF



