How are mid-size law firms measuring whether AI spend pays off?

Most are not, at least not with numbers. The majority run on user surveys, adoption rates, and hallway feedback, and over half admit their real reason for spending is competitive pressure rather than a return they have modeled. A small group has landed on hard metrics, and they all got there the same way: they started from the specific outcome they wanted AI to change, then measured that one thing.

What mid-size firms reportShare
Justify AI spend by competitive pressure, not a modeled return52%
Rely on user surveys or feedback as their main proxy for value48%
Track at least one hard, quantifiable AI metric16%

How we know this

Sidebar puts one question a week to legal management professionals at firms of 10 to 200 attorneys. Members are verified by title, employer, and firm size before they are admitted, and every reply is private. This page draws on every Sidebar cycle that has touched this question, and it is updated as new replies come in. We publish patterns across the group, never individual firms, and only once at least five members have replied to that question. Full methodology at gosidebar.ai/methodology.

Anecdote is standing in for evidence

The most common way firms judge their AI spend is soft: user surveys, focus groups, and informal feedback about whether people like the tools. As a first read that is reasonable, and it is honest about where most firms actually are. The problem starts when budgets grow and partners begin asking harder questions, because a good story is difficult to defend in a budget review. The firms that stay comfortable are the ones already converting those stories into a number. It is worth being precise about why this matters, because it is not really a measurement failure. Soft feedback is a perfectly good instrument for the question it answers, which is whether people like the tool and whether they will keep opening it. It becomes a governance problem only at the point where the spend is large enough that someone has to defend it to the partnership, and the gap between those two moments is usually about a year. Firms that use that year to convert the story into a number arrive at the conversation with something to say.

Usage is a floor, not proof

Adoption rate is the second most common proxy, and it is a real signal that a tool has not been abandoned. But continued use conflates habit with value. People keep using plenty of things that are not moving the business, and no partner reviewing the line item will accept high login counts as evidence of return. Usage tells you a tool is alive. It does not tell you it is paying off. The distinction matters most at renewal, when the only number on the table is a login count and the question on the table is whether to spend the same money again. A high number there proves the firm has built a habit, which is worth something, but it is silent on whether the habit is worth the invoice.

The firms with real numbers started narrow

The minority tracking genuine outcomes did not begin with a better dashboard. They picked one role and one result, then measured only that: paralegal overtime against billable hours, or a target such as a 5% lift in billable hours year over year. In practices where the outcome is naturally visible, a stronger settlement demand or a faster path to resolution does the measuring for them. The lesson is not which metric to copy. It is that the number follows a clear question about what AI was supposed to fix. The route those firms took is worth copying even though the metrics are not: they asked the people using the tool what they most wanted it to change, took the answer that came up most, and watched only that. Starting from the complaint rather than from the capability is what makes the resulting number defensible, because it measures the thing the firm set out to buy in the first place.

The learning curve is hiding the answer

There is a trap sitting underneath all of this that almost nobody accounts for. Firms that did measure early are finding that tasks take about as long as they used to, and concluding the tool is not working. The more likely explanation is that prompting is a skill and their people are still acquiring it. A measurement taken in month two is mostly measuring how well the firm has learned to ask, not what the tool can do once they have. That produces a specific and expensive failure: a firm quietly writes off a capable tool on evidence gathered before anyone was competent with it, and the renewal decision gets made on a number that was never about the software. The firms getting clean reads are the ones that let the curve flatten before they started counting, which usually means waiting a good deal longer than feels comfortable when a partner is asking what the money bought.

The wider record says the same thing

This is not unique to mid-size firms, or to law. A 2026 Thomson Reuters Institute survey of more than 1,500 professional-services respondents across 27 countries, spanning legal, tax and accounting, corporate functions, and government, found that only 18% knew their organization was tracking the return on its AI tools in any form. That is the same gap our members describe, measured on a much larger sample and well outside legal. What mid-size firms add is the view from the middle. The spend is large enough for partners to notice it and the firm is small enough to have no analytics function to hand the question to. Firms of every size struggle to measure this. Mid-size firms have fewer people to put on the problem, and a shorter runway before someone asks what the money bought.

Source: Thomson Reuters Institute, 2026 AI in Professional Services Report

Naming the bet takes the pressure off

Over half of firms are spending because they are worried about being left behind, not because they have modeled a return, and there is nothing wrong with that. It is a strategic bet on staying competitive, and it is defensible as long as the firm names it as one. The trouble comes from dressing a competitive decision up as an ROI calculation it was never based on. Call the bet a bet, and you can stop forcing a measurement story that does not yet exist. Then, when you are ready, pick the one outcome you actually want and start counting.

What to do with this

Do not start with a dashboard. Ask the people who use the tool what they most wanted it to fix, pick the single answer that comes up most, and watch only that. The firms with real numbers all took that route, and none of them began from a measurement framework. Then hold two dates in your head: the one where you captured a baseline, which has to be before rollout or it is not a baseline, and the one where you start counting, which should be late enough that prompting skill is no longer the variable you are measuring. In between, keep collecting the surveys and the hallway feedback. They are a perfectly good instrument for the question of whether people like the tool. They are simply not an answer to whether it paid for itself, and the moment a partner asks the second question is a bad moment to discover you have only measured the first. If none of that is possible yet, the fallback is honest and it is not nothing: write down, in one sentence, what you expected the tool to change. A firm that can produce that sentence a year later is in a far stronger position than one that cannot, whether or not it ever got to a number.

I don't know if they did amazing work with that, or if they just burned through tokens. Usage does not equal proficiency.

Debbie Foster, The perils of just turning it on.

Members get more. The breakdown by firm size, practice area, and tech stack. A new question every Tuesday, the outlier answers that cut against the pattern, and most weeks an expert take with one clear next step.

Frequently asked questions

How do law firms measure ROI on AI tools?
Most do not measure it with numbers at all. The common approaches are user surveys, adoption and login rates, and informal feedback about whether people like the tool. A minority track a hard outcome, and those that do almost always picked one role and one result rather than building a general dashboard.
What is a good AI ROI metric for a law firm?
The firms that got to a real number chose a single outcome tied to a single role: paralegal overtime against billable hours, time from intake to engagement letter, or a target lift in billable hours year over year. The metric matters less than starting from a clear statement of what the tool was supposed to change.
Why can most law firms not prove their AI spend is working?
Because the baseline was never captured. Measurement is usually attempted after rollout, when there is nothing to compare against, and firm-wide metrics are too noisy to isolate the effect of one tool. Adoption gets substituted for value because it is the number that happens to be available.
Is it acceptable to buy AI without a business case?
It is defensible as long as the firm names it for what it is. A competitive bet on not falling behind is a legitimate strategic decision. The problem is presenting that decision as an ROI calculation it was never based on, which is difficult to defend the first time a partner examines it closely.