From Data Hoards to Information Moats 

By Rupali Patel Shah, Head of Legal Solutions, DiliTrust

Fall is upon us, and the wild and crazy ride that has been 2026 continues to be twisty and full of thrills. Society is still grappling with the Age of AI and what it means for the future of work, society and the planet. Human displacement is already happening in some industries, although it remains unclear who the new lamplighters will be. Regulation of ownership, privacy and the veracity of information in the United States remains a patchwork, while the proliferation and use of data continues to expand. Most of us are scrambling to keep up.

The LegalTech world is no different. The initial hype around AI in legal has settled. The conversation is no longer only about efficiency gains, seat adoption or how quickly a tool can generate an answer. Increasingly, it is about the quality of the output and the return on investment for the client.

That shift is overdue. It is also not clear that the industry is ready to answer those questions, because the lack of information governance strikes again.

Every technology vendor encounters the same two dilemmas when it comes to adoption and metrics. First, the client’s current data state is often unstructured, unknown, stale, inaccurate or nonexistent. That means the technology is being implemented on a deeply questionable foundation. Second, it is rare that a client can describe or qualify its current state. Most of the time, the starting point is simply chaos. There is no reliable baseline against which to measure what automation changed, what it improved or what it made worse.

And without a measurable starting point, you cannot establish return on investment. So, yes: it really is the data that matters.

The honeymoon was short—and it is over

Most organizations experienced the same thing when they first introduced an AI tool. The initial results were exciting. The system was fast. The output looked polished. A task that used to take an afternoon appeared to take minutes. Everyone began imagining what else the tool could do.

That was most of 2025 and a good part of 2026.

Eventually, reality set in. The answer could not be trusted without checking the source material. The underlying records were incomplete or stale. Prompts became longer because users were trying to compensate for missing context, which consumed more tokens and increased cost. More people started using the system and performance slowed, cutting into the efficiency argument.

Then there was the output problem. Unverified summaries, analyses and drafts were saved somewhere: in a knowledge base, a matter file, a contract repository or a shared drive. That material was then shared back into the data moat, which became polluted with content that might be plausible, useful or completely wrong. It was not always clear which.

That is the vicious cycle: poor data produces irrelevant output; the irrelevant output becomes another data store; the additional data makes it harder to find what matters; and the next tool is expected to solve the problem.

AI did not create the data-quality problem, but it amplifies the problem and is adding to it at a rapid clip.

The latest shift in the market

For the past few years, much of the legal AI conversation focused on adoption, seats and time saved. Now the language is shifting toward interoperability, knowledge, workflow and return on investment. The market is beginning to move from efficiency theater toward accountability for whether a product can actually be used, trusted and tied to a business outcome.

Harvey, for example, now describes itself as “multi-model by design” and argues that customers should be able to route work among different frontier models rather than depend on a single provider. It has also emphasized connections with document and agreement systems, while its ROI framework makes the point that time saved matters only when it produces better client service, stronger margins or more meaningful work.

Those of us who have been lawyers for some time may find this shift obvious—and frankly, overdue. Once every firm and legal department can access capable models, the model itself will not provide much of an edge for long. The difference will come from the quality of the organization’s information, the processes built around it and the judgment applied to the result.The institutional knowledge behind complex legal work does not exist on the internet. It lives in people, precedents, risk tolerances, prior decisions and the way an organization actually operates.

Slowly but surely, LegalTech is shifting its focus toward the data layer: owning or controlling the source, establishing a defensible information moat, and providing agents that support adaptation, consistency and token containment. For vendors that began as wrappers around frontier models, the reality is setting in that technology is difficult to monetize when the underlying capability becomes a commodity. Differentiation has to come from the quality of the data, the context surrounding it and the way the technology fits into the work.

It has always been the data

Since the start of the Digital Age, companies have collected enormous amounts of data and treated the accumulation itself as a source of corporate value. They maintained these hoards because someone, someday, might want to access them. Then came the Age of AI, and the hoards began to seem strategically useful. Algorithms train on data, so more data must be better. Right?

Not necessarily. Much of that data is stale, corrupted or off limits because the organization does not have the rights to use it. It is scattered, unstructured and difficult to interrogate. Algorithms may train on data—and lots of it—but businesses run on information. Bad data produces bad information.

There is a scene in AMC’s The Audacity that captures the economics of the data business and explains the why behind data hoarding with uncomfortable precision. The CEO of a fictional data-mining company challenges the executives of a self-driving-car company over the personal information its vehicles collect. The drivers have technically consented, but the consent is buried in a 97-page document printed in seven-point font. Without accepting it, the car does not function properly.

One of the executives responds:

“The truth, we don’t sell cars. We sell data collection devices on wheels with heated seats. Cars are expensive. Data has profit margins that you can drive a car right through.”

It is funny because it is close enough to reality to make us uncomfortable. Data is produced as a byproduct of commerce and ordinary life. We generate it when we drive, shop, travel, sign a contract, attend a meeting or click a link. In many cases, participating in the digital economy requires us to hand it over.

There are important questions about privacy, security and meaningful consent in that exchange. We are already reckoning with some of the more insidious uses of personal information. That deserves its own discussion. For now, I am interested in what the company receiving the data actually possesses: usually, a very large pile of raw material. Raw material becomes valuable only after someone does the work of validating it, organizing it, connecting it to other facts and making it usable. That is the difference between data and information.

Information is data with context, structure, provenance, maintenance and meaning. It is the knowledge, analysis, statistics, creativity and know-how that allow an organization to operate and compete. Data may have almost no marginal cost to reproduce, but trustworthy information is not free. It requires investment.

ROI requires a baseline

The same data problem extends to the other measure of accountability LegalTech is now being asked to provide: return on investment.

Too many vendors still encourage large rollouts without helping customers establish a baseline. If you cannot describe how long the work takes today, what it costs, how much volume it involves, where it breaks down and how much rework it creates, you will struggle to prove that the new tool improved anything.

Usage is evidence that people logged in. It is not proof of productivity. Productivity is what people can accomplish. ROI asks a harder question: compared with the current state, what did the organization gain, what did it save and what did it avoid?

That means understanding the baseline before implementation. How much would the department otherwise spend? How long would the task take? How often would the work need to be corrected? What decisions would be delayed? What risks would remain invisible? Without those answers, both vendors and clients are guessing.

The baseline also has to account for the full cost of ownership: implementation, integrations, data migration and normalization, privacy and security reviews, training, change management, token and consumption charges, model evaluation, human validation, maintenance and governance. A product that reduces drafting time but creates a new review queue may be faster at one step and more expensive overall.

A genuinely consultative vendor should invest enough upfront to understand how the client works, what information it has, what condition that information is in and what success would look like. Building for use means starting there. It means designing for the enterprise, not simply demonstrating a compelling feature.

Make AI investment measurable

Get a realistic ROI based on industry standards when you partner with DiliTrust.

DiliTrust Suite ROI Calculator
Calculate your potential ROI

Consider this from the perspective of a legal department and the way it protects the enterprise. A contract buried in a shared drive is data. A reliable record of its obligations, owners, deadlines and commercial consequences is information. A set of board minutes is data. A traceable history of decisions, delegated authority and unresolved actions is information. A long email chain is data. A current, verified account of the decision made, the reason for it and the person responsible for acting is information.

The difference comes from context, structure, maintenance and judgment. It also comes from the discipline to decide which information matters. Knowledge creation requires human intervention. Otherwise, we just create restatements of the same data.

This is why legal departments are not simply another buyer of AI productivity tools. Their literal job is to turn facts, rules and circumstances into information that decision-makers can use. Presenting relevant information, in a timely manner, to the right decision-maker is the work. We have one primary tool in trade: judgment. Clients trust lawyers to evaluate facts, understand context and make a recommendation when the answer is unclear. A model may help us reach that point faster, but responsibility for the judgment remains with us. The differentiator for the profession is not access to a particular model. It is the quality of the judgment and the trust clients place in it. That is also the differentiator for LegalTech – clients will license technology and procure services from vendors they can trust because they put their integrity on the line every time they rely on output from a tool

The LegalTech industry is finally being pushed to take those questions seriously. That is a good thing. The value of AI is not simply that it can make work faster. Its value is that it can help people analyze more information, find what matters and turn data into something an organization can use.

More generated content does not create more knowledge. More data does not mean the organization knows more. And a faster answer is not necessarily a better answer.

What should legal teams assess before connecting generative AI to matter, contract or board data?

They should assess whether the underlying information is accurate, current, permissioned and traceable before connecting it to an AI workflow. The review should cover data ownership, access controls, provenance and how outputs will be validated, because a model can scale the weaknesses of the information layer as quickly as it scales useful insight.

How should legal teams govern AI-generated content before it re-enters the knowledge base?

AI-generated summaries, analyses and drafts should remain clearly marked as unverified until a responsible person checks the source, confirms the context and approves reuse. That validation step protects the knowledge base from becoming a second layer of unreliable information, especially when outputs are copied into matter files, repositories or shared drives.

When is a multi-model legal AI strategy worth the added governance effort?

A multi-model strategy is worthwhile when different models materially improve the quality, cost or speed of distinct legal workflows, and the organization can preserve consistent permissions, auditability and validation across them. If routing work between models adds complexity without a measurable business outcome, a single governed approach may be easier to trust and manage.

What evidence should a legal department require before expanding an AI pilot?

Before expanding, the department should require evidence from representative work: baseline task time, volume, correction rates, review effort, total implementation cost and the risks that remain. Adoption metrics alone are not enough. A credible case connects measured change to a business outcome, such as reduced rework, faster decisions or more reliable service.

Avatar photo
Author

Rupali Patel Shah

Head of Legal Solutions & Alliances for North America, DiliTrust

A strategist with a human touch, Rupali brings equal parts precision and heart to every conversation. Twenty years in — across in-house legal, Big Four, and legal tech — she's earned her fluency where AI governance, data privacy, and enterprise contracting collide, and she turns that complexity into bold, practical solutions. She delivers them from boardrooms to keynote stages, with the humility and humor that make her a force people actually want to follow.