Essay · AI & Defensibility

Data Is No Longer an Asset by Default: Rethinking Liability in the AI Era

For a decade, "data is the new oil" was the boardroom metaphor of choice, and it drove a real and mostly reasonable instinct to collect and retain as much data as an organization could reasonably manage. That instinct is now actively dangerous in ways most data strategies haven't caught up to, and the shift has moved faster over the past eighteen months than most executives have registered. Every dataset an organization holds is simultaneously a potential asset and a standing liability, and the AI era has moved the liability side of that ledger substantially, in ways that are now showing up in courtrooms rather than just in theoretical risk assessments.

The shift, made concrete by recent litigation

The shift is straightforward once you see it laid out. Data sitting inert in a warehouse is a liability mostly in the event of a breach — real risk, but a bounded and insurable one that most organizations have processes for. Data actively feeding AI systems is a liability continuously, because every model trained or fine-tuned on it inherits its quality problems, its bias, its provenance gaps, and its licensing ambiguity, and then expresses those problems at the speed and scale of automated decision-making rather than at the speed of individual human error.

This has stopped being a hypothetical risk. In February 2025, the Thomson Reuters v. Ross Intelligence case produced the first final judgment on AI training data copyright, with the court finding that Ross Intelligence's use of Westlaw headnotes to train its legal AI was not fair use, focused specifically on the fact that Ross's product competed directly with Westlaw's own research service.[1] Critically, that ruling established that using an AI tool trained on infringing data can create downstream legal liability for the company deploying the tool, not only for the company that built it. In May 2026, five major publishers and a well-known novelist filed a proposed class-action copyright suit against a large technology company over AI training data, seeking an injunction requiring destruction of the allegedly infringing copies along with monetary damages.[1] Separately, Anthropic agreed to pay $1.5 billion to settle claims from authors whose books were used, without permission, to train its models.[1] And one contract analysis found that ninety-two percent of AI vendor agreements claim data usage rights that extend beyond what's strictly necessary to deliver the contracted service, compared to a broader software market average of sixty-three percent.[1] A biased or improperly sourced dataset sitting quietly in a spreadsheet was always a bounded problem. The same dataset feeding an automated decision system at scale, or licensed into a vendor's model training pipeline without anyone reading the contract closely enough to notice, is a fundamentally different order of exposure — legal, reputational, and increasingly regulatory all at once.

A question most data strategies have never actually asked

This means the pre-AI data strategy question — how do we collect and retain more — needs to be replaced with a considerably harder one: which of our datasets are genuine assets under this new use, and which have quietly become liabilities that were never re-underwritten for it. Most organizations have never run this exercise, because the data in question was accumulated under a retention policy written for an entirely different threat model, years before anyone planned to feed it into a decision-making system or license it into a vendor's training pipeline.

Three practical shifts

Three practical shifts follow from taking this seriously, and none of them are exotic, which is part of why the organizations that have made them are pulling ahead quietly rather than dramatically. First, data provenance needs to become a first-class governance artifact, not a data engineering afterthought buried in a technical wiki nobody outside the team reads. The organization needs to be able to answer, for any dataset currently feeding a model, where it came from, what it's licensed for, and who is accountable for its quality, on demand and in writing, ideally before a regulator or plaintiff's attorney asks the same question in a less friendly setting.

Second, retention policies need to be re-underwritten specifically for AI use cases, because "we might need it someday" is a considerably weaker justification when "someday" now plausibly means "feeding an automated decision system that affects real people" rather than "answering a hypothetical future business query." The bar for justified retention has quietly gotten higher, even though most retention policies haven't been rewritten to reflect that.

Third, the data quality investment that used to be treated as optional — the genuinely unglamorous work of cleaning, labeling, and documenting data — has become close to a prerequisite rather than a nice-to-have. AI systems don't tolerate messy data the way skilled human analysts historically learned to compensate for it through judgment and context. They encode the mess directly into the model and then scale it, at exactly the speed and volume that turns a quiet internal data quality problem into a public and legally exposed one.

A useful historical parallel

This pattern has a clear precedent in how financial data governance evolved after Sarbanes-Oxley in the early 2000s. Before that regulatory shift, financial data quality was largely treated as an internal operational concern, handled informally by whichever team happened to own a given spreadsheet. Afterward, provenance, controls, and documented accountability became legally mandated, board-level concerns almost overnight, and organizations that had already been managing their financial data with that level of rigor absorbed the transition easily, while organizations that hadn't spent years and meaningful money retrofitting controls onto systems that were never designed to support them. Data governance for AI is very plausibly headed down the same road, on a considerably faster timeline given how quickly the relevant litigation has already materialized.

The honest counter-case

It's worth resisting the opposite overcorrection, which is treating all data as radioactive and retreating into excessive caution that forfeits real competitive advantage. Well-governed, properly licensed, high-quality data remains a genuine and durable asset, and organizations that overcorrect into hoarding nothing and using nothing distinctive will find themselves with AI systems trained on the same generic, widely available data as every competitor, which is its own strategic failure, just a quieter one than a lawsuit. The goal isn't less data. It's data whose provenance, licensing, and quality the organization can actually defend, in writing, under scrutiny.

The organizations approaching this well in 2026 aren't the ones with the most data. They're the ones who can state, for any dataset currently feeding a model, exactly what it's worth, what it costs to hold, and what it's actually licensed to do — and who have stopped treating "we have a lot of data" as a strategy in its own right. That phrase alone is closer to a risk disclosure than a competitive advantage this year. Could your organization answer that question today, for its highest-stakes dataset, in writing?

Sources

  1. AI training data litigation, including Thomson Reuters v. Ross Intelligence, the May 2026 publisher class action, the Anthropic settlement, and contract rights analysis, as summarized in AI Vortex, "AI Copyright Training Data Lawsuits 2026: Status, Timeline, Risk." https://www.aivortex.io/legal/guides/ai-copyright-training-data-2026-landscape/

Juan Vegarra is the author of An Outsider's Playbook (forthcoming). The views here are his own. More essays · Write me