The AI Industry Wants Congress to Create the Next 100-Year Radio Loophole

“Formal property’s contribution to mankind is not the protection of ownership… Property’s real breakthrough is that it radically improved the flow of communications about assets and their potential.”

Hernando de Soto, The Mystery of Capital.

Musicians and other creators are unfortunately familiar with many efforts by big business to extract the economic value of their authorship through expansive free-riding copyright loopholes that pretend property rights don’t exist. The current AI crisis did not originate with Big Tech—they learned it from Big Radio.  I distinctly recall having lunch with a Big Tech Washington lobbyist for XM radio (pre-merger) who had just found out that broadcast radio didn’t pay sound recording performances and wanted that same deal for satellite radio.  I had to put the quietus on that pronto.  And they didn’t even know how close they came to disaster. Sheesh.

In case you were wondering, Congress modernized copyright law in 1995 through the Digital Performance Right in Sound Recordings Act.  The 1995 law created the statutory framework that launched licensed webcasting while preserving the archaic terrestrial radio performance loophole—preserved due to lobbying by Big Radio.

For decades, terrestrial AM/FM broadcasters have relied on a statutory copyright exception that allows them to broadcast sound recordings without compensating the featured artists, session musicians, and backup singers whose performances attract listeners, or the record companies who bear the substantial costs of discovering, recording, marketing, and promoting those works. Despite years of bipartisan efforts to end that free ride through legislation like the American Music Fairness Act (AMFA) and its predecessor bills, broadcasters have vigorously defended the exemption with overwhelming money and utilization of the very broadcast license they abuse to feather their nests.  We have put excellent witnesses in front of Congress only to be outspent by smarmy swamp creatures from the National Association of Broadcasters.

AI disputes echo that familiar pattern. In the end, it all comes down to vast wealth accumulated through safe harbors of one kind or another.  Instead of relying on a terrestrial performance exemption, AI companies advance absurd interpretations of fair use and text-and-data-mining doctrines to justify the uncompensated use of stolen works for commercial model training “because China.” They use influence peddlers like White House AI Viceroy David Sacks to try to sneak retroactive safe harbors into the law through Congress in the form of groundless federal preemption of state and local regulation or executive orders that are clearly bought and paid for under the guise of “data center factories” which are not factories at all.   Although the legal theories differ between AI and broadcasting, the economic consequence is remarkably similar: sweeping commercial enterprises seek to build profitable businesses by lobbying or litigating (two sides of the same King’s shilling) to expand exceptions to the exclusive rights Congress granted creators, while forcing artists, musicians, writers, journalists, film makers and photographers to absorb the resulting loss in value.

That concern is no longer theoretical. In a recent Bloomberg podcast, SoundExchange President and CEO Michael Huppe—whose organization distributes more than $1 billion annually in digital performance royalties derived from rights created by that market-making 1995 legislation—described AI as “something that has a lot of danger, but also a lot of potential.” But he cautioned that “we need to make sure that human creators are protected” and that “there need to be guardrails so that [AI] doesn’t steamroll over the whole creative industry.” I couldn’t agree more. Rather than treating property rights as obstacles to AI, Congress should remember Hernando de Soto’s lesson that clearly defined ownership creates wealth—a principle it proved when licensing sound recordings gave birth to the webcasting industry largely thanks to SoundExchange and the infrastructure it brings to the table.

Huppe’s concerns are rooted in measurable economics rather than speculation. Streaming now accounts for approximately 85% of U.S. recorded music revenue, and streaming services distribute a finite, shared royalty pool among eligible recordings. Huppe noted that some services report receiving roughly 75,000 new recordings every day, with reports suggesting that more than 80% are AI-generated.

Whether those estimates ultimately prove higher or lower, the underlying economic principle is unavoidable: every AI-generated recording entering the marketplace competes for listener attention and, if streamed, competes for a share of the same finite, shared royalty pool. Huppe also warned that AI facilitates streaming fraud, allowing bad actors to generate AI recordings, deploy bots to inflate plays, and “siphon away payment from the pipeline that would otherwise go to real artists and real record labels.” His conclusion was unequivocal: “It’s fraud, basically. Straight-up fraud.”

Moreover, generative AI takes legitimate recorded performances to create competing works substituting for the originals themselves. This economic effect echoes Judge Vince Chhabria’s observations in the Kadrey v. Meta books litigation, where he suggested that flooding markets with AI-generated works competing against originals could constitute the type of market harm that would block a fair use defense to copyright infringement.

The explosion of AI-generated music that Mike Huppe cites therefore provides strong evidence of repeatable and measurable market harm identified by Judge Chhabria. Every AI-generated stream competes for listener attention while simultaneously reducing each human artist’s share of a finite, shared royalty pool. Unlike speculative claims of future injury, this dilution can be observed, quantified, and modeled using actual streaming and royalty distributions.

The economics become even more troubling when combined with large-scale scraping. As we have seen litigated in the cases against Udio, Anthropic and Meta (and I think will continue to see proven through all of the AI models including Suno),  AI has trained on enormous quantities of illegally acquired works without obtaining licenses or compensating the creators whose recordings, performances, writings, images, and other expressive works supplied the raw material that makes those models commercially valuable.  Sound familiar?

The same creative ecosystem that furnished the training corpus is then required to compete against a cascading and endless supply of AI-generated outputs while receiving no payment for either the training use or the resulting competition. Worse yet, because nothing says freedom like getting away with it, AI platforms connected to Google, Facebook and Amazon are so used to ignoring copyrights in their day jobs that they clearly planned to ignore our rights.

In music, the effect is especially stark: the recordings that taught music-generation systems how to produce theoretically commercially appealing songs also become the works displaced by those outputs in the marketplace. Creators are effectively asked to finance their own displacement. They suffer a double economic injury—first, uncompensated exploitation of their works to build commercial AI systems, and second, measurable erosion of their share of a finite, shared royalty pool as AI-generated recordings compete for the same listeners and revenues that streamers like Spotify seem unable to stop from invading the ecosystem.

Because of the insane pool allocation formula used for streaming mechanical royalties on interactive services like Spotify, Amazon, Apple and Deezer, songwriters are also subject to the same kind of dilution as artists.  Hopefully the Copyright Royalty Judges will address this new humiliation in the current statutory rate proceeding and clearly state that AI works are not eligible for the statutory license under Section 115.

This measurable dilution also helps illustrate the broader market-flooding concern identified by Judge Chhabria. If AI-generated outputs systematically occupy the same commercial markets as human-created works, reducing revenues through sheer volume rather than direct substitution alone, then streaming provides one of the first empirical laboratories for proving market harm for “the effect of the use upon the potential market for or value of the copyrighted work.”  Because streaming royalties are transparent, pooled, and data-driven, music offers unusually strong evidence that AI-generated competition can inflict repeatable, measurable, and scalable economic injury. If courts follow Judge Chhabria in recognizing this analysis, the same analytical framework could extend beyond music to books, journalism, visual art, film, software, and other creative industries in which AI-generated outputs compete for the same audiences, revenues, and licensing opportunities as human creators.

Against that backdrop, the American Music Fairness Act is no longer simply a current solution to a decades-old copyright reform proposal. If AI companies are correct that generative AI will place unprecedented pressure on the economics of human creativity, then Congress should strengthen—not further weaken—all of the economic foundations supporting human creators. AMFA would finally require terrestrial broadcasters to compensate featured artists, session musicians, and vocalists for the use of their sound recordings, just as streaming and satellite radio already do. It would also unlock reciprocal foreign performance royalties that American performers currently forfeit because the United States remains an international outlier. 

At a moment when AI is intensifying the struggle for creative labor to survive even while platforms seek broad legal exceptions for uncompensated training through lobbying and executive orders, eliminating one of copyright law’s oldest uncompensated uses would send an important signal: the future of artificial intelligence should not be financed by the continued erosion of the livelihoods of human creators.

The AI industry’s habit of predicting existential harm while aggressively commercializing the same technology presents a profound ethical contradiction that Professor Cal Newport calls “doom trolling” in a recent New York Times post.  This leads to a conclusion that AI companies cannot credibly claim their technology poses existential risks while continuing to accelerate its commercialization without meaningful restraint.

Newport gives this example reminiscent of my personal favorite, the exploding gas tank in Ford Pintos (not to pick on Ford):

Imagine if the Ford Motor Company put out a report saying that it feared its popular F-150 trucks might soon start bursting into flames, but that there was nothing the company could do about it because automotive technology was too inevitable and important to slow down. You’re probably struggling to picture this scenario because no reasonable consumer product company would ever act like this. 

The A.I. companies could start behaving the same way. To do so would require that they stop treating A.I. like some inevitable force that they’re struggling to steward. It’s not. It’s a collection of specific tools that these companies are choosing to design and sell according to specific business plans. Accordingly, they need to talk about their offerings like any other consumer product. This means explaining clearly whom these products are for, justifying their benefits and, critically, taking full responsibility for any harm they might cause. Just because A.I. currently enjoys a high-tech sheen doesn’t make it exceptional with respect to common-sense safety standards.

If these A.I. companies insist on continuing to pretend that they’re merely stoic observers of an unavoidable dystopian future, then perhaps it’s time to force the issue. As consumers, we can refuse to play the doom-trolling game. Next time Anthropic releases a dire report, or Sam Altman’s voice cracks as he imagines the disruption that OpenAI is unleashing, we can pivot back to the pragmatic: “OK, but what benefits am I getting by spending $1,000 a month on tokens?” If they continue to ratchet up the doom, then perhaps it’s time to transform dread into ridicule: The earnest pseudoscience of Anthropic’s white papers already borders on satire. The current zeitgeist surrounding A.I. encourages a fretful submission to these tech leaders, but this could rapidly change.

The AI industry cannot have it both ways. It cannot warn that generative AI will fundamentally transform—or even eliminate—millions of creative jobs while simultaneously insisting that the law should expand uncompensated access to the very works that make those systems possible. If AI companies genuinely believe their own predictions, then the appropriate public policy response is not to weaken copyright, broaden fair use, or create new exceptions for commercial training. It is to reinforce every remaining economic support for human creativity. 

The evidence emerging from music streaming already demonstrates why. AI-generated works are not merely theoretical substitutes; they compete for attention, streams, and revenue, measurably reducing each creator’s share of a finite, shared royalty pool. That provides some of the clearest real-world evidence yet of repeatable market harm from generative AI at commercial scale. Congress should take note. The question is no longer whether creators deserve compensation for their work. It is whether the United States will choose to finance the AI economy by systematically eroding the economic incentives that have sustained human creativity for generations—or whether it will insist that technological progress, like every other successful industry before it, pays its own way.

Perhaps the greatest lesson of the American Music Fairness Act is not about radio at all. It is about refusing to repeat yesterday’s policy mistakes in tomorrow’s technology. As Mike Huppe observed on Bloomberg, Congress should not be creating new copyright exceptions while it is still trying to fix old ones. That warning applies with even greater force to artificial intelligence. If policymakers know that generative AI is likely to place extraordinary pressure on the economics of human creativity—as many AI companies themselves readily acknowledge—then the answer cannot be to expand uncompensated uses of creative works in the name of innovation and unintended consequences be damned.

The webcasting revolution showed what Hernando de Soto long argued: respecting property rights doesn’t kill innovation—it gives innovators the legal foundation to build sustainable markets. AMFA is a cautionary tale: a narrow copyright exception adopted decades ago has deprived generations of American performers of compensation and remains difficult to unwind. Congress should learn from that history, not repeat it. The AI economy should be built by paying for the creative works that make it possible and respecting the rights of all creators—not by creating another exception that future generations will spend decades trying to reverse and an entrenched bureaucracy of the richest corporations in commercial history will oppose with all the resources they can muster.

The AI Subsidy Is Over. Or Maybe It’s Just Beginning.


The current narrative says the “AI subsidy era” is ending. Prices are rising. Rate limits are tightening. Ads are creeping in. Enterprise tiers are replacing all-you-can-eat plans. In short: users will finally start paying what AI actually costs.

Haydon Field writing in The Verge tells us:

Earlier this month, millions of OpenClaw users woke up to a sweeping mandate: The viral AI agent tool, which this year took the worldwide tech industry by storm, had been severely restricted by Anthropic.

Anthropic, like other leading AI labs, was under immense pressure to lessen the strain on its systems and start turning a profit. So if the users wanted its Claude AI to power their popular agents, they’d have to start paying handsomely for the privilege.

“Our subscriptions weren’t built for the usage patterns of these third-party tools,” wrote Boris Cherny, head of Claude Code, on X. “We want to be intentional in managing our growth to continue to serve our customers sustainably long-term. This change is a step toward that.”

The announcement was a sign of the times. Investors have poured hundreds of billions of dollars into companies like OpenAI and Anthropic to help them scale and build out their compute. Now, they’re expecting returns. After years of offering cheap or totally free access to advanced AI systems, the bill is starting to come due — and downstream, users are beginning to feel the pinch.

That’s true but it’s leaving out a lot.

Yes, the consumer subsidy—venture-backed underpricing of inference—may be winding down. But the broader subsidy system that made AI possible isn’t going away. It’s expanding. Just ask President Trump.

To understand why, you have to go back to the last great digital disruption.

From P2P to Streaming to AI

Start with Napster.

P2P didn’t just enable infringement. It rewired expectations. It taught users that all music should be available, instantly, for free. Why? Because there was gold in them long tails. Forget about supply and demand, we had infinite supply so demand would take care of itself.

It’s for sale

Every artist, songwriter, label and publisher in the history of recorded music were not compensated for this shift. They were its involuntary financiers. Their catalogs created the demand, the network effects, and the user adoption that built the early internet music economy.

Streaming—think Spotify—didn’t reverse that logic. It formalized it. (Remember, streaming saved us from piracy and we should all be so grateful.) It actually transferred that involuntary financing from the p2p balance sheet to Spotify’s, and took it public.


Streaming platforms accepted a new baseline: the entire world’s repertoire must be available at all times, regardless of demand. That is a costly and structurally inefficient mandate, but it became the price of competing in a market shaped by P2P expectations. Licensing systems like the Mechanical Licensing Collective (MLC) were built to support that scale, but the underlying premise remained: total availability first, compensation second.

AI changes the game again.

AI Doesn’t Just Distribute Works. It Consumes Them.

P2P distributed music. Streaming licensed it. AI models ingest it.

That’s the critical difference.

Generative AI systems are trained on massive corpora that include copyrighted works, performances, and what we might call personhood signals—voice, style, tone, phrasing, and creative identity. These inputs are not just indexed or streamed. They are transmogrified (see what I did there) into model weights that can generate new outputs that compete with, mimic, or substitute for the originals.

So the role of the artist evolves:
    •    In P2P: unpaid distributor subsidy
    •    In streaming: underpaid inventory supplier
    •    In AI: uncompensated production input
That is not a marginal shift. It is a structural one.

The Real Subsidy Stack

When people say the “AI subsidy era is over,” they are usually talking about one thing: cheap access to compute.
But AI has always depended on a multi-layered subsidy stack:

    Creators – supply training data, cultural value, and identity signals without compensation or consent
    Users – supply prompts, feedback, and behavioral data that improve the models
    Communities – absorb land use, water consumption, and environmental costs
    Ratepayers – fund grid upgrades, transmission, and reliability for data center demand
    Venture capital – underwrites early losses to drive adoption and scale

The shift we are seeing now is not the end of subsidies. It’s a reallocation. Or as a cynic might say, it’s rearranging the deck chairs to hide the lifeboats.

Users may start paying more. But creators still aren’t being paid for training. Communities are still being asked to host infrastructure. And the physical footprint of AI is accelerating. Just ask President Trump.

The World Turned Upside Down

What makes this moment different is the scale of the buildout.
We are not just talking about apps anymore. We are talking about an industrial transformation:
    •    New data centers the size of small cities
    •    High-voltage transmission lines
    •    Water-intensive cooling systems
    •    Semiconductor supply chains
    •    And even discussions of new nuclear capacity to support compute demand

This is infrastructure on the scale of a national project, or more like national mobilization. But it is being built on top of a premise that has not been resolved: the uncompensated use of human creative work as training input.

That is the inversion: We are building power plants for systems that depend on not paying the people whose work makes those systems possible.

A Better Frame

The cleanest way to understand this is as a continuum:

P2P turned infringement into consumer expectation.
Streaming turned that expectation into platform infrastructure.
AI turns uncompensated authorship into industrial feedstock.

Or more bluntly:
The AI free ride is not ending. It is being re-invoiced. Users may now see higher prices. But the deeper subsidies—creative, environmental, and civic—remain off the books.

What Comes Next

If the industry is serious about “pricing AI correctly,” it cannot stop at compute.

It has to address:
    •    Compensation frameworks for training data
    •    Attribution and provenance standards
    •    Licensing models for style and voice
    •    Infrastructure cost allocation (who pays for the grid?)
    •    Governance of large-scale compute deployment

Otherwise, we are not exiting the subsidy era. We are doing what Big Tech lives for.

We are scaling it.

And this time, instead of a few server racks in a dorm room, we are building an global energy system around it.

What Would Freud Do? The Unconscious Is Not a Database — and Humans Are Not Machines

What would Freud do?

It’s a strange question to ask about AI and copyright, but a useful one. When generative-AI fans insist that training models on copyrighted works is merely “learning like a human,” they rely on a metaphor that collapses under even minimal scrutiny. Psychoanalysis—whatever one thinks of Freud’s conclusions—begins from a premise that modern AI rhetoric quietly denies: the unconscious is not a database, and humans are not machines.

As Freud wrote in The Interpretation of Dreams, “Our memory has no guarantees at all, and yet we bow more often than is objectively justified to the compulsion to believe what it says.” No AI truthiness there.

Human learning does not involve storing perfect, retrievable copies of what we read, hear, or see. Memory is reconstructive, shaped by context, emotion, repression, and time. Dreams do not replay inputs; they transform them. What persists is meaning, not a file.

AI training works in the opposite direction—obviously. Training begins with high-fidelity copying at industrial scale. It converts human expressive works into durable statistical parameters designed for reuse, recall, and synthesis for eternity. Where the human mind forgets, distorts, and misremembers as a feature of cognition, models are engineered to remember as much as possible, as efficiently as possible, and to deploy those memories at superhuman speed. Nothing like humans.

Calling these two processes “the same kind of learning” is not analogy—it is misdirection. And that misdirection matters, because copyright law was built around the limits of human expression: scarcity, imperfection, and the fact that learning does not itself create substitute works at scale.

Dream-Work Is Not a Training Pipeline

Freud’s theory of dreams turns on a simple but powerful idea: the mind does not preserve experience intact. Instead, it subjects experience to dream-work—processes like condensation (many ideas collapsed into one image), displacement (emotional significance shifted from one object to another), and symbolization (one thing representing another, allowing humans to create meaning and understanding through symbols). The result is not a copy of reality but a distorted, overdetermined construction whose origins cannot be cleanly traced.

This matters because it shows what makes human learning human. We do not internalize works as stable assets. We metabolize them. Our memories are partial, fallible, and personal. Two people can read the same book and walk away with radically different understandings—and neither “contains” the book afterward in any meaningful sense. There is no Rashamon effect for an AI.

AI training is the inverse of dream-work. It depends on perfect copying at ingestion, retention of expressive regularities across vast parameter spaces, and repeatable reuse untethered from embodiment, biography, or forgetting. If Freud’s model describes learning as transformation through loss, AI training is transformation through compression without forgetting.

One produces meaning. The other produces capacity.

The Unconscious Is Not a Database

Psychoanalysis rejects the idea that memory functions like a filing cabinet. The unconscious is not a warehouse of intact records waiting to be retrieved. Memory is reconstructed each time it is recalled, reshaped by narrative, emotion, and social context. Forgetting is not a failure of the system; it is a defining feature.

AI systems are built on the opposite premise. Training assumes that more retention is better, that fidelity is a virtue, and that expressive regularities should remain available for reuse indefinitely. What human cognition resists by design—perfect recall at scale—machine learning seeks to maximize.

This distinction alone is fatal to the “AI learns like a human” claim. Human learning is inseparable from distortion, limitation, and individuality. AI training is inseparable from durability, scalability, and reuse.

In The Divided Self, R. D. Laing rejects the idea that the mind is a kind of internal machine storing stable representations of experience. What we encounter instead is a self that exists only precariously, defined by what Laing calls ontological security” or its absence—the sense of being real, continuous, and alive in relation to others. Experience, for Laing, is not an object that can be detached, stored, or replayed; it is lived, relational, and vulnerable to distortion. He warns repeatedly against confusing outward coherence with inner unity, emphasizing that a person may present a fluent, organized surface while remaining profoundly divided within. That distinction matters here: performance is not understanding, and intelligible output is not evidence of an interior life that has “learned” in any human sense.

Why “Unlearning” Is Not Forgetting

Once you understand this distinction, the problem with AI “unlearning” becomes obvious.

In human cognition, there is no clean undo. Memories are never stored as discrete objects that can be removed without consequence. They reappear in altered forms, entangled with other experiences. Freud’s entire thesis rests on the impossibility of clean erasure.

AI systems face the opposite dilemma. They begin with discrete, often unlawful copies, but once those works are distributed across parameters, they cannot be surgically removed with certainty. At best, developers can stop future use, delete datasets, retrain models, or apply partial mitigation techniques (none of which they are willing to even attempt). What they cannot do is prove that the expressive contribution of a particular work has been fully excised.

This is why promises (especially contractual promises) to “reverse” improper ingestion are so often overstated. The system was never designed for forgetting. It was designed for reuse.

Why This Matters for Fair Use and Market Harm

The “AI = human learning” analogy does real damage in copyright analysis because it smuggles conclusions into fair-use factor one (transformative purpose and character) and obscures factor four (market harm).

Learning has always been tolerated under copyright law because learning does not flood markets. Humans do not emerge from reading a novel with the ability to generate thousands of competing substitutes at scale. Generative models do exactly that—and only because they are trained through industrial-scale copying.

Copyright law is calibrated to human limits. When those limits disappear, the analysis must change with them. Treating AI training as merely “learning” collapses the very distinction that makes large-scale substitution legally and economically significant.

The Pensieve Fallacy

There is a world in which minds function like databases. It is a fictional one.

In Harry Potter and the Goblet of Fire, wizards can extract memories, store them in vials, and replay them perfectly using a Pensieve. Memories in that universe are discrete, stable, lossless objects. They can be removed, shared, duplicated, and inspected without distortion. As Dumbledore explained to Harry, “I use the Pensieve. One simply siphons the excess thoughts from one’s mind, pours them into the basin, and examines them at one’s leisure. It becomes easier to spot patterns and links, you understand, when they are in this form.”

That is precisely how AI advocates want us to imagine learning works.

But the Pensieve is magic because it violates everything we know about human cognition. Real memory is not extractable. It cannot be replayed faithfully. It cannot be separated from the person who experienced it. Arguably, Freud’s work exists because memory is unstable, interpretive, and shaped by conflict and context.

AI training, by contrast, operates far closer to the Pensieve than to the human mind. It depends on perfect copies, durable internal representations, and the ability to replay and recombine expressive material at will.

The irony is unavoidable: the metaphor that claims to make AI training ordinary only works by invoking fantasy.

Humans Forget. Machines Remember.

Freud would not have been persuaded by the claim that machines “learn like humans.” He would have rejected it as a category error. Human cognition is defined by imperfection, distortion, and forgetting. AI training is defined by reproduction, scale, and recall.

To believe AI learns like a human, you have to believe humans have Pensieves. They don’t. That’s why Pensieves appear in Harry Potter—not neuroscience, copyright law, or reality.

Schrödinger’s Training Clause: How Platforms Like WeTransfer Say They’re Not Using Your Files for AI—Until They Are

Tech companies want your content. Not just to host it, but for their training pipeline—to train models, refine algorithms, and “improve services” in ways that just happen to lead to new commercial AI products. But as public awareness catches up, we’ve entered a new phase: deniable ingestion.

Welcome to the world of the Schrödinger’s training clause—a legal paradox where your data is simultaneously not being used to train AI and fully licensed in case they decide to do so.

The Door That’s Always Open

Let’s take the WeTransfer case. For a brief period this month (in July 2025), their Terms of Service included an unmistakable clause: users granted them rights to use uploaded content to “improve the performance of machine learning models.” That language was direct. It caused backlash. And it disappeared.

Many mea culpas later, their TOS has been scrubbed clean of AI references. I appreciate the sentiment, really I do. But—and there’s always a but–the core license hasn’t changed. It’s still:

– Perpetual

– Worldwide

– Royalty-free

– Transferable

– Sub-licensable

They’ve simply returned the problem clause to its quantum box. No machine learning references. But nothing that stops it either.

 A Clause in Superposition

Platforms like WeTransfer—and others—have figured out the magic words: Don’t say you’re using data to train AI. Don’t say you’re not using it either. Instead, claim a sweeping license to do anything necessary to “develop or improve the service.”

That vague phrasing allows future pivots. It’s not a denial. It’s a delay. And to delay is to deny.

That’s what makes it Schrödinger’s training clause: Your content isn’t being used for AI. Unless it is. And you won’t know until someone leaks it, or a lawsuit makes discovery public.

The Scrape-Then-Scrub Scenario

Let’s reconstruct what could have happened–not saying it did happen, just could have–following the timeline in The Register:

1. Early July 2025: WeTransfer silently updates its Terms of Service to include AI training rights.

2. Users continue uploading sensitive or valuable content.

3. [Somebody’s] AI systems quickly ingest that data under the granted license.

4. Public backlash erupts mid-July.

5. WeTransfer removes the clause—but to my knowledge never revokes the license retroactively or promises to delete what was scraped. In fact, here’s their statement which includes this non-denial denial: “We don’t use machine learning or any form of AI to process content shared via WeTransfer.” OK, that’s nice but that wasn’t the question. And if their TOS was so clear, then why the amendment in the first place?

Here’s the Potential Legal Catch

Even if WeTransfer removed the clause later, any ingestion that occurred during the ‘AI clause window’ is arguably still valid under the terms then in force. As far as I know, they haven’t promised:

– To destroy any trained models

– To purge training data caches

– Or to prevent third-party partners from retaining data accessed lawfully at the time

What Would ‘Undoing’ Scraping Require?

– Audit logs to track what content was ingested and when

– Reversion of any models trained on user data

– Retroactive license revocation and sub-license termination

None of this has been offered that I have seen.

What ‘We Don’t Train on Your Data’ Actually Means

When companies say, “we don’t use your data to train AI,” ask:

– Do you have the technical means to prevent that?

– Is it contractually prohibited?

– Do you prohibit future sublicensing?

– Can I audit or opt out at the file level?

If the answer to those is “no,” then the denial is toothless.

How Creators Can Fight Back

1. Use platforms that require active opt-in for AI training.

2. Encrypt files before uploading.

3. Include counter-language in contracts or submission terms:

   “No content provided may be used, directly or indirectly, to train or fine-tune machine learning or artificial intelligence systems, unless separately and explicitly licensed for that purpose in writing” or something along those lines.

4. Call it out. If a platform uses Schrödinger’s language, name it. The only thing tech companies fear more than litigation is transparency.

What is to Be Done?

The most dangerous clauses aren’t the ones that scream “AI training.” They’re the ones that whisper, “We’re just improving the service.”

If you’re a creative, legal advisor, or rights advocate, remember: the future isn’t being stolen with force. It’s being licensed away in advance, one unchecked checkbox at a time.

And if a platform’s only defense is “we’re not doing that right now”—that’s not a commitment. That’s a pause.

That’s Schrödinger’s training clause.