A demand to delete models built on their works
Following the infringement suit by the Seattle Times and Newsday against OpenAI and Microsoft, I trace what the claim of works used for training and the demand for model destruction asks of governance.
Read the original paper
Patrick Rho · AI Research
I read the Verge piece titled Seattle Times and Newsday sue OpenAI and Microsoft for infringement, which reports that the Seattle Times and Newsday have filed a copyright infringement suit against OpenAI and Microsoft. The shape of the claim is that models were built on the plaintiffs' works, and the remedy sought is the destruction of the models built on those works. The original is linked under the title.
It is a short report, yet I found myself rereading it. Who sues, who is sued, what grounds the claim, and what remedy is demanded all fit inside a few lines. One lawsuit ties training data, model weights, and newsroom copy into a single scene. The part that held my attention was the form of the demand rather than any dollar figure. The lead ask is removal of something built, with no damages number placed up front. In this post I want to follow that single line slowly, in the language of technology and governance.
Why one line about destruction stayed with me
The sentence that catches the eye first is the one saying the plaintiffs want models built on their works destroyed. My hand stopped on that line. Dispute coverage usually starts with money. How much is demanded, how the harm is calculated, that is where the story normally centers. Here the shape of the demand differs. The lead is a request to remove something made, with compensation absent from the front of the story. When the form changes, the reading changes with it. Attention moves from how much to what exactly.
I think the structure behind the number matters more than the number itself. Destruction reads as a refusal to treat the dispute as a transaction. It says the issue is the continuing artifact, to be ended rather than paid for and left running. To be fair, money claims may well sit alongside in the actual filing, and I will not speculate beyond what the report carries. Still, the placement of destruction on the front line invites interpretation. It signals that the plaintiffs see this as a problem of an operating output, something that keeps running, rather than a one time unauthorized use.
From a systems perspective, the next question is obvious. What precisely would destruction remove. Copies held for training, the trained weights, or the traces of specific works inside those weights are three different targets. The report does not draw these distinctions. It stops at the level of destroying models built on the works. This is where I would be careful not to over-read the sentence. Reading one reported line as a technical specification leads to overreach. Reading it as carrying no direction misses the signal. I take destruction at this stage as a sentence that widens the dispute to the whole artifact, rather than a concrete deletion procedure.
How to read two plaintiffs standing side by side
The second feature of the story is the pairing of plaintiffs. The Seattle Times and Newsday appear together. Seeing those two names side by side, what came to mind first was that newsrooms of different regions and characters had stepped into the same sentence. A regional daily of the Pacific Northwest and a daily rooted in the New York area now direct the same kind of claim at the same defendants. This is not two lookalike organizations joining forces. It is two organizations raised on different ground pointing at the same spot.
A pairing like this changes the message of the suit. It no longer reads as one shop's grievance. It reads as a claim that the output of the news trade itself became training material in bulk. The two plaintiffs cannot share identical circumstances. They cover different ground, serve different readers, and write with different texture. That they stand in the same sentence despite those differences suggests the logic beneath the claim reaches past individual sentences and touches news production in general. What I felt in that passage was a sense that the unit of dispute had climbed from single stories to the archive as a whole.
Still, a boundary needs drawing. The facts carried by the report are that these two outlets are plaintiffs, that OpenAI and Microsoft are defendants, and that the skeleton of the claim is use of works plus a destruction demand. Why the two joined, and what internal reasoning led there, is outside what the report gives. I do not want to fill that blank with imagination. I would rather leave the questions that can be asked. How the works of different newsrooms mix inside one training process, and why that mixing supports a claim that reaches past any single plaintiff. Those questions connect directly to the structure of training, which the next sections take up.

What the phrase built on their works points to
The expression that rewards rereading is built on their works. It sounded to me like a legal sentence and a systems sentence at once. In legal terms it points to use. In systems terms it points to dependence. It asks whether the model would look the same without those works. Inside a structure where changing the data changes the result, the claim that a particular body of data forms the base of the output is easy to grasp intuitively. Yet intuition and proof stand apart.
Unpacked simply, the training pipeline collects works, cleans them, splits them into tokens, and feeds them as material for numerical updates. Individual sentences are not stored whole. Statistical tendencies drawn from sentences seep into arrays of numbers called weights. This is a picture I often sketch for colleagues. The model resembles a student who absorbed a sense for sentence building from vast reading, rather than a student who memorized books cover to cover. The analogy helps, and its limits are plain. Not memorizing something is a different matter from not using it. The dividing line moves from memory to permission.
An odd question follows. Which layer does built on point to. Copying at collection, reflection in weights during training, and reproduction at output are separate layers. At collection, the question is whether files were copied. At training, the question is what trace that copying left in the weights. At output, the question is whether the model emits particular works. The report binds these layers into one sentence without separating them. I take that binding as the nature of a short report. Brief news has no room to unfold layers. Readers like us have to unfold them instead. The evidence required and the remedies available differ by layer, which is why the unfolding matters.

The narrow path between training and infringement claim
Copyright disputes over training walk a narrow path. On one side stands the grievance that works were copied and used without permission. On the other stands the counter that a computational process called training may be a different thing from copying. When I walk that path, I start by tidying terms. Copying moves a file as is. Training updates numbers using that file as material. Output draws fresh sentences from the trained numbers. Three acts live inside one pipeline, and they are not the same act. Claim and rebuttal often point at different acts while using the same words. That is why the conversation slips.
The plaintiffs' claim, as reported, stands at the start of the pipeline. Their works became material. For that claim to hold, the fact of becoming material has to be addressed first. What was collected, when, and through which route is a matter of records. I use the word records deliberately. Without provenance in the training pipeline, there is nothing to examine later, and every account of use becomes reconstruction after the fact. Where the crawl came from, which filters applied, which dataset version included it: unwritten, all of it turns into competing stories.
The next part is harder. Becoming material does not conclude infringement by itself. Permission, room for unlicensed use, and whether outputs substitute for the originals are separate inquiries. I want to avoid verdicts at this point. What the report carries is the skeleton of a claim, not a court finding. No defendant response and no judicial determination sits inside this story's scope. So this post does not take a side. It maps the terrain of the narrow path instead. Which layer will be contested, and what records each layer will require: sketching that in advance serves readers better than picking a winner.
Why destruction aims at structure rather than money
Returning to destruction, I can unpack why the demand points at structure. Money settles the past. Destruction settles the present and the future. Behind the word sits the view that the contested state continues for as long as the model keeps serving. I think this framing changes the character of the dispute. The venue moves from contesting a single unauthorized use to contesting an artifact that keeps operating.
When my team looks at a new architecture, we do not stop at announced numbers. We look at what actually changed and what cost structure the change creates in a running system. Apply the same lens to the destruction demand and its cost structure comes into view. Removing weights wholesale voids the compute, time, and data that training consumed. Removing only the influence of specific works starts with the hard task of measuring that influence. Neither path is light. The weight of the demand sits exactly there. A light demand would not lead with destruction.
So how should we actually receive this. I read destruction as a reference point for negotiation and judgment, rather than a final procedure. Leading with the strongest ask is common practice in litigation. Destruction up front does not mean whole models will vanish. What remedies a court can order and what is technically feasible belong to a separate stage. The reported sentence is the starting point of that stage, not its conclusion. What can be read now is that the starting point sits at the strong end, and that its placement defines the dispute as structural. The rest belongs to records and procedure.

Where a regional daily and a metro tabloid meet
Looking a little longer at the two names, the Seattle Times and Newsday, the terrain of news comes into view. One is a regional daily anchored in a Pacific coast metro, the other a daily dug deep into the New York area. Recalling their pages, the common thread I found was that neither lives as a mere relay of national news. Both are reporting organizations with feet on their own ground. They chase city hall, the statehouse, the courts, and the schools every day. Those records pile into archives, and archives become the memory of their places. Wire copy does not replace that work.
What does that common thread mean from the training data side. Local news is a scarce record. National topics get covered in overlapping layers by many outlets, while steady coverage of one area's meetings and rulings often exists only in that area's paper. From the data perspective, such records are hard to substitute. Common sentences can be learned from text available anywhere, yet answers about specific local facts call for local records. Reading this passage, I felt the plaintiffs' claim could be heard as a dispute over records reaching past disputes over sentence craft. Who wrote down the life of a place, and how that writing seeps into model answers.
Some caution belongs here. That local records carry weight does not prove any particular output copied them. Value sits on a different layer from infringement. Weight can explain motive for dispute, while the establishment of each use must be weighed separately. I want to keep that separation clean. The meeting point of the two papers sits at the level of shared concern, not shared evidence. That two organizations with the same concern stand in one sentence only previews the direction of the coming evidentiary fight. Measuring the distance between preview and proof remains the work ahead.
Placing this beside earlier publisher suits
Readers meeting this story fresh will want context. Disputes between publishers and platforms, and between publishers and model builders, did not start here. Reading this report, I recalled the flow of recent years. First a suit by a national paper against a model builder, then a run of similar claims from publishers and authors. Defendants, plaintiffs, and the texture of claims differ case by case. The question running through them stays put. Whether gathering published text at massive scale for training is permitted without asking, and if permitted, where the boundary lies.
Within that flow, two features mark this story's place. First, the widening of plaintiffs, from national flagships toward regional dailies. Second, the pitch of the remedy, with destruction placed on the front line. I read the two together. The parties widen while the demanded form hardens. Cases differ in court, jurisdiction, and issues, so stitching them into one story calls for care. Even so, the grain of the outside trend is legible. Disputes over training data are not fading. They continue while carrying more concrete demands.
From a systems perspective, the next question is obvious. How does parallel litigation change data management inside model organizations. I will not fix an answer, but I can mark the spots that must change. Recording collection sources, documenting filter criteria, and preparing procedures to retrain with a source excluded. Before each suit resolves, organizations face pressure to build pipelines that withstand dispute. Whether pressure turns into good design is a separate matter. If keeping records looks like creating ammunition, teams may drift toward keeping fewer records, and governance thins instead. I want to read the trend without believing it bends only toward the good.

What deleting a model asks in technical terms
Translated into technical language, destruction releases a flood of questions. Writing this section, I sketched the flow in my head. Removing a source from the dataset, lifting its influence out of trained weights, and halting a serving model to swap in a replacement each surfaced in turn. Each step carries its own difficulty. The first is fairly direct. Drop the source from the collection list and delete stored copies. The second is hard. Training blends the influence of countless sentences into numbers, so no clean method exists to pick out one bundle's contribution. The third is operational. What fills the gap at the moment of shutdown.
Why the second step resists solution deserves a slower pass. Weights are not a cabinet with one drawer per source. Traces of every sentence blend into every number a little. Wanting one paper's stories removed does not open one drawer to empty. Retraining from scratch without the source is the surest route, and its price is steep. Approximate techniques for dampening specific influence get studied, yet guaranteeing how much was removed stays difficult. At this point I see the languages of law and technology diverge. Destruction in law is binary. Gone or remaining. Deletion in technology is a spectrum. How much was lifted, and what measures it, stays open.
So I read the destruction demand on two layers at once. One is the layer of declaration. Treating the whole artifact as the problem. The other is the layer of implementation. Carrying out the declaration requires deciding what gets measured and what stays. The reported sentence rests on the declarative layer. The implementation layer stands empty. Emptiness here does not signal failure. A short report does its job by delivering the declaration. Filling the blank belongs to courts and engineering. In this post I only want to sketch the blank. That it runs large, and that filling it costs more than money.

Conditions I check first in governance documents
Shifting the angle, imagine reading a model organization's internal documents. When I read governance material, I look for conditions rather than declarations. How far source records for training data reach, what process handles an exclusion request, whether a model version without a given source can actually be built. Anyone can write declarations. Only teams that ran the procedures can write conditions. In disputes where destruction enters the picture, the presence of these conditions becomes readiness itself.
The first check is source records. Where material came from, when it arrived, which filters applied. Records narrow disputes. With them, a team can answer whether a given source in a given window entered the mix. Without them, every account turns into reconstruction. The second check is exclusion and retraining procedure. Past removing a source from the dataset, the question is how model versions get swapped when an exclusion request arrives. The third check is output side control. Whether machinery exists to block long verbatim passages of specific works, and how that machinery gets evaluated. I think dispute costs diverge widely between organizations that carry these three and those that do not.
My team applies the same yardstick to architectures. Past whether a structure looks tidy, we ask whether a way back exists when trouble hits. A pipeline without a return path is not fast. It is hazardous. Rolling back to a known version, or removing a data source, has to be possible for governance to stand. This story reads as a question about those return paths too. The plaintiffs challenge the road their works took in, and with the word destruction they demand a road out. When both roads stay recorded and bound into procedure, disputes turn from wars of attrition into solvable problems.

What I want to see next is records and options
Finally, let me gather what remains past this report. The facts it carries run short. Two papers sued two companies, alleging use of works and demanding destruction of models. On that brief skeleton I want to leave coming attractions in question form. Which layers of the claim a court will probe, what records the defendants will produce, how both sides will draw the bounds of destruction. I will not write answers in advance. Still, the yardstick for judging whatever answers arrive can be set now.
What I would watch next is twofold. First, records. Collection, filtering, and version records needed to settle what became material. With records, disputes narrow. Without them, disputes widen. Second, options. What middle paths can be drawn between the poles of destruction and retention. Excluding specific sources, swapping specific versions, tightening output side controls: whether such midpoints exist in practice, and what would verify them, draws my curiosity. I think governance levels show less in declarations than in the design of these midpoints.
This result pleases me, yet the data stays too thin to generalize from. One report cannot settle the whole training data dispute. Even so, one shift stands clear. The language of dispute moves from money toward structure. Faced with suits that ask what should be removed, building models becomes work that must stay explainable end to end. Where material came from, how it mixed, how to reverse course when trouble hits. When the next story lands, I plan to pull out this post's questions again. Whether records appeared, whether options widened, whether demands aimed at structure met answers built from structure.