this post was submitted on 30 Jul 2026
1529 points (98.4% liked)

Fuck AI

7790 readers
1832 users here now

"We did it, Patrick! We made a technological breakthrough!"

A place for all those who loathe AI to discuss things, post articles, and ridicule the AI hype. Proud supporter of working people. And proud booer of SXSW 2024.

AI, in this case, refers to LLMs, GPT technology, and anything listed as "AI" meant to increase market valuations.

founded 2 years ago
MODERATORS
 
you are viewing a single comment's thread
view the rest of the comments
[–] ricecake@sh.itjust.works 40 points 1 day ago (5 children)

I agree with the conclusion, but that rationale is wrong. First, you can digitize a DVD. Second, it's not a double standard. You can grab copies of random stuff and jam it in an AI model too.
Our laws are written such that it's making a copy outside of reasonable use that's illegal, and AI training only makes a copy incidentally to what they're doing and then it's deleted. It's the same standard that makes viewing a photo on an artists website legal.

It's not bullshit because they're breaking the law, but because we need to refine the law to make it clear training an AI model isn't a reasonable usage anymore than a public broadcast of a DVD is a reasonable use.
Trying to shoehorn it into the existing laws will just create a nightmare of loopholes and complications.

[–] bss03@infosec.pub 12 points 1 day ago (3 children)

AI training only makes a copy incidentally to what they’re doing and then it’s deleted. It’s the same standard that makes viewing a photo on an artists website legal.

That's not what the U.S. Copyright office says about training. They hold that it does implicate the copyright of reproduction. Meaning: If you train on a protected work without a license you are violating copyright, and if that's not a fair use then you are breaking the law.

Training ~ viewing might be an analogy used by "AI" brands, but it is not legal reality.

[–] ricecake@sh.itjust.works 2 points 2 hours ago (3 children)

I'm not sure that's been extensively tested in courts. The document you referenced below appears to be as-yet not officially published, so I don't believe it actually qualifies as an official position yet, but the bigger issue is that it's untested in court.

This thread is a response to an AI court case where the ruling was that training on copy written works is fair use.

https://www.reuters.com/sustainability/boards-policy-regulation/us-judge-approves-15-billion-anthropic-copyright-settlement-with-authors-2025-09-25/

Alsup ruled in June that Anthropic made fair use of the authors' work to train Claude, but found that the company violated their rights by saving more than 7 million pirated books to a "central library" that would not necessarily be used for that purpose

Regardless, you do make good points and I think we agree that the end state is "they shouldn't be able to do that". I have concerns that using existing standards that take copying too literally results in some unintended ambiguity, and situations where AI training is incidentally blocked, but so is stuff like "opening a news article on a computer", which does the same things the copyright office report highlights as infringement.
I think we'd be in a much more agreeable place if we just legally state that a commercial AI tools training isn't fair use. That lets you have nuance like "search engine? It's a statistical model, but not generative: allowed. AI agent? Statistical model that's generating content as opposed to classification or ranking: not allowed".

[–] bss03@infosec.pub 1 points 1 hour ago

I think we’d be in a much more agreeable place if we just legally state that a commercial AI tools training isn’t fair use.

Good luck getting any federal law changes through before 2028, at best. So, for at least a couple of years, we get to use the existing laws to address generative "AI"'s use of works still under copyright protection.

Any change in status before then we be policy changes by the U.S. Copyright Office, but they have already come down against training (mostly; the publications are long because there's a lot of nuance).

[–] bss03@infosec.pub 1 points 1 hour ago

Alsup ruled in June that Anthropic made fair use of the authors’ work to train

But, Kadrey v. Meta Platforms, Inc. (Judge Chhabria, June 25, 2025) states that “in most cases,” training LLMs on copyrighted works without permission is likely infringing and not fair use and "this ruling does not stand for the proposition that [...] use of copyrighted materials to train its language models is lawful."

The courts are divided, but the copyright office is not.

[–] bss03@infosec.pub 1 points 1 hour ago* (last edited 1 hour ago)

The document you referenced below appears to be as-yet not officially published

Parts 1 and 2 are officially published.

On May 9, 2025, the Office released a pre-publication version of Part 3 [...] A final version of Part 3 will be published in the future, without any substantive changes expected in the analysis or conclusions.

[–] cavitationfetishist01@quokk.au 3 points 21 hours ago (1 children)

Its a slopper who wants to project these spreadsheets as 'conscious' when what's happening is they're essentially being transcoded into statistical models.

[–] bss03@infosec.pub 4 points 21 hours ago (1 children)

Yup. Give a Markov chain multi-billion parameters and you can get some surprisingly cogent results.

I will freely admit that current LLM architectures include several innovations that make them not actually Markov chains, but it's still statistics and linear algebra. I don't know what thought is, but I'm quite unconvinced that LLMs (or any current generative AI architecture) is doing it.

[–] RogueJello@lemmy.world 2 points 20 hours ago (2 children)

I don’t know what thought is, but I’m quite unconvinced that LLMs (or any current generative AI architecture) is doing it.

Okay, why not? I also don't know what thought is, so I don't think it's possible to say if an LLM is or is not doing it. And giving wrong or incoherent answers doesn't invalidate it as thought or your local stoner buddy would be considered brain dead.

[–] bss03@infosec.pub 5 points 18 hours ago* (last edited 17 hours ago)

I've not seen evidence of it in any of my interactions with generative AI, which have pretty universally been bad. I feel like it has something to do with autonomous spontaneity. I recognize it in animals I can't communicate well with, but I found it lacking in the LLM that I tried to play a TTRPG with. It would be easier for me to be convinced, if I really had a better understanding of what thought is. It's hard for me to be convinced because while I understand LLMs and diffusion networks better than most people*, I don't think I understand thought so I recognize the gap.

Also I'm not sure I agree with your final assertion, the stoner buddy is plenty wrong, but there is a coherency there. When coherency disappears entirely from human thought that's usually a seizure or stroke. Even as confusing as they are dreams and acid trips often have a coherency while you are in them, if not one that's easily described when recalling the experience.

*: My formal AI training ended before big data met ML, so it's woefully out of date. I am quite the computer geek tho, it's just my passion tends toward languages, type systems, and proof assistants. So, better than most, but not an expert by any means.

[–] cavitationfetishist01@quokk.au -1 points 19 hours ago* (last edited 19 hours ago)

Nah. Shut the fuck up slopper. If you're going to insult us and then ask chatgpt to win the argument, which you always do and it always misses the point, I'm not going to answer your question.

The fact is these systems can do what a lot of humans do. That's not because the matrix multiplication is identical to thinking, but because most of these humans have never thought in their lives, do not have interiority, and are not people in any way that matters. Prove you're conscious if you want me to address you as such, fucking slopper.

[–] JackbyDev@programming.dev 4 points 23 hours ago (1 children)

I could've sworn a court case decided otherwise. Literally EVERY AI model in existence right now is commiting copyright theft on a massive scale if that's the interpretation the courts took. Which is why I have a hard time buying it. I fear it's reached the idea of normalcy in people's minds and we'll never see it illegal.

[–] bss03@infosec.pub 5 points 22 hours ago* (last edited 22 hours ago)

There's been a couple court cases (that I know of / at least), and one judge was accepting on the argument that model training was a "fair use" while the other was not. I think both of those rulings came down prior to the publication of the U.S. Copyright Office guidelines.

Also, I'm not 100% sure that the U.S. Copyright Office is an authority here. The DOJ and/or Federal judiciary would have the authority to interpret the copyright laws: The DOJ to decide to prosecute, and the judiciary to make binding rulings and/or advise juries. I'm sure both the DOJ and the judiciary will give a lot of weight to the guidelines, but the guidelines aren't actually the law.

In any case, you can read the guidelines and make your own decisions: https://www.copyright.gov/ai/ Part 3 is about training, and I think the damning bits are III, B and D. Part 2 is about outputs, and I think the damning bits are II, B and D.2. (My summaries: 1. Training infringes 2. Outputs that are substantially similar infringe 3. models get no copyright 4. prompts are NOT 'human creative effort' and thus are insufficient to establish copyright 5. human creative effort still gets copyright protections, even when generative AI is used as a tool in the creative process.)

It is likely that commercial generative AI is in violation of a lot of copyrights, yes. Research projects are fair use, but only as long as they stay research projects.

[–] cmhe@lemmy.world 2 points 17 hours ago* (last edited 17 hours ago) (2 children)

First, you can digitize a DVD.

Nitpick: Why would anyone do that, DVDs are already digital mediums with a filesystem. So they first have to convert to analog media first... And that introduces losses...

You just put a DVD in your drive and now you can copy files from it to you HD... If copying files is now called 'digitizing' we live in a strange world...

[–] ricecake@sh.itjust.works 1 points 4 hours ago

It definitely does have the meaning of making an analog signal discrete, and DVDs are digital in that sense.
There's also the sense of serializing or enumerating something into a format more flexible or amenable to computer manipulation. More briefly: to put it on a computer.

For some things it's super obvious: a vinyl record has a groove you can see, and you digitize it by sampling the output of the wiggle: analog.
The contents of an SD card are invisible and the only way to observe them is via a computer, where they're already presented in a flexible format: not analog.

Then you have fuzzy things: a printed picture that was taken with a digital camera. Grab a magnifying glass and you can see that it's got blocks of quantized color, and isn't continuous. You digitize the photo by shining light on it and capturing the bounce.
Optical media is a disc with a groove in it that you read by shining a light on it to capture the bounce. Instead of a horizontal wiggle though it's pits creating an interference pattern with a laser.
It's digital because it's quantized information that it stores. It needs to be digitized to work with easily because the format is often read-only or write once. It's analog because the pits themselves are continuous and the system spends surprising amount of nuance correcting errors from things like fingerprints and dust.

Books are plainly analog, even though they're composed almost entirely of discrete, quantized units of structured information. Sometimes even with a lookup table and index for faster search!

Magnetic tape stores analog audio by varying the strength of a magnet with the audio signal. It stores digital information by only setting the magnet to specific strengths. Same for VHS or cable Internet.

To make a long ramble short: digitize also means to stick it on a computer, owing to efforts to digitize things often being focused on items that aren't continuous, like text.

[–] kamen@lemmy.world 4 points 17 hours ago (2 children)

If you have a large number of DVDs, digitising them makes them easier to browse and decouples you from the physical medium (so you don't have to bring them with you everywhere in order to watch them).

As for the conversion, maybe you're confusing this with ripping vinyl or tape. Ripping a CD or a DVD is easy and reproducible and gives you a 1:1 digital copy of it.

[–] Kazumara@discuss.tchncs.de 5 points 16 hours ago (1 children)

You're missing his point, he's unhappy with the word choice. Ripping a DVD is not digitizing it. Digitizing means specifically turning an analog signal digital.

[–] kamen@lemmy.world 2 points 16 hours ago

Ah, my bad - now that I read it again, you're right. In the same line of thought though "digitising" is sometimes wrongly used in the place of ripping.

[–] cmhe@lemmy.world 2 points 14 hours ago

Maybe because I'm not a native english speaker, but to me 'digitising' means converting an analog medium to a digital one. Like with vinyl, VHS, etc. And 'ripping' is about an extraction process. Like overcoming a protection or more difficult to access mediums like Audio CDs, which don't have a real filesystem, and into a easily accessible single file on a harddrive.

DVDs are already digital, and if they don't have a copy protection, which you have to rip through, you can just copy the files to a hardrive and then, if you want reencode them into one more portable file...

[–] testaccount789@sh.itjust.works 6 points 1 day ago* (last edited 1 day ago) (1 children)

There may be a limitation if the DVD is copy-protected, as is usually the case. There's too much to read for a single comment in DMCA: https://www.congress.gov/105/plaws/publ304/PLAW-105publ304.pdf

But it does fall under this definition (§1201):

‘‘(a)(3) As used in this subsection—
‘‘(A) to ‘circumvent a technological measure’ means to descramble a scrambled work, to decrypt an encrypted work, or otherwise to avoid, bypass, remove, deactivate, or impair a technological measure, without the authority of the copyright owner;

and

‘‘(a) VIOLATIONS REGARDING CIRCUMVENTION OF TECHNOLOGICAL MEASURES.—(1)(A) No person shall circumvent a technological measure that effectively controls access to a work protected under this title.

[–] skisnow@lemmy.ca 2 points 1 day ago (1 children)

They hyphenated techno-logical both times in the second paragraph? Was it a line break both times, or was the moron who wrote it really that out of touch with the topic he was writing laws for?

Line break, I'll fix it.

[–] T156@lemmy.world 4 points 1 day ago (1 children)

The larger part of the infringement is probably its use commercially. I doubt that there would have been such a fuss if it was a fully-open, low-profit operation.

But as-is, the commercial products are being used to make money for the AI company in an unauthorised way.

Similar to how it's generally frowned upon for fan media to make money, because it starts being infringement. You can have a "support the fan media maker" button, but you generally can't do things like put your fan media behind a pay wall. The IP owners will come down hard on you for that.

[–] bss03@infosec.pub 5 points 22 hours ago

Commercializing a work virtually guarantees it's creation isn't "fair use".

But also, "fair use" is actually quite a bit more narrow than just non-commercial.

[–] Hawke@lemmy.world 3 points 1 day ago (2 children)

How do you digitize a DVD when it’s already digital?

[–] ricecake@sh.itjust.works 2 points 18 hours ago

Har har.

Different senses of the word digital. The dvd is digital as in "made discrete and not analog".

I meant in the sense of "to move off of fixed use physical media and translate to a format more agnostic to storage medium or conducive to transfer and immediate processing".

More succinctly: to copy something to a storage medium that's harder to loose under the couch.