TechTakes

1977 readers

56 users here now

Big brain tech dude got yet another clueless take over at HackerNews etc? Here's the place to vent. Orange site, VC foolishness, all welcome.

This is not debate club. Unless it’s amusing debate.

For actually-good tech, you want our NotAwfulTech community

founded 2 years ago

MODERATORS

dgerard@awful.systems

The AI bill Newsom didn’t veto — AI devs must list models’ training data (pivot-to-ai.com)

submitted 8 months ago by dgerard@awful.systems to c/techtakes@awful.systems

16 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[–] OhNoMoreLemmy@lemmy.ml 20 points 8 months ago (2 children)

The other reason they don't do it is because many models are trained on a large corpus of pirated texts, and documenting this would be a confession.

Not just in an 'I scraped the new york times without permission' kind of way, but in a 'I illegally downloaded a torrent containing bestsellers from the last 30 years' kind of way.

[–] Tar_alcaran@sh.itjust.works 11 points 8 months ago (1 children)

Exactly. It's not that they can't, or that it's too expensive, it's that doing so will reveal their crimes.

[–] imadabouzu@awful.systems 9 points 8 months ago

In a sense, to me, it is the same thing. If your business is built upon repurposing everyone else's inputs indiscriminately to your benefit and their detriment, it is, too expensive, to reveal that simple truth.

[–] Soyweiser@awful.systems 3 points 8 months ago (1 children)

Bestsellers? There used to be torrents of basically all releases. My provider blocks torrent sites and I dont use a vpn so im not sure if people still do this, but downloading basically all books (in english) at once released in a certain period was possible

[–] skillissuer@discuss.tchncs.de 4 points 8 months ago

occasionally i see this for music (weekly new tracks)