this post was submitted on 01 Aug 2026
577 points (99.1% liked)

Programmer Humor

32589 readers
518 users here now

Welcome to Programmer Humor!

This is a place where you can post jokes, memes, humor, etc. related to programming!

For sharing awful code theres also Programming Horror.

Rules

founded 3 years ago
MODERATORS
top 50 comments
sorted by: hot top controversial new old
[–] SirDimples@programming.dev 2 points 11 hours ago

Wow, got about 9 repos of mine there and a few from my startup's, only public ones are scraped so I say fair enough. Happy to have switched to running my own git infra a year ago

[–] Bieren@lemmy.today 8 points 1 day ago

If someone wants to scrap the shitty ass code I have on GitHub, have at it. Talk about poisoning AI

[–] rounding_error@lemmy.today 3 points 1 day ago* (last edited 1 day ago)

Out of my 33 public repositories, they only scraped the 7 oldest, pointless ones. Weird.

[–] onlinepersona@programming.dev 9 points 1 day ago (2 children)

Get off of Github if you think this is a problem 🤷 There are alternatives like Forgejo (Codeberg), Gitlab, and Radicle (decentralised).

[–] cultist@feddit.dk 2 points 15 hours ago

I am personally on SourceHut, but do have a codeberg account.

[–] OsrsNeedsF2P@lemmy.ml 6 points 1 day ago (1 children)

Those are obviously scraped as well?

[–] onlinepersona@programming.dev 1 points 15 hours ago

Sure, but they can't discover them all. Especially radicle is not easy to crawl, due to it's decentralised nature. It can even be hosted on TOR and I2P

[–] mindbleach@sh.itjust.works 24 points 1 day ago (1 children)

Genuinely surprised it's even opt-out.

These companies train on Disney DVDs. Permission is not a factor. Training is transformative use, as much for counting letter frequency as for building a chatbot that can sort of code.

[–] WhyJiffie@sh.itjust.works 24 points 1 day ago (1 children)

no, officer, you misunderstand! I'm not pirating this movie, I'm just training my intelligence on it! it is transformative use, see, I can now write you this summary!

[–] mindbleach@sh.itjust.works 2 points 1 day ago (1 children)

Quoting one sentence from a book is fair use even if you shoplifted that book.

[–] Wiz@midwest.social 1 points 10 hours ago (1 children)

I would say this is a gray area of law. It hasn't been tested yet. There are a few factors in determining "fair use". One of those factors is commercialization, which could nullify fair use. Another is the amount you're using.

[–] mindbleach@sh.itjust.works 1 points 7 hours ago* (last edited 7 hours ago)

Commercial parody is commonplace fair use. If it surpasses the original... oh well.

The amount being used becomes negligible as the corpus grows. Filtering a billion webpages into one gigabyte means their average contribution is one byte. Play with the numbers all you like, it's gonna come out to less than a paragraph.

If we went full copyright maximalist, and criminally banned using anything but public-domain / BSD / CC0 works, the commercial impact would not be much different. The technology itself shifts the economics of text, images, video, and code.

[–] the16bitgamer@piefed.ca 177 points 2 days ago (13 children)

Looks for my username. Sees my 8 open repositories in there. Sees my poorly coded Uni projects are also in there.

Oh lord my code is actively helping making AI worse.

[–] umbraroze@slrpnk.net 2 points 1 day ago

Oh lord my code is actively helping making AI worse.

I checked out, it has some of my repositories. They crawled this stuff in 2025 and they probably won't update it. Whoever uses this dataset will have to deal with some super garbage, I tell ya.

[–] iammike@programming.dev 68 points 2 days ago (2 children)

Mission failed successfully!

Glad to be part of the crew with shitty code in Github to taint them plagiarism machines!

load more comments (2 replies)
[–] Rentlar@lemmy.ca 22 points 2 days ago (1 children)

Woo! My crappy code and group projects are in there too!

May all AI generated code be in one giant main loop thanks to my influence 😈

load more comments (1 replies)
load more comments (10 replies)
[–] brucethemoose@lemmy.world 63 points 2 days ago (19 children)

Also, all this reminds me of drama in the Skyrim and Minecraft modding scenes, when devs publish stuff under Apache or MIT or whatever.

Then they find out they don’t like what others are doing with their code. Drama ensues.


…That’s kinda the deal with permissive licenses.Or posting publicly, like here on Lemmy. People will do things you don’t like with your code or content.

[–] kibiz0r@midwest.social 59 points 2 days ago (6 children)

Eh, I don’t think it’s hypocritical to contribute to a commons and then get mad when someone comes along and tries to use the commons to undermine the commons.

Like yes, the commons is there to be used… but not to kill the commons.

https://www.citationneeded.news/free-and-open-access-in-the-age-of-generative-ai/

[–] lemmyman@lemmy.world 15 points 2 days ago (1 children)

It's kind of a tragedy if you think about it

load more comments (1 replies)
load more comments (5 replies)
[–] Toga77@lemmy.world 29 points 2 days ago

Nah if you're a massive AI company and you scrape without contribution, you're a huge piece of shit.

It's just stealing plain and simple like anything else.

They're not a small user getting open source software, they're scraping what is already done to try and make you obsolete.

load more comments (17 replies)
[–] charonn0@startrek.website 13 points 1 day ago (5 children)

A number of my repos are listed.

But the weird part is that it also lists a repo I don't recognize. The repo does actually exist on my github account, but it's marked as "ignored", and the description says it was automatically exported from Google Code. The code seems to be a MacOS shareware file encryption tool called "BitClamp", published circa 2008.

No idea how it got on my account.

load more comments (5 replies)
[–] ProbablyUnwise@anarchist.nexus 18 points 2 days ago* (last edited 2 days ago)

meanwhile I'm just here scraping GitHub repos for API keys and credentials 🤷‍♂️

for legal reasons I must insist this is a joke, and in Minecraft.

[–] dextro@feddit.org 22 points 2 days ago* (last edited 2 days ago) (2 children)

I don’t see a problem as long as they stick to AGPL when building a product out of it

Edit: Oh they also scraped my proprietary code 🧐

[–] Axolotl_cpp@feddit.it 1 points 1 day ago

They also scrapped my GPL code, apparently

load more comments (1 replies)
[–] heliotrope@retrofed.com 32 points 2 days ago (2 children)

Bad News: My old GitHub repos are there.

Good News: I wrote that shit when I was 12. The code runs, but it's not good and not inventive.

load more comments (2 replies)
[–] TeaWithDani@lemmy.world 8 points 1 day ago (3 children)

In theory, it's all supposed to be permissively licenced code and the opt out is more than other models give. I saw Starcoder as one of the more ethical models. I thought the underlying principals to be fair at least.

I'm interested in this gut hostility to it regardless. Kind of shows how you can't present LLMs in a positive angle no matter what.

[–] Viking_Hippie@lemmy.dbzer0.com 18 points 1 day ago (18 children)

Kind of shows how you can't present ~~LLMs~~ harvesting peoples data without consent or even warning and making it difficult to impossible for people to avoid it in a positive angle no matter what.

Fixed it for you.

An LLM built from only consensually provided data would be perfectly fine, as long as it works without environmentally ruinous data centers.

In fact, that was how EVERY LLM was to begin with, until regulatory capture and corporate impunity reached the current crescendo.

load more comments (18 replies)
[–] JackbyDev@programming.dev 7 points 1 day ago (4 children)

You can host copyleft as well as all rights reserved code on GitHub. It's not like Codeberg.

load more comments (4 replies)
load more comments (1 replies)
[–] carotte@lemmy.blahaj.zone 14 points 2 days ago (1 children)

…and now, suddenly, github being overrun by vibecoded garbage isn’t so bad

load more comments (1 replies)
[–] brucethemoose@lemmy.world 24 points 2 days ago* (last edited 2 days ago) (6 children)

Well… I’d rather the dataset be public and there, with an ostensible centralized opt-out, instead of every AI startup frantically rescraping the same things their predecessors did.

load more comments (6 replies)
load more comments
view more: next ›