this post was submitted on 01 Aug 2026
17 points (100.0% liked)
LemmyToday
318 readers
28 users here now
If you experience issues or problems with this instance (lemmy.today), this is the place to discuss them. Or if you just want to ask questions about how something works. Anything related to the instance or lemmy itself.
founded 2 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
Yeah, I don't want the whole world of social media to wind up in a "get an account everywhere to view anything" situation.
I think that it's a reasonable stopgap fix from an admin standpoint, because there aren't a whole lot of great levers for admins to pull to try to deal with bots beating the hell out of a server. And one can't just have one's server become unusable.
But it should really be something that is addressed on the dev side. Like, we need a real, long-term solution for the Web.
Also, while improving server performance and failure modes under load would help, the "badly-written, aggressive scraper bots that ignore robots.txt and are given loads of network resources clobbering servers" is something that affects many, many different Web servers out there. This isn't a Lemmy problem, nor even just social media problem. It's a Web problem.
I think what really needs to happen is some form of mechanism that makes it hard for purely-automated systems to hammer a server, something that segregates "bad bots" and humans. Maybe a CAPTCHA or other proof-of-resource thing or something like that.
It's going to also create server admin problems. Think of things like Web search engines indexing sites
server admins normally want their servers to be indexed by "good" bots
so it's going to create additional headache in that even after successfully creating a system that is able to segregate bots and humans, admins have to be able to identify "good" bots
maybe "good" IP ranges or have "good" bots use public keys or something
and then whitelist those. Maybe have some mechanism to distribute "good bot" whitelists and an Apache mod that downloads them occasionally, lets an admin subscribe to a whitelist or something.
I also kind of like the ability to occasionally
wget -rsections of sites, and a solution of the above sort will break that without logging in in a browser and handing off browser cookies towget, which is not ideal.And there's some resource cost to humans
just lower
for proof-of-resource solutions, and human time cost for CAPTCHAs, all of which are undesirable.
considers
One unorthodox solution might be making it easier to snapshot a site. Like, okay. One of the things that irritates the hell out of me is that the bots are doing this to a number of websites that go out of their way to make it really easy to distribute the data that they're after.
Like, GitHub
which has lots of open-source code, which is useful to train coding AI models
ran into load problems with bots scraping it. That was induced by the same data being massively downloaded over via inefficient Web views. Bots would follow every single link, many of which displayed the same data formatted slightly differently. But...you can already get all of the code far more efficiently, which would be better for GitHub and better for bot operators, by just using
git cloneto pull the code (and all of its history!). Bot operators could just say "oh, this is GitHub" and justgit cloneeverything and the load would be completely ignorable. I'm sure that people have mass-git-cloned GitHub lots of times before. The reason bot operators presumably aren't doing it is because it takes some amount of time and dev effort on their side compared to just "build generic Web spider and clobber everything".Same thing for the Threadiverse. The Threadiverse will let you set up a Threadiverse instance and subscribe to everything, efficiently feed all the posts and comments you want to your instance, the moment they come in. In nice, machine-readable form, rather than in something intended for humans that you have to scrape and post-process. But...it takes more dev effort to set up something specific to the Threadiverse than to just treat it like another website.
If there were some sort of widely-adopted API for "request site snapshot" or "request snapshot of changes since time X", widely-enough that it were worth using, maybe bot operators would use that instead. If they don't have to write a "detect GitHub site and
git clone" and "detect Lemmy and subscribe to posts and comments" system, but just have a single universal "dump changes" API, they might use that; less dev effort on their end.considers
I guess the problem is that some website operators might treat the "snapshot API" itself as an opportunity to discriminate between bots and humans, and just return garbage or nothing as a snapshot. Some websites don't want to be scraped at all. If many websites returned incorrect data in response to such an API request, that'd kill the incentive of bot operators to use such an API.
100%. Everyone gets this problem and then they put their site behind cloudflare. But what happens when the entire Internet is behind an american company who can decide which sites it likes and which it doesnt?
Yeah, the web was intended to be this universal format that would always work on any platform and any operating system. But the downside is clear today. Some estimates say that we already have much more bot activity than human activity on the internet. And we have AI creating content for the last few years...