It's probably not directly related to these changes if you're only doing it Saturday and it's only on old.lemmy.com, but if you're looking into load problems, might be indirectly relevant. In the past...I think two days, using the standard Lemmy Web UI, I've seen a high rate of the server returning a page just reading:
Error!
There was an error on the server. Try refreshing your browser. If that doesn't work, come back at a later time. If the problem persists, you can seek help in the Lemmy support community or Lemmy Matrix room.
Sometimes, but not always, the above error text is also followed by:
The server returned this error: Error. This may be useful for admins and developers to diagnose and fix the error
This has not, as far as I can tell, shown up when trying to view lists of posts (like, when I view all subscribed posts) but it does show up when trying to view the page for a post. It appears to affect all posts for me, not just specific ones.
I initially thought that its just that the post had been deleted, but I could view said posts fine on the original, remote instance.
One such example that I just hit, tried to load a number of times unsuccessfully:
https://lemmy.today/post/57605044
Which has, as a remote post:
https://nord.pub/c/games@lemmy.world/p/326761/is-routine-worth-playing
I reloaded five or six times, and eventually it came up locally on lemmy.today.
When it happens, it seems to affect all of the post pages that I try to load.
When I see this, the error page comes up pretty quickly, maybe in a second, so it's not some (long) timeout. Maybe hitting some limit on concurrent database queries or something like that?
tries loading page again
And now that post is back to showing an error. Also, I saw it without doing a cache-invalidating reload on Firefox (shift-reload rather than a normal reload), so whatever the server is trying to generate, and failing, it's not driven by the server trying to respond to something that my browser will have cached.
When the problem is happening, the same problem also shows up when I try to view my user profile page (so doesn't as best I can tell, affect lists of posts, either "all subscribed" or viewing posts in a community, but does affect individual post pages and does affect user profile pages).
Yeah, I don't want the whole world of social media to wind up in a "get an account everywhere to view anything" situation.
I think that it's a reasonable stopgap fix from an admin standpoint, because there aren't a whole lot of great levers for admins to pull to try to deal with bots beating the hell out of a server. And one can't just have one's server become unusable.
But it should really be something that is addressed on the dev side. Like, we need a real, long-term solution for the Web.
Also, while improving server performance and failure modes under load would help, the "badly-written, aggressive scraper bots that ignore robots.txt and are given loads of network resources clobbering servers" is something that affects many, many different Web servers out there. This isn't a Lemmy problem, nor even just social media problem. It's a Web problem.
I think what really needs to happen is some form of mechanism that makes it hard for purely-automated systems to hammer a server, something that segregates "bad bots" and humans. Maybe a CAPTCHA or other proof-of-resource thing or something like that.
It's going to also create server admin problems. Think of things like Web search engines indexing sites
server admins normally want their servers to be indexed by "good" bots
so it's going to create additional headache in that even after successfully creating a system that is able to segregate bots and humans, admins have to be able to identify "good" bots
maybe "good" IP ranges or have "good" bots use public keys or something
and then whitelist those. Maybe have some mechanism to distribute "good bot" whitelists and an Apache mod that downloads them occasionally, lets an admin subscribe to a whitelist or something.
I also kind of like the ability to occasionally
wget -rsections of sites, and a solution of the above sort will break that without logging in in a browser and handing off browser cookies towget, which is not ideal.And there's some resource cost to humans
just lower
for proof-of-resource solutions, and human time cost for CAPTCHAs, all of which are undesirable.
considers
One unorthodox solution might be making it easier to snapshot a site. Like, okay. One of the things that irritates the hell out of me is that the bots are doing this to a number of websites that go out of their way to make it really easy to distribute the data that they're after.
Like, GitHub
which has lots of open-source code, which is useful to train coding AI models
ran into load problems with bots scraping it. That was induced by the same data being massively downloaded over via inefficient Web views. Bots would follow every single link, many of which displayed the same data formatted slightly differently. But...you can already get all of the code far more efficiently, which would be better for GitHub and better for bot operators, by just using
git cloneto pull the code (and all of its history!). Bot operators could just say "oh, this is GitHub" and justgit cloneeverything and the load would be completely ignorable. I'm sure that people have mass-git-cloned GitHub lots of times before. The reason bot operators presumably aren't doing it is because it takes some amount of time and dev effort on their side compared to just "build generic Web spider and clobber everything".Same thing for the Threadiverse. The Threadiverse will let you set up a Threadiverse instance and subscribe to everything, efficiently feed all the posts and comments you want to your instance, the moment they come in. In nice, machine-readable form, rather than in something intended for humans that you have to scrape and post-process. But...it takes more dev effort to set up something specific to the Threadiverse than to just treat it like another website.
If there were some sort of widely-adopted API for "request site snapshot" or "request snapshot of changes since time X", widely-enough that it were worth using, maybe bot operators would use that instead. If they don't have to write a "detect GitHub site and
git clone" and "detect Lemmy and subscribe to posts and comments" system, but just have a single universal "dump changes" API, they might use that; less dev effort on their end.considers
I guess the problem is that some website operators might treat the "snapshot API" itself as an opportunity to discriminate between bots and humans, and just return garbage or nothing as a snapshot. Some websites don't want to be scraped at all. If many websites returned incorrect data in response to such an API request, that'd kill the incentive of bot operators to use such an API.