Unfortunately, these crawlers instead try to read every single page from Codeberg, no matter if it makes sense. This includes all the different issue filter variants, Git history, as well as the actual files at any point in Git history - even if they are still equal.
I'm surprised they don't lock some things behind account logins. Is it that hard to decide what would be acceptable to no longer serve without an account?
Does the full commit and change history have to be available without an account, under these circumstances? If we think about what would be minimally enough:
- Current branch head tree
- Tag trees (you can link to release source state and [potentially/manually] compare between releases)
Not having a change log history nor code change diff seems like a big loss to me, but you have to draw the line somewhere, and that seems acceptable to me.
To me, Codeberg is in a better position to do so than smaller instances. I already have an account because it has many [relevant/significant] projects, is a home of public good, and is under an appropriate org.