Over time I have discovered that for me, the scoring of The Storygraph (which also tends to match Goodreads) aligns reasonably well with my own views.
That being said, there are almost literally no books with 5* and the universally acclaimed masterpieces sit at around 4.3-4.4.
So I avoid anything below 3, and only read below 3.5 if I'm really in the mood for that. That way I'm filtering pretty bad stuff, but not letting my filter stand in the way of many (hopefully any) books I would have enjoyed.
And yeah, there's cases where you don't agree with the rating but that's why there are other mechanisms (like reviews) to supplement it.
Unfortunately AI doesn't work like that. Any way to explain it would be an oversimplification but I can try.
The training data (songs) are used to create the weights. This is a bunch of numbers that are on their own meaningless - they don't map to specific songs, but to attributes such as "tone" "rhythm"... And like that but many (millions of) abstract attributes that don't make sense as people, but make sense to computers.
So if the thing makes a song that is very "rhythmic" but also "tonal", there's no specific training song that contributed to that - all did, and it's a mess to decypher how much each contributed to each attribute. Except the resulting song doesn't use two parameters, uses many millions, so it's essentially impossible to know.