so? It was never advertised as intelligent and capable of solving any task other than that one.
Meanwhile slop generators are capable of doing a lot of things and reasoning.
One claims to be good at chess. The other claims to be good at everything.
opposite or not, they are both tasks that the fixed-matrix-multiplications can utterly fail at. It's not a regulation thing. It's a math thing: this cannot possibly work.
If you could get the checker to be correct all of the time, then you could just do that on the model it's "checking" because it is literally the same thing, with the same failure modes, and the same lack of any real authority in anything it spits