Sakana AI’s Multi-Layered Review catches 73% of core-claim errors in a new 1,164-error benchmark for LLM-assisted peer review ...