- Community aggregator
- Country: United States
Small LLM Judges Approved 11% and 41% of Wrong Answers. Then I Fixed My Own Pairwise Test.
Grading small open judges against deterministic oracles, plus the harness that makes most judge calls unnecessary. Scope note, first paragraph as promised: every judge here is ≤3B parameters running locally. Nothing below carries over to frontier judges. Synthetic corruptions and natural errors are…