genrm
an archive of posts with this tag
| Aug 11, 2026 | One Token to Fool: GenRM도 결국 뚫린다 |
|---|---|
| Aug 11, 2026 | DeepSeek-GRM: reward model이 평가 기준을 스스로 만든다 |
| Aug 11, 2026 | Generative Reward Models: GenRM과 선호 학습을 잇다 |
| Aug 11, 2026 | Generative Verifiers: reward를 분류가 아니라 생성으로 풀다 |