The opponent problem

A reinforcement-learning agent is only ever as good as whatever it was measured against. These are 76 training runs in Gin Rummy, each scored three ways at once — against a random player, against the previous champion, and against a fixed gold-standard expert. The three numbers disagree violently, and which one you report decides what your paper claims.

Every value is read from this repository's own sweep results. Accompanies the AIIDE 2026 paper.

Gin Rummy · masked PPO/TRPO · PettingZoo / RLCard · 412 GPU-hours across the sweep Project report · Repository · arXiv