Multi-turn Reinforcement Learning with Preference Human Feedback

Lior Shani, Aviv Rosenberg, Asaf B. Cassel, Oran Lang, Daniele Calandriello +8 more
2/3/2026
Semantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. 5 weaker matches are hidden by default — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{shani2026multiturn,
  title  = {Multi-turn Reinforcement Learning with Preference Human Feedback},
  author = {Lior Shani and Aviv Rosenberg and Asaf B. Cassel and Oran Lang and Daniele Calandriello and Avital Zipori and Hila Noga and Orgad Keller and Bilal Piot and Idan Szpektor and Avinatan Hassidim and Yossi Matias and Rémi Munos},
  year   = {2026},
  url    = {https://www.semanticscholar.org/paper/d623c09b3cf03a0feeecd3ba97894c2ecd16a15a},
  journal = {NEURIPS 2024 2024}
}

Discussion