Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Tengyang Xie, Dylan J. Foster, Akshay Krishnamurthy, Corby Rosset, A. Awadallah +1 more
2/12/2026
Semantic Scholar

Code Implementations

No confident code match yet

We couldn't find an author-owned or strongly-evidenced community implementation for this paper. Any repos shown below are weak matches — verify before relying on them.

No code implementations found yet.

Know of an implementation? Let us know in the comments below!

Cite this paper

@article{xie2026exploratory,
  title  = {Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF},
  author = {Tengyang Xie and Dylan J. Foster and Akshay Krishnamurthy and Corby Rosset and A. Awadallah and Alexander Rakhlin},
  year   = {2026},
  url    = {https://api.semanticscholar.org/CorpusID:ef6a1aed21f3193f6b27440ea714a9d6870a9b5c},
  journal = {ICLR 2025 2025}
}

Discussion