Plain-text files are at risk
Lobsters
Fulvid
Product Hunt
A standalone desktop editor for Markdown and MDX
FeedsBar
Product Hunt
A quiet, continuous news/media ticker for your Mac desktop
PewCB
Product Hunt
Vibe-routing is coming. Desktop PCB Factory is here.
2BA.AI
Product Hunt
Stop waiting for tokens, and start shipping
Googlebook
Product Hunt
The laptop your Android phone has been waiting for
The Softness of Metal
Hacker NewsComments
MUNI Heritage Weekend in San Francisco
Hacker NewsComments
Tight Sample Complexity Bounds for Entropic Best Policy Identification
arXiv cs.LGarXiv:2605.13717v2 Announce Type: replace
Abstract: We study best-policy identification for finite-horizon risk-sensitive reinforcement learning under the entropic risk measure. Recent work established a constant gap in the exponential horizon dependence between lower and upper bounds on the number of samples required to identify an approximately optimal policy. Precisely, known lower bounds scale in $\Omega(e^{|\beta| H})$ where $H$ is the horizon of the MDP, while the state-of-the-art upper bound achieves at best $O(e^{2|\beta| H})$ (arXiv:2506.00286v2) using a generative model. We show that this extra exponential factor can be traced to overly loose concentration control for exponential utilities. To close this open gap, we revisit the analysis of this problem through a forward-model based algorithm building on KL-based exploration bonuses that we adapt to the entropic criterion. The improvement we get is due to two main novel technical innovations. We leverage the smoothness properties of the exponential utility to derive sharper concentration bounds, and we propose a new stopping rule that exploits further this tightness to obtain a sample complexity that matches the lower bound.
继续滚动加载更多…