KL-Regularized Contextual Bandits Achieve Logarithmic Regret via Greedy Sampling
A new arXiv paper analyzes KL-regularized contextual bandits under both reward and preference feedback. The authors show that a greedy sampling approach attains logarithmic regret without an explicit dependence on the eluder dimension. The work covers regret guarantees for the reward-feedback setting and extends the analysis to preference-based feedback.