papersTODAY 04:00 UTC
Semi-Bandit Algorithm Selects k Paths to Cut Worst-Case Transmission Time
A new arXiv paper studies an online learning problem where a system must repeatedly choose k paths through a network to keep the slowest path's transmission time as low as possible. The authors formalize this as a stochastic semi-bandit problem, where feedback is observed only for the paths actually selected. They propose and analyze algorithms for minimizing the longest path length under uncertainty.