LIVE PULSE
4.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.2 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.0 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src1.8 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.4 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.1 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.1 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.1 Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election1 src1.1 Study Audits Misalignment in Multi-Modal World Models1 src1.1 Retrieval-Grounded Reasoning Approach Proposed for Universal Multimodal Embeddings1 src4.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.2 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.0 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src1.8 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.4 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.1 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.1 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.1 Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election1 src1.1 Study Audits Misalignment in Multi-Modal World Models1 src1.1 Retrieval-Grounded Reasoning Approach Proposed for Universal Multimodal Embeddings1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

online learning

topic7 events
papersTODAY 04:00 UTC

Explicit Solution Derived for Five-Expert Prediction PDE and COMB Optimality Set

A new preprint presents a closed-form solution to the stationary prediction-with-expert-advice partial differential equation in the case of five experts. The solution is split into three regions, with the first two given by the four-expert result plus a single integral term. The paper also characterizes the exact set on which the COMB aggregation strategy is optimal.

papersTODAY 04:00 UTC

Optimal Switching Regret Bounds for Multi-Armed Bandits Against Oblivious Adversaries

This paper studies adversarial multi-armed bandit problems in which the benchmark arm sequence may change up to S times over the course of play, a setting known as switching regret. It reviews and develops regret guarantees of order the square root of (S+1)KT, which prior work showed is achievable when S is known in advance. The work aims to pin down the optimal achievable rate under an oblivious adversary.

papersTODAY 04:00 UTC

Semi-Bandit Algorithm Selects k Paths to Cut Worst-Case Transmission Time

A new arXiv paper studies an online learning problem where a system must repeatedly choose k paths through a network to keep the slowest path's transmission time as low as possible. The authors formalize this as a stochastic semi-bandit problem, where feedback is observed only for the paths actually selected. They propose and analyze algorithms for minimizing the longest path length under uncertainty.

papersSEP 10 04:00 UTC

Exact-form regret analysis for gradient descent, mirror descent, and follow-the-regularized-leader

A newly posted arXiv paper investigates how online learning methods such as gradient descent, mirror descent, and follow-the-regularized-leader behave when measured against more demanding, action-dependent benchmarks rather than fixed comparison points. Moving past the standard external regret framing, the authors pursue a geometric account of these deviations and derive closed-form expressions for the resulting regret bounds.

papersSEP 10 04:00 UTC

Online Learning of Scale Parameters in Score-Driven Filters

A new research paper addresses how to learn the gain, the scale parameter that multiplies the scaled log-likelihood score in score-driven filters, directly online. Rather than fixing this coefficient beforehand, the method treats each admissible gain as selecting a reachable next state given the current state and the realized scaled score. This allows the filter's update step to adapt during operation.

papersSEP 10 04:00 UTC

Constant-regret algorithm for online inverse integer linear optimization proposed

Researchers address online inverse linear optimization, where a learner predicts weights, observes the agent's optimal action, and updates its estimate each round. Their new small-gradient skipping technique achieves constant regret and only a finite number of mistakes for integer linear optimization problems. This improves on prior bounds that left a logarithmic gap between upper and lower regret limits.

papersSEP 10 04:00 UTC

Researchers study sequence prediction when the oracle can lie

A new machine learning theory paper on arXiv examines how a learner can predict elements of a sequence when the oracle supplying feedback may misreport outcomes. The authors frame the task as a repeated interaction in which the environment picks an outcome from a finite alphabet and the learner must commit to a probability distribution without reliable ground truth. The work analyzes what prediction performance can still be guaranteed in this adversarial setting.