4.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.2 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.0 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src1.8 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.4 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.1 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.1 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.1 Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election — 1 src1.1 Study Audits Misalignment in Multi-Modal World Models — 1 src1.1 Retrieval-Grounded Reasoning Approach Proposed for Universal Multimodal Embeddings — 1 src4.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules — 11 src2.2 Agility Robotics unveils Digit 5 humanoid for warehouses and factories — 2 src2.0 Apple ships rebuilt Siri with Google Gemini, but not in the EU — 2 src1.8 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions — 2 src1.4 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns — 5 src1.1 OpenAI contractors review real ChatGPT conversations to rate responses, report says — 2 src1.1 Anthropic data retention policy prompts firms to limit Claude use for sensitive work — 1 src1.1 Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election — 1 src1.1 Study Audits Misalignment in Multi-Modal World Models — 1 src1.1 Retrieval-Grounded Reasoning Approach Proposed for Universal Multimodal Embeddings — 1 src
OpenAI researchers propose prediction-based curiosity method for RL exploration
Researchers at OpenAI developed Random Network Distillation, a technique that rewards reinforcement learning agents for encountering unfamiliar states as a way to drive exploration. The approach uses prediction errors from a randomly initialized neural network as an intrinsic reward signal. Agents trained with this method surpassed average human scores on Montezuma's Revenge for the first time, a game known for being difficult to explore.
WHY IT MATTERS ↘Sparse-reward exploration has been a core bottleneck keeping RL confined to games and simulations, so a general intrinsic-reward mechanism that needs no task-specific reward engineering makes real-world deployment meaningfully cheaper. It also strengthens OpenAI's position in the basic-research layer that underlies agent capabilities, where such methods tend to diffuse quickly across the field rather than remain proprietary.
COVERAGE · 1 REPORT · LINKS GO TO THE ORIGINAL OUTLETS