LIVE PULSE
3.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.1 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.0 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src1.7 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.4 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.1 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.1 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.1 Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election1 src1.1 Study Audits Misalignment in Multi-Modal World Models1 src1.1 Retrieval-Grounded Reasoning Approach Proposed for Universal Multimodal Embeddings1 src3.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.1 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.0 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src1.7 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.4 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.1 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.1 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.1 Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election1 src1.1 Study Audits Misalignment in Multi-Modal World Models1 src1.1 Retrieval-Grounded Reasoning Approach Proposed for Universal Multimodal Embeddings1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

AI benchmarking

topic9 events
papersTODAY 04:00 UTC

PAI-Bench: Benchmark Measures Persistent Identity in Deployed AI Agents

A new arXiv paper introduces PAI-Bench, a provider-neutral benchmark designed to test how faithfully AI agents adhere to a versioned identity contract that can be updated under governance rules. The work argues that existing evaluations conflate an agent's ability to recall identity facts with its ability to express and act on them. The benchmark separates recall from expression and enactment, aiming to give a clearer picture of identity persistence in deployed agents.

papersTODAY 04:00 UTC

FaithfulBench benchmark measures how well AI advice matches users' religious beliefs

Researchers introduced FaithfulBench, described as the first benchmark for evaluating AI moral counsel based on how closely it aligns with a user's stated faith. It scores assistant responses across multiple religious traditions using scenarios built around moral dilemmas. The work appears on arXiv in the cs.AI and cs.CL categories.

papersTODAY 04:00 UTC

Turkish MMLU Pro Benchmark Examines Limits of Adding Answer Options

A new arXiv paper introduces Turkish MMLU Pro, a benchmark built from 12,000 Turkish-language questions spanning 58 sections. Each item keeps its original question stem and five answer choices, allowing researchers to test whether adding more options actually improves measurement quality. The authors argue that extra options can reduce scores without making the assessment more valid.

papersTODAY 04:00 UTC

Paper argues behavioral consistency is a distinct, measurable property of LLM agents

A revised arXiv preprint contends that current agent evaluations lean almost exclusively on outcome measures like success rate, which show whether an agent finishes a task but not how uniformly it behaves. The authors propose treating consistency of behavior across different tasks as its own measurable characteristic. Their work offers a way to evaluate agents beyond simple pass/fail scores.

papersSEP 12 04:00 UTC

Study Evaluates Edge-Deployable Vision-Language Models for Species ID

A new arXiv paper argues that species identification from camera traps should be assessed using small vision-language models that can run locally on edge hardware, rather than frontier-scale systems. The authors note that field deployments often have weak or no network connectivity, which makes compact, on-device models the realistic option to study. The work positions this evaluation setting as the practically relevant benchmark for the task.

papersSEP 10 04:00 UTC

Psychometric audit finds MMLU aggregate scores mainly measure factual retrieval, not reasoning

A new arXiv paper applies psychometric methods to the MMLU benchmark, analyzing how question difficulty is distributed across its aggregate score. The authors conclude that the headline number primarily reflects a model's ability to recall facts, providing limited signal about reasoning skill. The results caution against relying on MMLU alone as a measure of general AI capability.

papersSEP 10 04:00 UTC

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

A revised arXiv paper introduces a benchmark designed to test AI agents on tasks outside the well-known applications that dominate current evaluations. The authors contend that testing in familiar, comparatively simple settings can mask how poorly agents generalize to novel situations. The work aims to give a more accurate picture of how agentic systems will behave in real-world deployment.

papersSEP 10 04:00 UTC

Researchers Propose Discovery Certification Protocol for Auditing AI Research Agents

An arXiv paper argues that benchmark scores by themselves are not sufficient evidence that AI research agents have made genuine scientific discoveries. The authors introduce the Discovery Certification Protocol, which converts an agent's claimed results into executable recovery and feedback tests that can be independently run. The first of its gates checks that reported results can actually be reproduced from the recorded evidence.