LIVE PULSE
4.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.2 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.0 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src1.8 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.4 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.1 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.1 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.1 Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election1 src1.1 Study Audits Misalignment in Multi-Modal World Models1 src1.1 Retrieval-Grounded Reasoning Approach Proposed for Universal Multimodal Embeddings1 src4.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.2 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.0 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src1.8 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.4 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.1 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.1 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.1 Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election1 src1.1 Study Audits Misalignment in Multi-Modal World Models1 src1.1 Retrieval-Grounded Reasoning Approach Proposed for Universal Multimodal Embeddings1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

ai-benchmarks

topic24 events
papersTODAY 04:00 UTC

TwinICL Benchmark Tests Multimodal In-Context Learning With Paired Counterfactuals

Researchers released TwinICL, a procedurally generated benchmark that provides matched text and image versions of the same tasks, allowing direct comparison of in-context learning across modalities. The paired counterfactual design is intended to isolate how much a model's few-shot performance depends on the input format rather than the task itself. The work appears on arXiv under cs.LG.

papersTODAY 04:00 UTC

K-Bench: clinician-calibrated benchmark for LLM safety in high-risk mental health chats

Researchers introduced K-Bench, a benchmark designed with clinician input to assess how large language models handle high-risk mental health conversations that escalate over time. The work addresses the limited understanding of LLM safety in these evolving support dialogues, where users increasingly turn for help. The benchmark provides a protected evaluation framework for measuring model performance in these sensitive settings.

papersTODAY 04:00 UTC

Study Examines Limits of Agentic ICD-10-CM Coding Benchmarks

A new arXiv paper analyzes how well agentic systems perform on ICD-10-CM medical coding, the alphanumeric codes used in the US for diagnoses, billing, and epidemiology. The authors argue that standard benchmarks rely on aggregate scores that hide poor performance on harder coding scenarios. The work aims to expose where current evaluation practices fall short for complex cases.

papersTODAY 04:00 UTC

Robusto-2 Benchmark Tests Vision-Language Models for Self-Driving in Lima and New York

A new arXiv paper introduces Robusto-2, a benchmark evaluating both humans and vision-language models on autonomous driving tasks in Lima, Peru and New York City. The work targets how well multi-modal systems generalize when deployed in unfamiliar, out-of-distribution urban environments. It is a cross-listed replacement submission on arXiv cs.AI.

papersTODAY 04:00 UTC

Benchmark Measures AI Agents' Ability to Locate Security Flaws in Code Repos

A new arXiv paper introduces a benchmark that tests whether language-model agents can pinpoint the specific code responsible for a vulnerability across an entire software repository. Existing cybersecurity evaluations mostly check if agents can detect, reproduce, or patch flaws, leaving location ability largely unmeasured. The work targets repository-scale settings, where agents must search large codebases rather than isolated snippets.

papersTODAY 04:00 UTC

NoteVQA Benchmark Targets Everyday Visual Questions From Human Communities

Researchers introduce NoteVQA, a benchmark that evaluates vision-language models on visual questions drawn from real human communities rather than pre-defined task categories. The work argues that current benchmarks focus on narrow capabilities such as multi-hop retrieval and miss the variety of questions users ask in daily life, including consumer AI search. It is published as an arXiv preprint.

papersTODAY 04:00 UTC

DepthBenchCAD Examines Whether More Auditing Checks Improve Generative CAD Evaluations

A new arXiv preprint introduces DepthBenchCAD, a benchmark studying how the number of edit checks affects the reliability of evaluations for generative CAD models. The work focuses on behavioral correctness after parameter edits and asks whether auditing more programs under a fixed budget actually leads to firmer conclusions. It questions the common assumption that adding edit checks is a straightforward path to more trustworthy evaluation.

papersTODAY 04:00 UTC

GroundBench benchmark aims to pinpoint where vision-language models fail on affordance tasks

A new arXiv paper introduces GroundBench, described as a factorized, counterfactual benchmark for identifying the specific points at which vision-language models break down on affordance tasks. The work cites a companion evaluation in which explicitly naming the target part in a manipulation prompt improved action accuracy by 0.32 to 0.63 across eight vision-language models, and no model exceeded a constant baseline before that part was named. The benchmark is intended to isolate these failures rather than report only aggregate scores.

papersYESTERDAY 16:32 UTC

Discussion: Why machine learning research agents do not overfit

A Hacker News thread explores why autonomous agents that carry out machine learning research tend not to overfit their results, unlike typical human-run experiment loops. Commenters compare how these systems generate, evaluate, and discard candidate models, and question whether current benchmarks hide overfitting. The conversation also considers how evaluation harnesses and search procedures shape the conclusions agents report.

papersSEP 12 04:00 UTC

arXiv paper proposes evaluating AI agents on resilience across repeated interactions

A new arXiv preprint argues that measuring whether an agent completes a single task is insufficient for judging fitness in long-running deployments. The authors propose evaluating agents on how well they hold up as challenges accumulate, including shifting conditions, repeated interactions, and reliance on human collaborators in shared workflows. The work frames resilience and considerate participation as dimensions that need dedicated benchmarks.

papersSEP 12 04:00 UTC

Paper proposes measuring implicit conventions in cooperative AI evaluation

A new arXiv paper argues that benchmarks testing cooperation between AI agents may overlook the unwritten conventions that let humans infer meaning beyond literal messages. The authors introduce the concept of a "convention gap" and outline an approach for quantifying implicit communication in cooperative AI evaluations. The work is a research proposal rather than a released model or tool.

papersSEP 11 04:00 UTC

arXiv Paper Examines AI Inference Optimization Across Deployment Stack

A new arXiv preprint argues that AI deployment performance depends on how compression methods, compiler transformations, and serving policies interact, rather than on model architecture alone. It notes that existing benchmarks often report latency and throughput under conditions that cannot be directly compared, which limits practical conclusions. The work appears to be a cross-listed submission surveying the inference deployment stack.

papersSEP 10 04:00 UTC

S3-Bench: New Benchmark Tests Speech Models as Scientific Voice Assistants

Researchers have released S3-Bench, a benchmark aimed at measuring how well speech interaction models function as voice assistants for scientific work. The evaluation focuses on multimodal large language models, examining whether their conversational strengths extend beyond general-purpose assistant tasks to domain-specific spoken interactions.

papersSEP 10 04:00 UTC

LexAgentHallu: a hierarchical benchmark for hallucinations in legal AI agents

Researchers have introduced LexAgentHallu, a new benchmark for measuring how tool-augmented legal AI agents hallucinate. It uses a hierarchical structure to trace how errors in tool calls and reasoning cascade into fabricated case holdings and miscited legal authority. The benchmark aims to fill a gap left by existing legal evaluations that do not capture agentic workflows.

papersSEP 10 04:00 UTC

MetroLLM-Bench: New Benchmark Tests LLMs as Public Transit Kiosk Assistants

Researchers have released MetroLLM-Bench, a set of 955 test cases that measures how well language models can act as the decision-making layer of a metro station kiosk. The benchmark draws on six real subway networks of varying sizes and covers eleven task types, including route planning, fare computation, and responding to service disruptions. The paper was posted to arXiv and cross-listed in the AI, computation and language, and machine learning categories.

papersSEP 10 04:00 UTC

Sim2Signal: Sim-to-Real Benchmarks for RL-Based Traffic Signal Control

A new arXiv paper introduces Sim2Signal, a benchmark suite for measuring how reinforcement learning policies for traffic signal control transfer from simulators to real-world deployment. The work targets the sim-to-real gap, where agents that perform well in simulation often fail once applied in practice. The benchmarks aim to give researchers a standardized way to evaluate transferability before deployment.

papersSEP 10 04:00 UTC

Benchmarking Hybrid Deep Research Across Database Querying and Web Search

A new arXiv paper presents a benchmark that evaluates AI research agents on tasks requiring both structured database querying and open-web exploration. The authors note that real analytical work rarely stays within a single environment, so their setup measures how well agents combine data retrieval from databases with web-based information gathering. The result offers a standardized testbed for comparing hybrid deep-research capabilities.

papersSEP 10 04:00 UTC

AgentAudit: Open Framework Evaluates AI Agents Across Their Full Lifecycle

A new arXiv paper introduces AgentAudit, an open and extensible framework for auditing the trustworthiness of AI agents. The authors contend that today's benchmarks examine only slices of agent behavior, such as task success or robustness against attacks, and instead propose measuring every stage of an agent's operation, including planning, tool use, memory, and reasoning. The design is meant to be extendable so that new evaluation checks can be added over time.

papersSEP 10 04:00 UTC

EVA-Bench: An End-to-End Framework for Evaluating Voice Agents

A new research paper introduces EVA-Bench, a benchmark for assessing voice agents across the entire interaction pipeline. It combines simulated conversations that mimic real usage with metrics tailored to voice-specific behaviors, filling a gap left by earlier evaluation tools that handled these aspects separately. The work responds to the growing deployment of voice agents in enterprise applications.

papersSEP 10 04:00 UTC

New protocol IBIB scores enterprise AI deployments by serving route rather than model identifier

A cross-listed arXiv paper introduces IBIB, a protocol for evaluating AI systems as they are actually deployed inside enterprises rather than as bare model checkpoints. The authors argue that real-world capability emerges from the combination of weights, serving configuration, precision, output contract, and harness, so scoring an advertised model name alone is a measurement error. After auditing 18 existing benchmarks and finding that every one grades model identifiers, the paper proposes routing-based measurement instead.

papersSEP 10 04:00 UTC

FrontierChallenge: New Benchmark for Evaluating AI Agents on Scientific Workflows

Researchers have introduced FrontierChallenge, a benchmark of 300 tasks spanning multiple scientific disciplines that measures whether AI agents can carry out complete research workflows. It goes beyond existing evaluations that score only final answers, standalone programs, or work within a single field, instead assessing capabilities like data processing, coding, and producing research artifacts.

papersSEP 10 04:00 UTC

Reference-based method audits LLM bias via relative representations of hidden states

An arXiv paper in cs.AI introduces a technique for auditing bias in large language models by analyzing internal hidden states instead of relying on generated outputs. By comparing a model's representations against those of a reference model using relative representations, the approach aims to detect internal bias shifts that output-based benchmarks or judge models could miss. The authors frame it as a cheaper alternative to benchmark-heavy or judge-dependent auditing pipelines.