LIVE PULSE
3.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.1 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.0 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src1.7 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.4 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.1 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.1 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.1 Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election1 src1.1 Study Audits Misalignment in Multi-Modal World Models1 src1.1 Retrieval-Grounded Reasoning Approach Proposed for Universal Multimodal Embeddings1 src3.9 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.1 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.0 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src1.7 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.4 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.1 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.1 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.1 Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election1 src1.1 Study Audits Misalignment in Multi-Modal World Models1 src1.1 Retrieval-Grounded Reasoning Approach Proposed for Universal Multimodal Embeddings1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

multimodal-llms

topic11 events
papersTODAY 04:00 UTC

MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding

Researchers propose MarKey, a method that picks keyframes for long-video question answering using a greedy strategy driven by marginal utility rather than dense encoding or uniform sampling. The approach targets multimodal large language models, where encoding every frame is costly and even sampling can overlook brief but important moments. The paper is an arXiv preprint and appears under a cross-listing.

papersTODAY 04:00 UTC

VideoScout Agent Explores Long Videos With Adaptive Reasoning Pacing

A new arXiv paper introduces VideoScout, a method that lets multimodal language models actively explore long videos instead of relying on uniform frame sampling. The approach pairs agentic search with adaptive reasoning pacing to work around limited visual context windows. It targets the problem of long-video understanding, where current models still lag behind their short-video performance.

papersTODAY 04:00 UTC

FriendBench Benchmark Tests Whether AI Can Tell Friends From Strangers

Researchers introduced FriendBench, a benchmark that evaluates how well humans and multimodal large language models can judge whether two people in a short video clip are already acquainted or meeting for the first time. The task uses 20-second recordings of ice-breaker conversations, where cues come from behavior and body language rather than spoken content alone. The work aims to measure social perception abilities that go beyond text-based reasoning.

papersTODAY 04:00 UTC

TimeThink: Method Aims to Improve Compositional Reasoning in Time-Series LLMs

A new arXiv paper introduces TimeThink, a technique intended to help time-series multimodal large language models reason more compositionally. The authors note that such models often struggle to capture dynamic temporal patterns when answering questions. The work focuses on eliciting stronger reasoning behavior from these models rather than treating forecasting as pure pattern matching.

papersTODAY 04:00 UTC

ChartAnno Benchmark Tests Multimodal LLMs on Chart Annotation Generation

A new research benchmark called ChartAnno evaluates how well multimodal large language models can generate annotations for charts, a task that helps explain data and highlight key findings in visualizations. The work examines whether these models can automate annotation authoring, which is normally done by hand. It is presented as an arXiv paper revision.

papersTODAY 04:00 UTC

Func-R1: Method Aims to Improve Mathematical Function Reasoning in Multimodal LLMs

A new arXiv paper introduces Func-R1, an approach aimed at strengthening mathematical function reasoning in multimodal large language models. The work targets the challenge of combining visual perception with symbolic logic when solving math problems from images. The abstract frames deliberate mathematical reasoning in visual settings as an indicator of advanced multimodal model capability.

papersSEP 12 04:00 UTC

ReactHuman Benchmark Tests Reactive Decision-Making in Embodied Multimodal LLMs

A new arXiv paper introduces ReactHuman, a physics-grounded benchmark designed to evaluate how well embodied multimodal large language models handle sudden physical hazards. The tasks include scenarios such as catching a slipping plate or dodging a falling knife, which the authors frame as both a test of embodied intelligence and a prerequisite for using MLLMs as decision cores in household robots. The work is listed as a cross-submission announcement in arXiv's cs.AI category.

papersSEP 10 04:00 UTC

S3-Bench: New Benchmark Tests Speech Models as Scientific Voice Assistants

Researchers have released S3-Bench, a benchmark aimed at measuring how well speech interaction models function as voice assistants for scientific work. The evaluation focuses on multimodal large language models, examining whether their conversational strengths extend beyond general-purpose assistant tasks to domain-specific spoken interactions.

papersSEP 10 04:00 UTC

Sample-Adaptive Strategy Routing Improves Vision Token Pruning in Multimodal LLMs

A new arXiv preprint proposes choosing among vision token pruning strategies on a per-image basis rather than applying one fixed policy to every input. By adapting the pruning approach to each sample, the method aims to lower the heavy inference costs that multimodal language models incur from processing large numbers of visual tokens while maintaining output quality.

papersSEP 10 04:00 UTC

Survey Reviews Inference-Efficiency Methods for Video and Audiovisual LLMs

A new survey on arXiv examines mechanisms for reducing inference costs in video large language models, which pair video representations with pretrained LLMs to generate responses from text prompts. The paper addresses why video understanding remains computationally expensive and organizes existing efficiency techniques across video and audiovisual tasks.

papersSEP 10 04:00 UTC

Predicting Middle-Layer Attention in Multimodal LLMs for Efficient Visual Token Pruning

Multimodal large language models spend significant compute processing large numbers of visual tokens, and effective pruning depends on knowing which tokens actually matter. This paper introduces a learned approach that predicts attention at middle layers, enabling models to identify and drop less relevant visual tokens. The method aims to cut inference costs while maintaining performance across vision-language tasks.