LIVE PULSE
4.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.2 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.0 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src1.8 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.4 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.1 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.1 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.1 Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election1 src1.1 Study Audits Misalignment in Multi-Modal World Models1 src1.1 Retrieval-Grounded Reasoning Approach Proposed for Universal Multimodal Embeddings1 src4.0 Anthropic CEO Amodei calls for slower AI development and shared safety rules11 src2.2 Agility Robotics unveils Digit 5 humanoid for warehouses and factories2 src2.0 Apple ships rebuilt Siri with Google Gemini, but not in the EU2 src1.8 Siri AI in macOS 27 Golden Gate: FAQ, Germany availability, privacy questions2 src1.4 Sam Altman says OpenAI will not go public in 2026, citing AI safety concerns5 src1.1 OpenAI contractors review real ChatGPT conversations to rate responses, report says2 src1.1 Anthropic data retention policy prompts firms to limit Claude use for sensitive work1 src1.1 Study examines issue bias in LLMs used as writing assistants before Swedish 2026 election1 src1.1 Study Audits Misalignment in Multi-Modal World Models1 src1.1 Retrieval-Grounded Reasoning Approach Proposed for Universal Multimodal Embeddings1 src
HEATPULSEAI MAGAZINES
FLIP · FOLLOW · SAVE

cs.AI

topic34 events
papersTODAY 04:00 UTC

arXiv paper targets controllable speech generation with nonverbal vocalizations

A new arXiv preprint addresses the difficulty of synthesizing nonverbal vocalizations such as laughs, sighs, and coughs in controllable speech generation. The authors attribute the challenge to the acoustic variety of these sounds and their uneven representation in existing speech corpora, and propose modeling, scaling, and decoding methods to improve them. The work is listed under the cs.AI cross-list announcement.

papersTODAY 04:00 UTC

Paper Proposes Eliciting Skill Routing Directly from a Frozen LLM

A new arXiv paper argues that current agent frameworks pick skills by loading all skill metadata into the context window, which spreads the model's attention thin and limits how many skills can be offered. The authors instead describe a method for surfacing routing behavior that already exists inside a frozen language model, avoiding that metadata overhead. The work appears under arXiv identifiers 2609.15982v1 in both the cs.AI and cs.LG listings.

papersTODAY 04:00 UTC

arXiv Paper Proposes Sandboxed Execution Environment for AI Agents Handling Private Data

A new arXiv paper describes a sandboxed execution environment designed to let AI agents use personal and financial data without exposing it to the underlying model. The approach aims to limit leakage and misuse by isolating agent operations from raw user information. It is framed as a cross-listed replacement submission on arXiv's cs.AI category.

papersTODAY 04:00 UTC

arXiv Paper Benchmarks Intra-Patient 3D Deformable Multimodal Image Registration

A new arXiv preprint introduces a benchmarking study for 3D deformable multimodal image registration focused on aligning scans from the same patient. The work addresses the difficulty that matching anatomical structures appear with very different intensities across imaging modalities, which complicates clinical registration workflows. It is announced as a cross-listed submission in the cs.AI category.

papersTODAY 04:00 UTC

CLQT benchmark targets diagnostic evaluation of LLM portfolio-management agents

A new arXiv paper introduces CLQT, a closed-loop, cost-aware and strategy-consistent benchmark for evaluating LLM agents that manage investment portfolios. The authors argue that ranking agents by returns over a fixed window fails to show whether their process is sound or their performance durable, and propose a diagnostic alternative. The work is cross-listed in cs.AI and cs.LG as a replacement submission.

papersTODAY 04:00 UTC

Paper Tackles Lost-in-the-Middle Problem in Long-Text Generation

A new arXiv paper addresses how large language models tend to ignore information placed in the middle of long contexts, a problem studied mostly for retrieval tasks rather than long-input-to-long-output generation. The authors introduce a synthetic dataset and evaluation framework for this setting and propose a mitigation approach. The work is a revised cross-listing (v2) on arXiv cs.AI.

papersTODAY 04:00 UTC

Study Tests Whether Manager Agents Should Have Power to Reject Worker Output

A new arXiv paper reports a paired experiment comparing flat and hierarchical coordination in multi-agent LLM teams, where a Manager agent can review and return worker output for revision. The authors draw on classical organizational theory to examine how such loop-back authority affects team performance. The work is announced as a cross-listing on arXiv cs.AI.

papersTODAY 04:00 UTC

arXiv Paper Surveys Diffusion Language Models for Code Generation

A new arXiv preprint reviews how diffusion-based large language models can be applied to code generation, an area currently dominated by left-to-right autoregressive decoding. The authors examine the limitations of standard autoregressive generation and assess whether diffusion approaches offer advantages for producing source code. The work is a replacement submission (v3) to the cs.AI category.

papersTODAY 04:00 UTC

Paper Examines Adequacy-Fluency Tradeoff in MT Meta-Evaluation

A new arXiv paper analyzes how meta-evaluation of machine translation must balance alignment with adequacy versus fluency, noting that the preferred balance shifts depending on which translation systems are included in the evaluation set. Because those system sets are typically small and filtered, the authors propose parameterizing this balance explicitly. The work appears in the cs.AI and cs.LG cross-listings.

papersTODAY 04:00 UTC

Paper Introduces Robust Communication Method for Multi-Agent Reinforcement Learning

A new arXiv preprint presents a method for making the messages exchanged between agents in multi-agent reinforcement learning both informative and resilient to physical constraints. The work targets distributed intelligence settings where learned communication must stay reliable under real-world limitations. It is listed under both cs.AI and cs.LG.

papersTODAY 04:00 UTC

New benchmark tests multi-turn prompt injection attacks on LLM agents

Researchers released a 21-scenario benchmark for evaluating how well LLM agents resist adaptive, cross-session attacks from an autonomous LLM attacker. The setup pits an attacking model against defenders that start each session fresh, targeting prompt injection and multi-turn manipulation risks. The work appears on arXiv as a cross-listing in cs.AI and cs.LG.

papersTODAY 04:00 UTC

arXiv Paper Proposes Efficient Personalization for Generative User Interfaces

A new arXiv paper examines how generative user interfaces (GenUIs), which build interface layouts on demand, can be tailored to individual users. The authors note that conventional personalization through predefined settings is impractical because screens are created dynamically rather than designed in advance. Their work proposes an efficient approach to personalizing these generated interfaces, and the preprint is cross-listed under cs.AI and cs.LG.

papersTODAY 04:00 UTC

arXiv Paper Diagnoses and Improves Visual Chain-of-Thought for Geometry Solvers

A revised arXiv preprint argues that multimodal models need active visual assistance, such as drawing auxiliary lines, to handle complex geometry problems. The authors examine shortcomings in current evaluation of visual chain-of-thought methods and propose ways to strengthen how models reason with diagrams. The work falls under cs.AI and focuses on diagnosing and improving these visual reasoning pipelines.

papersTODAY 04:00 UTC

LLaTSA: general-purpose transient stability analysis aligned with LLMs

A new arXiv paper introduces LLaTSA, a method that adapts large language models to predict dynamic trajectories for power-system transient stability analysis. Most existing data-driven predictors are tied to a specific grid and need retraining when network topology or generation mix changes, whereas this approach aims to generalize across configurations. The work is a cross-listing on cs.AI and falls under research rather than a released product.

papersTODAY 04:00 UTC

arXiv Paper Reexamines Human Feedback for Robot Preference Learning

A new arXiv paper examines how robots typically build reward models of human preferences, a process that starts with collecting limited direct feedback such as positive or negative signals. The authors argue that the assumptions behind this standard three-step pipeline deserve reconsideration in the context of human-robot collaboration. The work is cross-listed in the cs.AI category.

papersTODAY 04:00 UTC

arXiv Paper Trains Humanoid Robot to Play Badminton with Human-Like Skills

A research team has developed a method that lets a humanoid robot acquire badminton skills resembling human play. The work addresses the difficulty of combining fast, explosive movement with precise racket control, which differs from ordinary walking or stationary manipulation tasks. The paper is posted on arXiv as a cross-listing replacement in the cs.AI and cs.LG categories.

papersTODAY 04:00 UTC

Valley3: Omni Multimodal LLM Targets Global E-commerce Tasks

Researchers introduce Valley3, a multimodal large language model designed for e-commerce applications across different markets. The model handles text, images, video, and audio within a single framework, aiming to combine understanding and reasoning abilities across those modalities. The paper is a revised arXiv preprint, with an updated version posted to the cs.AI category.

papersTODAY 04:00 UTC

Paper Proposes Using Model Internals to Predict Behavior on Unseen Data

A new arXiv paper reframes interpretability research around predicting how a model will respond to previously unseen inputs, rather than only to targeted mechanistic interventions. The authors use a model's internal representations to forecast its out-of-distribution behavior. The work appears in two arXiv listings, cs.AI and cs.LG, as a replacement submission.

papersTODAY 04:00 UTC

arXiv Paper Proposes Forking Garden Framework for Whole-Game Generation

A new arXiv paper argues that generating complete games is mainly an orchestration challenge, since narrative, level design, encounters, objectives, rewards, and visuals all need to share a common intent. The authors introduce Forking Garden, which threads narrative archetypes through gameplay planning as a semantic signal to keep these layers aligned. The work appears as a cross-listing replacement on arXiv cs.AI.

papersSEP 12 04:00 UTC

arXiv paper proposes importance weighting for unlabeled-unlabeled learning under distribution shift

A new arXiv preprint addresses unlabeled-unlabeled (UU) learning, a setting where a binary classifier is trained from two unlabeled datasets that have differing class priors. The authors introduce importance weighting to handle distribution shift in this framework, which generalizes approaches such as positive-unlabeled learning. The work is listed as a cross-submission in the cs.AI category.

papersSEP 12 04:00 UTC

Continuous-Time TTS Acoustic Modelling with Neural Controlled Differential Equations

This preprint proposes modelling text-to-speech acoustics in continuous time using neural controlled differential equations, rather than the usual approach of stretching phone-level encoder states to frame-level decoder inputs via predicted durations. The authors argue that length regulation fixes alignment structurally but leaves duration handling as a separate, discrete step. The work is a cross-listed arXiv submission in the cs.AI category.

papersSEP 12 04:00 UTC

AI Model Estimates Flare Stack Combustion Efficiency

A new arXiv paper presents an AI-based approach for estimating combustion efficiency in industrial flare stacks, a measure tied to regulatory compliance and hydrocarbon emissions. The authors note that conventional tools such as gas analyzers and hyperspectral cameras are costly, motivating a lower-cost alternative. The work appears as a new submission in the cs.AI category.

papersSEP 12 04:00 UTC

Benchmark and Method Proposed for Think-with-Video Reasoning in Generative Models

A new arXiv paper argues that while video generation models now produce convincing and temporally consistent output, it is unclear whether they can reason through video by following symbolic rules, obeying physics, and working toward defined goals. The authors introduce a benchmark for measuring this think-with-video ability and propose an approach for improving it. The work is listed under the cs.AI cross-submission category.

papersSEP 12 04:00 UTC

ExpTest Uses Loss-Curve Hypothesis Testing to Pick Learning Rates Automatically

A research paper proposes ExpTest, a method that selects learning rates for deep neural networks on its own by statistically testing the shape of the training loss curve. The approach aims to reduce the manual searching and expensive grid searches that currently make hyperparameter tuning costly and less accessible. It is presented as part of ongoing work on arXiv in the cs.AI category.

papersSEP 12 04:00 UTC

Study tests whether LLMs can reverse fact-preserving news framing

A new arXiv paper introduces a controlled inversion test that checks whether large language models can undo fact-preserving framing in news text, rather than just generating or detecting it. The authors argue prior framing research focuses on generation, detection, or apparent neutrality gains, and does not directly establish whether models can invert a framing transformation. The work is posted as a cross-listing on arXiv cs.AI.

papersSEP 12 04:00 UTC

arXiv paper introduces GitSkills, a dataset of agent skills collected from GitHub

A new arXiv paper presents GitSkills, a dataset built from GitHub repositories that package agent skills as folders containing a SKILL.md instruction file, sometimes with helper scripts and reference material. The work focuses on skills that language-model agents load when they decide a task matches a skill's description. It is a replacement submission (v2) in the cs.AI category.

papersSEP 10 04:00 UTC

Paper Examines Equity-Aware Online Allocation of Scarce Resources by Nonprofits

A revised arXiv preprint explores how nonprofit bodies, such as government agencies, can distribute limited resources in real time as demand arrives. The study emphasizes internal equity, aiming to ensure fair treatment across the parties seeking resources. The updated version (v3) is cross-listed in the cs.AI category.

papersSEP 10 04:00 UTC

MOONWALK: AI framework for intent-evidence-action alignment in animation/VFX review workflows

A new arXiv preprint presents MOONWALK, a framework for animation and VFX pre-production that turns loosely defined creative direction into revisions junior artists can execute. It mediates the junior-supervisor review loop by aligning stated intent, supporting evidence, and resulting actions to cut down on repeated clarification. The work is cross-listed in the cs.AI category.

papersSEP 10 04:00 UTC

HiRAD: A Flexible Routing System for Large-Scale AGV Fleets

Researchers have introduced HiRAD, a routing system aimed at coordinating large fleets of Automatic Guided Vehicles in warehouse environments. The preprint addresses the combinatorial blowup and super-quadratic runtimes that limit classical multi-agent pathfinding solvers at scale. The work is cross-listed on arXiv under the cs.AI category.

papersSEP 10 04:00 UTC

Teacher Geometry Shapes Learnability in Teacher-Student Networks

An arXiv preprint studies teacher-student frameworks, where one neural network produces training data for another network that must learn to reproduce its behavior, a standard abstraction in learning theory. The authors find that the geometric structure of the teacher network plays a decisive role in determining whether and how well the student can learn the target function. The paper was posted as a new submission and cross-listed in the cs.AI and cs.LG categories.

papersSEP 10 04:00 UTC

Paper Compares Subagents and Agent Skills for Long-Horizon Agentic Tasks

A new arXiv paper investigates how language model agents can draw on libraries of reusable knowledge when tackling long-horizon tasks. It contrasts two approaches—subagents and agent skills, where skills are packaged as multi-file bundles—and examines which executes such knowledge more effectively. The study was announced in the cs.AI category and cross-listed in cs.CL and cs.LG.

papersSEP 10 04:00 UTC

Research Paper Addresses Query Brand Entity Linking for E-Commerce Search

A new version of an arXiv paper (2502.01555) explores how to connect short user search queries with the correct brand entities during e-commerce product retrieval. The authors highlight the difficulty of the task, since queries average only three to four words and lack grammatical structure. The work is cross-listed in the cs.AI and cs.LG categories.