papersTODAY 04:00 UTC
PhysMent benchmark evaluates LLM physics reasoning through interactive experiments
Researchers introduced PhysMent, a benchmark designed to test how well large language models reason about physical systems by running experiments rather than answering static questions. The work argues that strong scores on existing science benchmarks do not show whether models can actively probe the physical world. The abstract notes that this ability remains poorly understood.