papersSEP 10 04:00 UTC
Researchers propose grounded evaluation and repair for LLM-generated PDDL planning problems
A new arXiv paper examines how large language models convert natural-language planning descriptions into PDDL problem instances, arguing that common checks like syntactic validity or planner success can overstate actual quality. The authors introduce an evaluation and repair framework that grounds assessment more firmly in the underlying planning task to better catch and fix flawed outputs.