papersSEP 10 04:00 UTC
Self-play in code distills a text harness for black-box optimization, arXiv study finds
Researchers explore whether a language-model agent can acquire a numerical search strategy through executable practice and then transfer it as plain text. Targeting low-budget black-box optimization, where unaided LLMs fall short of strong classical optimizers, the approach uses self-play in code to automatically build the harness. The work suggests learned optimization behavior can be distilled into reusable text prompts instead of manually engineered ones.