papersYESTERDAY 16:18 UTC
MIT researchers develop method to make AI models follow safety rules
Researchers at MIT have introduced a technique intended to keep AI systems compliant with prescribed safety rules rather than evading or overriding them. The work sits within ongoing alignment and guardrail research, which seeks ways to constrain model behavior within defined limits.