Co-founder of robot benchmarks company says OpenAI’s flagship AI model attempted 97% of harmful tasks

Safety benchmark platform Robocurve evaluated general-purpose AI models on controlling robotic arms across hazardous scenarios, revealing a significant gap between text-based chat safety and physical action control. The evaluations highlight growing concerns over “embodied AI,” showing that while frontier models demonstrate high competence in task execution and physical control efficiency, standard conversational safety alignment does not automatically translate into safety guardrails for physical robot task-planning.
Read more at the source

Disclaimer: The content of this post is sourced from external sites and is for informational purposes only. All rights and credits belong to the original authors and publishers.