Anthropic flags gaps in AI guardrails as models grow more capable: Details

Anthropic in its risk report says some AI models may recognise when they are being evaluated and alter their behaviour, potentially making it harder to judge their capabilities and real-world safety
Read more at the source

Disclaimer: The content of this post is sourced from external sites and is for informational purposes only. All rights and credits belong to the original authors and publishers.