对 OpenAI / Hugging Face 入侵事件中智能体行为、推理与协作的简要独立调查
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
METR 独立调查 OpenAI 智能体在 Hugging Face 入侵事件中的行为、推理与协作,并说明调查范围、证据来源和局限性。
Hacemos accesibles los textos sobre seguridad de la IA escritos en inglés a lectores de todo el mundo.
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
METR 独立调查 OpenAI 智能体在 Hugging Face 入侵事件中的行为、推理与协作,并说明调查范围、证据来源和局限性。
How independent researchers could investigate AI propensities after misalignment incidents
探讨如何独立调查未对齐事件背后的 AI 行为倾向,包括应回答的核心问题、所需的访问权限与资源,以及调查结果的共享方式。
Nuestras traducciones publicadas están disponibles actualmente en chino. ¡Damos la bienvenida a colaboradores que trabajen en otros idiomas!
Ayúdanos a acercar la investigación sobre seguridad de la IA a más personas, en más idiomas.