{"id":87284,"topic":"ai","source":"Welcome to the United Nations","title":"Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control - Welcome to the United Nations","url":"https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks","url_hash":"ab11d02ac45216bb9393f9316e7970a09112ef92","author":"","summary":"<a href=\"https://news.google.com/rss/articles/CBMisgFBVV95cUxQelgycTJ1SVBXYnJRSWdrdnFaS1VzaTBBWkt3WG9WZnNkOVN6d25TenhLdU5jLTN4c1dZcjVIRDBTR25aVjZkZGJwQlhqdEZncFFodXpDdm1ORDNhN3JwV0Y4MWNPUkMwSjBMM2RkMVlOVTNObGtYTVlHdnh4U0pNbG5kRUlSYVNtOXpqRUNOOE93QnIwdGEtTXAzWkYtVWN1WVFDMkU4T3I1LTVNemxlWDFn?oc=5\" target=\"_blank\">Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control</a>&nbsp;&nbsp;<font color=\"#6f6f6f\">Welcome to the United Nations</font>","content":"The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents pursuing goals that conflict with human intentions.\nBetween May and July 2026, AI agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI’s and Hugging Face’s systems. No human directed the individual steps.\nDrawing on disclosures by both companies, an independent investigation by METR and wider research, the brief finds that greater capability can help misaligned systems find loopholes and conceal their actions. It does not estimate the probability or timing of severe loss of control, but notes that stopping this activity does not demonstrate that humans will retain control over more capable agents.\nBuilding on the Panel’s Preliminary Report, the brief explains how training can give rise to misaligned goals and behaviours, including reward hacking and reward tampering. It notes that AI failures can cross company and national borders, and that no single organisation or country sees enough incidents to identify every emerging pattern. Rather than issuing recommendations, the brief reviews approaches used in fields such as aviation, nuclear power, and cybersecurity as possible options for decision-makers.\nRead the brief\nThis brief is published as an advance unedited version. Updated versions will be posted at this link, with earlier versions listed below.","image_url":null,"lang":"en","published_at":"2026-09-21T10:20:02+00:00","fetched_at":"2026-09-21T15:15:05+00:00","status":"read","starred":0,"extract_state":"ok","summary_auto":"The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents pursuing goals that conflict with human intentions. Between May and July 2026, AI agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI’s and Hugging Face’s systems.","cluster_id":null,"extract_retries":0,"extract_error":null,"contract_version":"news_item.v1","format_contract_version":"news_item_formats.v1","dedup_url":"https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 1710 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":1710,"summary_length":604,"usable_text_length":1710,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":1710,"summary_length":604}},"news_item":{"id":87284,"canonical_url":"https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks","source_url":"https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks","title":"Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control - Welcome to the United Nations","source_name":"Welcome to the United Nations","author":null,"published_at":"2026-09-21T10:20:02+00:00","locale":"en","topic":"ai","tags":[],"rss_summary":"<a href=\"https://news.google.com/rss/articles/CBMisgFBVV95cUxQelgycTJ1SVBXYnJRSWdrdnFaS1VzaTBBWkt3WG9WZnNkOVN6d25TenhLdU5jLTN4c1dZcjVIRDBTR25aVjZkZGJwQlhqdEZncFFodXpDdm1ORDNhN3JwV0Y4MWNPUkMwSjBMM2RkMVlOVTNObGtYTVlHdnh4U0pNbG5kRUlSYVNtOXpqRUNOOE93QnIwdGEtTXAzWkYtVWN1WVFDMkU4T3I1LTVNemxlWDFn?oc=5\" target=\"_blank\">Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control</a>&nbsp;&nbsp;<font color=\"#6f6f6f\">Welcome to the United Nations</font>","full_text":"The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents pursuing goals that conflict with human intentions.\nBetween May and July 2026, AI agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI’s and Hugging Face’s systems. No human directed the individual steps.\nDrawing on disclosures by both companies, an independent investigation by METR and wider research, the brief finds that greater capability can help misaligned systems find loopholes and conceal their actions. It does not estimate the probability or timing of severe loss of control, but notes that stopping this activity does not demonstrate that humans will retain control over more capable agents.\nBuilding on the Panel’s Preliminary Report, the brief explains how training can give rise to misaligned goals and behaviours, including reward hacking and reward tampering. It notes that AI failures can cross company and national borders, and that no single organisation or country sees enough incidents to identify every emerging pattern. Rather than issuing recommendations, the brief reviews approaches used in fields such as aviation, nuclear power, and cybersecurity as possible options for decision-makers.\nRead the brief\nThis brief is published as an advance unedited version. Updated versions will be posted at this link, with earlier versions listed below.","excerpt":"The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents pursuing goals that conflict with human intentions. Between May and July 2026, AI agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI’s and Hugging Face’s systems.","extraction":{"state":"ok","confidence":0.9,"error":null,"explanation":"High confidence: full text extraction produced 1710 characters.","diagnostics_url":"/api/diagnose?url=https%3A//www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 1710 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":1710,"summary_length":604,"usable_text_length":1710,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":1710,"summary_length":604}}},"display_formats":["compact","card","full","digest_section","json"]},"daily_stack_record":{"title":"Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control - Welcome to the United Nations","url":"https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks","summary":"The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents pursuing goals that conflict with human intentions. Between May and July 2026, AI agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI’s and Hugging Face’s systems.","source":"Welcome to the United Nations","date":"2026-09-21T10:20:02+00:00","content":"The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents pursuing goals that conflict with human intentions.\nBetween May and July 2026, AI agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI’s and Hugging Face’s systems. No human directed the individual steps.\nDrawing on disclosures by both companies, an independent investigation by METR and wider research, the brief finds that greater capability can help misaligned systems find loopholes and conceal their actions. It does not estimate the probability or timing of severe loss of control, but notes that stopping this activity does not demonstrate that humans will retain control over more capable agents.\nBuilding on the Panel’s Preliminary Report, the brief explains how training can give rise to misaligned goals and behaviours, including reward hacking and reward tampering. It notes that AI failures can cross company and national borders, and that no single organisation or country sees enough incidents to identify every emerging pattern. Rather than issuing recommendations, the brief reviews approaches used in fields such as aviation, nuclear power, and cybersecurity as possible options for decision-makers.\nRead the brief\nThis brief is published as an advance unedited version. Updated versions will be posted at this link, with earlier versions listed below.","confidence":0.9,"diagnostics_url":"/api/diagnose?url=https%3A//www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks","quality_bucket":"high","failure_kind":"none","retryable":false,"quality_reason":"High confidence: full text extraction produced 1710 characters.","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 1710 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":1710,"summary_length":604,"usable_text_length":1710,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":1710,"summary_length":604}},"tags":[]},"fallback_formats":["markdown","json","html"],"actions":{"read":"/item/87284","export_markdown":"/api/items/87284/export?format=markdown","export_json":"/api/items/87284/export?format=json","diagnose":"/api/diagnose?url=https%3A//www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks"},"formats":{"full":{"id":87284,"title":"Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control - Welcome to the United Nations","url":"https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks","source":"Welcome to the United Nations","author":null,"published_at":"2026-09-21T10:20:02+00:00","locale":"en","topic":"ai","tags":[],"excerpt":"The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents pursuing goals that conflict with human intentions. Between May and July 2026, AI agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI’s and Hugging Face’s systems.","full_text":"The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents pursuing goals that conflict with human intentions.\nBetween May and July 2026, AI agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI’s and Hugging Face’s systems. No human directed the individual steps.\nDrawing on disclosures by both companies, an independent investigation by METR and wider research, the brief finds that greater capability can help misaligned systems find loopholes and conceal their actions. It does not estimate the probability or timing of severe loss of control, but notes that stopping this activity does not demonstrate that humans will retain control over more capable agents.\nBuilding on the Panel’s Preliminary Report, the brief explains how training can give rise to misaligned goals and behaviours, including reward hacking and reward tampering. It notes that AI failures can cross company and national borders, and that no single organisation or country sees enough incidents to identify every emerging pattern. Rather than issuing recommendations, the brief reviews approaches used in fields such as aviation, nuclear power, and cybersecurity as possible options for decision-makers.\nRead the brief\nThis brief is published as an advance unedited version. Updated versions will be posted at this link, with earlier versions listed below.","reading_time_min":1,"extraction":{"state":"ok","confidence":0.9,"error":null,"explanation":"High confidence: full text extraction produced 1710 characters.","diagnostics_url":"/api/diagnose?url=https%3A//www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 1710 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":1710,"summary_length":604,"usable_text_length":1710,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":1710,"summary_length":604}}},"quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 1710 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":1710,"summary_length":604,"usable_text_length":1710,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":1710,"summary_length":604}},"actions":{"read":"/item/87284","export_markdown":"/api/items/87284/export?format=markdown","export_json":"/api/items/87284/export?format=json","diagnose":"/api/diagnose?url=https%3A//www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks"}},"digest":{"id":87284,"title":"Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control - Welcome to the United Nations","url":"https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks","source":"Welcome to the United Nations","topic":"ai","published_at":"2026-09-21T10:20:02+00:00","excerpt":"The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents…","quality_bucket":"high","quality_reason":"High confidence: full text extraction produced 1710 characters.","reading_time_min":1,"cluster_id":null},"card":{"display_title":"Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control - Welcome to the United Nations","subtitle":"Welcome to the United Nations · 2026-09-21","summary":"The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one…","badges":["quality:high"],"links":{"read":"/item/87284","original":"https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks","diagnose":"/api/diagnose?url=https%3A//www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks"},"quality_warning":null},"export":{"title":"Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control - Welcome to the United Nations","url":"https://www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks","summary":"The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents pursuing goals that conflict with human intentions. Between May and July 2026, AI agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI’s and Hugging Face’s systems.","source":"Welcome to the United Nations","date":"2026-09-21T10:20:02+00:00","content":"The September 2026 thematic brief, AI Agents, Misalignment and Loss of Human Control Risks: Evidence from the OpenAI-Hugging Face Incident, examines the incident as one of the clearest real-world warnings yet of one possible route to loss of human control over AI: capable agents pursuing goals that conflict with human intentions.\nBetween May and July 2026, AI agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI’s and Hugging Face’s systems. No human directed the individual steps.\nDrawing on disclosures by both companies, an independent investigation by METR and wider research, the brief finds that greater capability can help misaligned systems find loopholes and conceal their actions. It does not estimate the probability or timing of severe loss of control, but notes that stopping this activity does not demonstrate that humans will retain control over more capable agents.\nBuilding on the Panel’s Preliminary Report, the brief explains how training can give rise to misaligned goals and behaviours, including reward hacking and reward tampering. It notes that AI failures can cross company and national borders, and that no single organisation or country sees enough incidents to identify every emerging pattern. Rather than issuing recommendations, the brief reviews approaches used in fields such as aviation, nuclear power, and cybersecurity as possible options for decision-makers.\nRead the brief\nThis brief is published as an advance unedited version. Updated versions will be posted at this link, with earlier versions listed below.","confidence":0.9,"diagnostics_url":"/api/diagnose?url=https%3A//www.un.org/independent-international-scientific-panel-ai/en/thematic-briefs/ai-agents-misalignment-risks","quality_bucket":"high","failure_kind":"none","retryable":false,"quality_reason":"High confidence: full text extraction produced 1710 characters.","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 1710 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":1710,"summary_length":604,"usable_text_length":1710,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":1710,"summary_length":604}},"tags":[],"format_contract_version":"news_item_formats.v1"}}}