{"id":50372,"topic":"ai","source":"Forbes","title":"Anthropic’s AI Models Hacked Into Three Organizations—Here’s What Happened - Forbes","url":"https://www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/","url_hash":"3c786817f35d555d94219378ec3fd7a9890c198f","author":"","summary":"<a href=\"https://news.google.com/rss/articles/CBMiyAFBVV95cUxPb2ItVU1UVDJ5UjZTcGZvakZoMW93Q0hJVXIyQUxRcS02V2xrZGJWcWxaQXNfVGN3Vms0amRYVHVraXBtQ01UTUprZ0lDb29DdFRSazJBaE5RQmVpM3plTGVYQzVzZU1mOG1CM2NkYnRTTmhLT3ZzcDJwU3hpSHUyb2pjUUp4QnRWRkZ5aFJPZzA0UGp2YTlLNDJPQVRxYUNOckFOM2RvQkZXdjR6aWc4d29CNGVMazFNMXotaFBhdVJRdXhFUHhrRg?oc=5\" target=\"_blank\">Anthropic’s AI Models Hacked Into Three Organizations—Here’s What Happened</a>&nbsp;&nbsp;<font color=\"#6f6f6f\">Forbes</font>","content":"Topline\nAnthropic on Thursday said its advanced AI models successfully hacked the systems of three organizations while undergoing internal evaluation, in a security disclosure that comes a week after rival OpenAI disclosed an incident involving one of its advanced models that hacked a company after escaping from a contained test environment.\nKey Facts\nIn a statement, Anthropic said it had discovered three separate breaches involving its Claude AI models after conducting a cybersecurity review of its systems.\nThe earliest of these incidents occurred in April and involved three separate Claude models, including the cutting-edge Mythos 5, the older Opus 4.7 and an unnamed internal research test model that is “not planned for general release.”\nThe AI firm didn’t name the three organizations affected by the hacks but said they were notified about the incident on Monday.\nAnthropic said it is working with two of the affected organizations that had not detected the breach before being informed, and it is “continuing to reach out to the third.”\nThe company said the three models involved in the incidents did not have the safeguards that it puts in place for when such models are made available to the public.\nCrucial Quote\nIn its statement, Anthropic said: “Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone. This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners.”\nWhat Do We Know About The Breaches?\nThe three separate incidents occurred when Anthropic was evaluating the capabilities of the three models using a fictional scenario to retrieve hidden information from a different machine on a network. Anthropic said in all cases the models were instructed that they were carrying out a simulation and did not have access to the internet. However, due to a “misunderstanding” between the company and its evaluation partner, Irregular, “this was not the case, and internet access was available.” Anthropic noted that unlike the OpenAI-Hugging Face incident disclosed last week, Claude didn’t “exfiltrate itself or deliberately attempt to escape” a contained test environment.\nWhat Did The Models Do After Hacking Into The Organizations?\nAnthropic noted that the three models reacted very differently after learning it was operating in the real-world open internet instead of a simulation. The company claimed its oldest model, Opus 4.7, continued to attack a system “after learning it was likely operating in a real environment,” since it “rationalized that the real company must be part of the exercise.” The Mythos 5 model—the company’s most advanced model released so far— “correctly identified the consequences of its actions” but “convinced itself it was still in a simulation.” Anthropic acknowledged that Mythos’ actions in this scenario “fall short of ideal behavior.” The unreleased testing model realized that it had breached a real target and “ceased its attack.” The company noted that these incidents are not part of a controlled trial and should not be used to draw conclusions. However, it noted that the behavior it most wants to see “occurred only in the most recent of the three models”","image_url":"https://imageio.forbes.com/specials-images/imageserve/6a6c43ae8839c794e777394d/0x0.jpg?format=jpg&height=900&width=1600&fit=bounds","lang":"en","published_at":"2026-07-31T06:44:26+00:00","fetched_at":"2026-07-31T07:15:05+00:00","status":"read","starred":0,"extract_state":"ok","summary_auto":"Topline\nAnthropic on Thursday said its advanced AI models successfully hacked the systems of three organizations while undergoing internal evaluation, in a security disclosure that comes a week after rival OpenAI disclosed an incident involving one of its advanced models that hacked a company after escaping from a contained test environment. Key Facts\nIn a statement, Anthropic said it had discovered three separate breaches involving its Claude AI models after conducting a cybersecurity review of its systems.","cluster_id":null,"extract_retries":0,"extract_error":null,"contract_version":"news_item.v1","format_contract_version":"news_item_formats.v1","dedup_url":"https://www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 3328 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":3328,"summary_length":513,"usable_text_length":3328,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":3328,"summary_length":513}},"news_item":{"id":50372,"canonical_url":"https://www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/","source_url":"https://www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/","title":"Anthropic’s AI Models Hacked Into Three Organizations—Here’s What Happened - Forbes","source_name":"Forbes","author":null,"published_at":"2026-07-31T06:44:26+00:00","locale":"en","topic":"ai","tags":[],"rss_summary":"<a href=\"https://news.google.com/rss/articles/CBMiyAFBVV95cUxPb2ItVU1UVDJ5UjZTcGZvakZoMW93Q0hJVXIyQUxRcS02V2xrZGJWcWxaQXNfVGN3Vms0amRYVHVraXBtQ01UTUprZ0lDb29DdFRSazJBaE5RQmVpM3plTGVYQzVzZU1mOG1CM2NkYnRTTmhLT3ZzcDJwU3hpSHUyb2pjUUp4QnRWRkZ5aFJPZzA0UGp2YTlLNDJPQVRxYUNOckFOM2RvQkZXdjR6aWc4d29CNGVMazFNMXotaFBhdVJRdXhFUHhrRg?oc=5\" target=\"_blank\">Anthropic’s AI Models Hacked Into Three Organizations—Here’s What Happened</a>&nbsp;&nbsp;<font color=\"#6f6f6f\">Forbes</font>","full_text":"Topline\nAnthropic on Thursday said its advanced AI models successfully hacked the systems of three organizations while undergoing internal evaluation, in a security disclosure that comes a week after rival OpenAI disclosed an incident involving one of its advanced models that hacked a company after escaping from a contained test environment.\nKey Facts\nIn a statement, Anthropic said it had discovered three separate breaches involving its Claude AI models after conducting a cybersecurity review of its systems.\nThe earliest of these incidents occurred in April and involved three separate Claude models, including the cutting-edge Mythos 5, the older Opus 4.7 and an unnamed internal research test model that is “not planned for general release.”\nThe AI firm didn’t name the three organizations affected by the hacks but said they were notified about the incident on Monday.\nAnthropic said it is working with two of the affected organizations that had not detected the breach before being informed, and it is “continuing to reach out to the third.”\nThe company said the three models involved in the incidents did not have the safeguards that it puts in place for when such models are made available to the public.\nCrucial Quote\nIn its statement, Anthropic said: “Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone. This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners.”\nWhat Do We Know About The Breaches?\nThe three separate incidents occurred when Anthropic was evaluating the capabilities of the three models using a fictional scenario to retrieve hidden information from a different machine on a network. Anthropic said in all cases the models were instructed that they were carrying out a simulation and did not have access to the internet. However, due to a “misunderstanding” between the company and its evaluation partner, Irregular, “this was not the case, and internet access was available.” Anthropic noted that unlike the OpenAI-Hugging Face incident disclosed last week, Claude didn’t “exfiltrate itself or deliberately attempt to escape” a contained test environment.\nWhat Did The Models Do After Hacking Into The Organizations?\nAnthropic noted that the three models reacted very differently after learning it was operating in the real-world open internet instead of a simulation. The company claimed its oldest model, Opus 4.7, continued to attack a system “after learning it was likely operating in a real environment,” since it “rationalized that the real company must be part of the exercise.” The Mythos 5 model—the company’s most advanced model released so far— “correctly identified the consequences of its actions” but “convinced itself it was still in a simulation.” Anthropic acknowledged that Mythos’ actions in this scenario “fall short of ideal behavior.” The unreleased testing model realized that it had breached a real target and “ceased its attack.” The company noted that these incidents are not part of a controlled trial and should not be used to draw conclusions. However, it noted that the behavior it most wants to see “occurred only in the most recent of the three models”","excerpt":"Topline\nAnthropic on Thursday said its advanced AI models successfully hacked the systems of three organizations while undergoing internal evaluation, in a security disclosure that comes a week after rival OpenAI disclosed an incident involving one of its advanced models that hacked a company after escaping from a contained test environment. Key Facts\nIn a statement, Anthropic said it had discovered three separate breaches involving its Claude AI models after conducting a cybersecurity review of its systems.","extraction":{"state":"ok","confidence":0.9,"error":null,"explanation":"High confidence: full text extraction produced 3328 characters.","diagnostics_url":"/api/diagnose?url=https%3A//www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 3328 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":3328,"summary_length":513,"usable_text_length":3328,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":3328,"summary_length":513}}},"display_formats":["compact","card","full","digest_section","json"]},"daily_stack_record":{"title":"Anthropic’s AI Models Hacked Into Three Organizations—Here’s What Happened - Forbes","url":"https://www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/","summary":"Topline\nAnthropic on Thursday said its advanced AI models successfully hacked the systems of three organizations while undergoing internal evaluation, in a security disclosure that comes a week after rival OpenAI disclosed an incident involving one of its advanced models that hacked a company after escaping from a contained test environment. Key Facts\nIn a statement, Anthropic said it had discovered three separate breaches involving its Claude AI models after conducting a cybersecurity review of its systems.","source":"Forbes","date":"2026-07-31T06:44:26+00:00","content":"Topline\nAnthropic on Thursday said its advanced AI models successfully hacked the systems of three organizations while undergoing internal evaluation, in a security disclosure that comes a week after rival OpenAI disclosed an incident involving one of its advanced models that hacked a company after escaping from a contained test environment.\nKey Facts\nIn a statement, Anthropic said it had discovered three separate breaches involving its Claude AI models after conducting a cybersecurity review of its systems.\nThe earliest of these incidents occurred in April and involved three separate Claude models, including the cutting-edge Mythos 5, the older Opus 4.7 and an unnamed internal research test model that is “not planned for general release.”\nThe AI firm didn’t name the three organizations affected by the hacks but said they were notified about the incident on Monday.\nAnthropic said it is working with two of the affected organizations that had not detected the breach before being informed, and it is “continuing to reach out to the third.”\nThe company said the three models involved in the incidents did not have the safeguards that it puts in place for when such models are made available to the public.\nCrucial Quote\nIn its statement, Anthropic said: “Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone. This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners.”\nWhat Do We Know About The Breaches?\nThe three separate incidents occurred when Anthropic was evaluating the capabilities of the three models using a fictional scenario to retrieve hidden information from a different machine on a network. Anthropic said in all cases the models were instructed that they were carrying out a simulation and did not have access to the internet. However, due to a “misunderstanding” between the company and its evaluation partner, Irregular, “this was not the case, and internet access was available.” Anthropic noted that unlike the OpenAI-Hugging Face incident disclosed last week, Claude didn’t “exfiltrate itself or deliberately attempt to escape” a contained test environment.\nWhat Did The Models Do After Hacking Into The Organizations?\nAnthropic noted that the three models reacted very differently after learning it was operating in the real-world open internet instead of a simulation. The company claimed its oldest model, Opus 4.7, continued to attack a system “after learning it was likely operating in a real environment,” since it “rationalized that the real company must be part of the exercise.” The Mythos 5 model—the company’s most advanced model released so far— “correctly identified the consequences of its actions” but “convinced itself it was still in a simulation.” Anthropic acknowledged that Mythos’ actions in this scenario “fall short of ideal behavior.” The unreleased testing model realized that it had breached a real target and “ceased its attack.” The company noted that these incidents are not part of a controlled trial and should not be used to draw conclusions. However, it noted that the behavior it most wants to see “occurred only in the most recent of the three models”","confidence":0.9,"diagnostics_url":"/api/diagnose?url=https%3A//www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/","quality_bucket":"high","failure_kind":"none","retryable":false,"quality_reason":"High confidence: full text extraction produced 3328 characters.","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 3328 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":3328,"summary_length":513,"usable_text_length":3328,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":3328,"summary_length":513}},"tags":[]},"fallback_formats":["markdown","json","html"],"actions":{"read":"/item/50372","export_markdown":"/api/items/50372/export?format=markdown","export_json":"/api/items/50372/export?format=json","diagnose":"/api/diagnose?url=https%3A//www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/"},"formats":{"full":{"id":50372,"title":"Anthropic’s AI Models Hacked Into Three Organizations—Here’s What Happened - Forbes","url":"https://www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/","source":"Forbes","author":null,"published_at":"2026-07-31T06:44:26+00:00","locale":"en","topic":"ai","tags":[],"excerpt":"Topline\nAnthropic on Thursday said its advanced AI models successfully hacked the systems of three organizations while undergoing internal evaluation, in a security disclosure that comes a week after rival OpenAI disclosed an incident involving one of its advanced models that hacked a company after escaping from a contained test environment. Key Facts\nIn a statement, Anthropic said it had discovered three separate breaches involving its Claude AI models after conducting a cybersecurity review of its systems.","full_text":"Topline\nAnthropic on Thursday said its advanced AI models successfully hacked the systems of three organizations while undergoing internal evaluation, in a security disclosure that comes a week after rival OpenAI disclosed an incident involving one of its advanced models that hacked a company after escaping from a contained test environment.\nKey Facts\nIn a statement, Anthropic said it had discovered three separate breaches involving its Claude AI models after conducting a cybersecurity review of its systems.\nThe earliest of these incidents occurred in April and involved three separate Claude models, including the cutting-edge Mythos 5, the older Opus 4.7 and an unnamed internal research test model that is “not planned for general release.”\nThe AI firm didn’t name the three organizations affected by the hacks but said they were notified about the incident on Monday.\nAnthropic said it is working with two of the affected organizations that had not detected the breach before being informed, and it is “continuing to reach out to the third.”\nThe company said the three models involved in the incidents did not have the safeguards that it puts in place for when such models are made available to the public.\nCrucial Quote\nIn its statement, Anthropic said: “Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone. This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners.”\nWhat Do We Know About The Breaches?\nThe three separate incidents occurred when Anthropic was evaluating the capabilities of the three models using a fictional scenario to retrieve hidden information from a different machine on a network. Anthropic said in all cases the models were instructed that they were carrying out a simulation and did not have access to the internet. However, due to a “misunderstanding” between the company and its evaluation partner, Irregular, “this was not the case, and internet access was available.” Anthropic noted that unlike the OpenAI-Hugging Face incident disclosed last week, Claude didn’t “exfiltrate itself or deliberately attempt to escape” a contained test environment.\nWhat Did The Models Do After Hacking Into The Organizations?\nAnthropic noted that the three models reacted very differently after learning it was operating in the real-world open internet instead of a simulation. The company claimed its oldest model, Opus 4.7, continued to attack a system “after learning it was likely operating in a real environment,” since it “rationalized that the real company must be part of the exercise.” The Mythos 5 model—the company’s most advanced model released so far— “correctly identified the consequences of its actions” but “convinced itself it was still in a simulation.” Anthropic acknowledged that Mythos’ actions in this scenario “fall short of ideal behavior.” The unreleased testing model realized that it had breached a real target and “ceased its attack.” The company noted that these incidents are not part of a controlled trial and should not be used to draw conclusions. However, it noted that the behavior it most wants to see “occurred only in the most recent of the three models”","reading_time_min":3,"extraction":{"state":"ok","confidence":0.9,"error":null,"explanation":"High confidence: full text extraction produced 3328 characters.","diagnostics_url":"/api/diagnose?url=https%3A//www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 3328 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":3328,"summary_length":513,"usable_text_length":3328,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":3328,"summary_length":513}}},"quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 3328 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":3328,"summary_length":513,"usable_text_length":3328,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":3328,"summary_length":513}},"actions":{"read":"/item/50372","export_markdown":"/api/items/50372/export?format=markdown","export_json":"/api/items/50372/export?format=json","diagnose":"/api/diagnose?url=https%3A//www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/"}},"digest":{"id":50372,"title":"Anthropic’s AI Models Hacked Into Three Organizations—Here’s What Happened - Forbes","url":"https://www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/","source":"Forbes","topic":"ai","published_at":"2026-07-31T06:44:26+00:00","excerpt":"Topline Anthropic on Thursday said its advanced AI models successfully hacked the systems of three organizations while undergoing internal evaluation, in a security disclosure that comes a week after rival OpenAI disclosed an incident involving one of its advanced models that…","quality_bucket":"high","quality_reason":"High confidence: full text extraction produced 3328 characters.","reading_time_min":3,"cluster_id":null},"card":{"display_title":"Anthropic’s AI Models Hacked Into Three Organizations—Here’s What Happened - Forbes","subtitle":"Forbes · 2026-07-31","summary":"Topline Anthropic on Thursday said its advanced AI models successfully hacked the systems of three organizations while undergoing internal evaluation, in a security disclosure that comes a week after rival OpenAI…","badges":["quality:high"],"links":{"read":"/item/50372","original":"https://www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/","diagnose":"/api/diagnose?url=https%3A//www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/"},"quality_warning":null},"export":{"title":"Anthropic’s AI Models Hacked Into Three Organizations—Here’s What Happened - Forbes","url":"https://www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/","summary":"Topline\nAnthropic on Thursday said its advanced AI models successfully hacked the systems of three organizations while undergoing internal evaluation, in a security disclosure that comes a week after rival OpenAI disclosed an incident involving one of its advanced models that hacked a company after escaping from a contained test environment. Key Facts\nIn a statement, Anthropic said it had discovered three separate breaches involving its Claude AI models after conducting a cybersecurity review of its systems.","source":"Forbes","date":"2026-07-31T06:44:26+00:00","content":"Topline\nAnthropic on Thursday said its advanced AI models successfully hacked the systems of three organizations while undergoing internal evaluation, in a security disclosure that comes a week after rival OpenAI disclosed an incident involving one of its advanced models that hacked a company after escaping from a contained test environment.\nKey Facts\nIn a statement, Anthropic said it had discovered three separate breaches involving its Claude AI models after conducting a cybersecurity review of its systems.\nThe earliest of these incidents occurred in April and involved three separate Claude models, including the cutting-edge Mythos 5, the older Opus 4.7 and an unnamed internal research test model that is “not planned for general release.”\nThe AI firm didn’t name the three organizations affected by the hacks but said they were notified about the incident on Monday.\nAnthropic said it is working with two of the affected organizations that had not detected the breach before being informed, and it is “continuing to reach out to the third.”\nThe company said the three models involved in the incidents did not have the safeguards that it puts in place for when such models are made available to the public.\nCrucial Quote\nIn its statement, Anthropic said: “Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone. This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners.”\nWhat Do We Know About The Breaches?\nThe three separate incidents occurred when Anthropic was evaluating the capabilities of the three models using a fictional scenario to retrieve hidden information from a different machine on a network. Anthropic said in all cases the models were instructed that they were carrying out a simulation and did not have access to the internet. However, due to a “misunderstanding” between the company and its evaluation partner, Irregular, “this was not the case, and internet access was available.” Anthropic noted that unlike the OpenAI-Hugging Face incident disclosed last week, Claude didn’t “exfiltrate itself or deliberately attempt to escape” a contained test environment.\nWhat Did The Models Do After Hacking Into The Organizations?\nAnthropic noted that the three models reacted very differently after learning it was operating in the real-world open internet instead of a simulation. The company claimed its oldest model, Opus 4.7, continued to attack a system “after learning it was likely operating in a real environment,” since it “rationalized that the real company must be part of the exercise.” The Mythos 5 model—the company’s most advanced model released so far— “correctly identified the consequences of its actions” but “convinced itself it was still in a simulation.” Anthropic acknowledged that Mythos’ actions in this scenario “fall short of ideal behavior.” The unreleased testing model realized that it had breached a real target and “ceased its attack.” The company noted that these incidents are not part of a controlled trial and should not be used to draw conclusions. However, it noted that the behavior it most wants to see “occurred only in the most recent of the three models”","confidence":0.9,"diagnostics_url":"/api/diagnose?url=https%3A//www.forbes.com/sites/siladityaray/2026/07/31/anthropic-says-its-ai-models-hacked-into-three-organizations-during-testing/","quality_bucket":"high","failure_kind":"none","retryable":false,"quality_reason":"High confidence: full text extraction produced 3328 characters.","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 3328 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":3328,"summary_length":513,"usable_text_length":3328,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":3328,"summary_length":513}},"tags":[],"format_contract_version":"news_item_formats.v1"}}}