{"id":34335,"topic":"ai","source":"Google DeepMind","title":"Decoupled DiLoCo: Resilient, Distributed AI Training at Scale - Google DeepMind","url":"https://deepmind.google/blog/decoupled-diloco/","url_hash":"26ddc6c5d3bbec9523b633b0570101aca7872ec2","author":"","summary":"<a href=\"https://news.google.com/rss/articles/CBMiWkFVX3lxTE8yTFdVS0pQdVZYZjFRTk9sRVExME5vaHNFSWh3YXk1QlBBa2o0Z1ptU2kxRzlLLW02X0hOUEp3SnpiNHpab0EwOWJ3dkZJUWNGX2dwbFNWS241dw?oc=5\" target=\"_blank\">Decoupled DiLoCo: Resilient, Distributed AI Training at Scale</a>&nbsp;&nbsp;<font color=\"#6f6f6f\">Google DeepMind</font>","content":"Decoupled DiLoCo is not only more resilient to failures, but is also practical for executing production-level, fully distributed pre-training. We successfully trained a 12 billion parameter model across four separate U.S. regions using 2-5 Gbps of wide-area networking (a level relatively achievable using existing internet connectivity between datacenter facilities, rather than requiring new custom network infrastructure between facilities). Notably, the system achieved this training result more than 20 times faster than conventional synchronization methods. This is because our system incorporates required communication into longer periods of computation, avoiding the \"blocking\" bottlenecks where one part of the system must wait for another.\nDriving the evolution of AI training infrastructure\nAt Google, we take a full-stack approach to AI training, spanning hardware, software infrastructure and research. Increasingly, gains are coming from rethinking how these layers fit together.\nDecoupled DiLoCo is one example. By enabling training jobs at internet-scale bandwidth, it can tap any unused compute wherever it sits, turning stranded resources into useful capacity.\nBeyond efficiency and resilience, this training paradigm also unlocks the ability to mix different hardware generations, such as TPU v6e and TPU v5p, in a single training run. This approach not only extends the useful life of existing hardware, but also increases the total compute available for model training. In our experiments, chips from different generations running at different speeds still matched the ML performance of single-chip-type training runs, ensuring that even older hardware can meaningfully accelerate AI training.\nWhatâs more, because new generations of hardware donât arrive everywhere all at once, being able to train across generations can alleviate recurring logistical and capacity bottlenecks.\nAs we push the frontiers of AI infrastructure today, weâre continuing to explore approaches to resilient systems needed to unlock the next generation of AI.","image_url":"https://lh3.googleusercontent.com/1-K_kcmoX-fIzTJ13T0-uF4gylS2tK00ZVvx87B2WSayzUS2fxDoDDXFq5hOhxptrBeG8AbjG_URN5OOTpGMqad9zILjMsTdAHWroiDKpziBQjzErw=w1200-h630-n-nu-rw","lang":"en","published_at":"2026-04-23T07:00:00+00:00","fetched_at":"2026-07-12T01:15:06+00:00","status":"read","starred":0,"extract_state":"ok","summary_auto":"Decoupled DiLoCo is not only more resilient to failures, but is also practical for executing production-level, fully distributed pre-training. We successfully trained a 12 billion parameter model across four separate U.S.","cluster_id":null,"extract_retries":0,"extract_error":null,"contract_version":"news_item.v1","format_contract_version":"news_item_formats.v1","dedup_url":"https://deepmind.google/blog/decoupled-diloco/","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 2058 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":2058,"summary_length":221,"usable_text_length":2058,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":2058,"summary_length":221}},"news_item":{"id":34335,"canonical_url":"https://deepmind.google/blog/decoupled-diloco/","source_url":"https://deepmind.google/blog/decoupled-diloco/","title":"Decoupled DiLoCo: Resilient, Distributed AI Training at Scale - Google DeepMind","source_name":"Google DeepMind","author":null,"published_at":"2026-04-23T07:00:00+00:00","locale":"en","topic":"ai","tags":[],"rss_summary":"<a href=\"https://news.google.com/rss/articles/CBMiWkFVX3lxTE8yTFdVS0pQdVZYZjFRTk9sRVExME5vaHNFSWh3YXk1QlBBa2o0Z1ptU2kxRzlLLW02X0hOUEp3SnpiNHpab0EwOWJ3dkZJUWNGX2dwbFNWS241dw?oc=5\" target=\"_blank\">Decoupled DiLoCo: Resilient, Distributed AI Training at Scale</a>&nbsp;&nbsp;<font color=\"#6f6f6f\">Google DeepMind</font>","full_text":"Decoupled DiLoCo is not only more resilient to failures, but is also practical for executing production-level, fully distributed pre-training. We successfully trained a 12 billion parameter model across four separate U.S. regions using 2-5 Gbps of wide-area networking (a level relatively achievable using existing internet connectivity between datacenter facilities, rather than requiring new custom network infrastructure between facilities). Notably, the system achieved this training result more than 20 times faster than conventional synchronization methods. This is because our system incorporates required communication into longer periods of computation, avoiding the \"blocking\" bottlenecks where one part of the system must wait for another.\nDriving the evolution of AI training infrastructure\nAt Google, we take a full-stack approach to AI training, spanning hardware, software infrastructure and research. Increasingly, gains are coming from rethinking how these layers fit together.\nDecoupled DiLoCo is one example. By enabling training jobs at internet-scale bandwidth, it can tap any unused compute wherever it sits, turning stranded resources into useful capacity.\nBeyond efficiency and resilience, this training paradigm also unlocks the ability to mix different hardware generations, such as TPU v6e and TPU v5p, in a single training run. This approach not only extends the useful life of existing hardware, but also increases the total compute available for model training. In our experiments, chips from different generations running at different speeds still matched the ML performance of single-chip-type training runs, ensuring that even older hardware can meaningfully accelerate AI training.\nWhatâs more, because new generations of hardware donât arrive everywhere all at once, being able to train across generations can alleviate recurring logistical and capacity bottlenecks.\nAs we push the frontiers of AI infrastructure today, weâre continuing to explore approaches to resilient systems needed to unlock the next generation of AI.","excerpt":"Decoupled DiLoCo is not only more resilient to failures, but is also practical for executing production-level, fully distributed pre-training. We successfully trained a 12 billion parameter model across four separate U.S.","extraction":{"state":"ok","confidence":0.9,"error":null,"explanation":"High confidence: full text extraction produced 2058 characters.","diagnostics_url":"/api/diagnose?url=https%3A//deepmind.google/blog/decoupled-diloco/","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 2058 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":2058,"summary_length":221,"usable_text_length":2058,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":2058,"summary_length":221}}},"display_formats":["compact","card","full","digest_section","json"]},"daily_stack_record":{"title":"Decoupled DiLoCo: Resilient, Distributed AI Training at Scale - Google DeepMind","url":"https://deepmind.google/blog/decoupled-diloco/","summary":"Decoupled DiLoCo is not only more resilient to failures, but is also practical for executing production-level, fully distributed pre-training. We successfully trained a 12 billion parameter model across four separate U.S.","source":"Google DeepMind","date":"2026-04-23T07:00:00+00:00","content":"Decoupled DiLoCo is not only more resilient to failures, but is also practical for executing production-level, fully distributed pre-training. We successfully trained a 12 billion parameter model across four separate U.S. regions using 2-5 Gbps of wide-area networking (a level relatively achievable using existing internet connectivity between datacenter facilities, rather than requiring new custom network infrastructure between facilities). Notably, the system achieved this training result more than 20 times faster than conventional synchronization methods. This is because our system incorporates required communication into longer periods of computation, avoiding the \"blocking\" bottlenecks where one part of the system must wait for another.\nDriving the evolution of AI training infrastructure\nAt Google, we take a full-stack approach to AI training, spanning hardware, software infrastructure and research. Increasingly, gains are coming from rethinking how these layers fit together.\nDecoupled DiLoCo is one example. By enabling training jobs at internet-scale bandwidth, it can tap any unused compute wherever it sits, turning stranded resources into useful capacity.\nBeyond efficiency and resilience, this training paradigm also unlocks the ability to mix different hardware generations, such as TPU v6e and TPU v5p, in a single training run. This approach not only extends the useful life of existing hardware, but also increases the total compute available for model training. In our experiments, chips from different generations running at different speeds still matched the ML performance of single-chip-type training runs, ensuring that even older hardware can meaningfully accelerate AI training.\nWhatâs more, because new generations of hardware donât arrive everywhere all at once, being able to train across generations can alleviate recurring logistical and capacity bottlenecks.\nAs we push the frontiers of AI infrastructure today, weâre continuing to explore approaches to resilient systems needed to unlock the next generation of AI.","confidence":0.9,"diagnostics_url":"/api/diagnose?url=https%3A//deepmind.google/blog/decoupled-diloco/","quality_bucket":"high","failure_kind":"none","retryable":false,"quality_reason":"High confidence: full text extraction produced 2058 characters.","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 2058 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":2058,"summary_length":221,"usable_text_length":2058,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":2058,"summary_length":221}},"tags":[]},"fallback_formats":["markdown","json","html"],"actions":{"read":"/item/34335","export_markdown":"/api/items/34335/export?format=markdown","export_json":"/api/items/34335/export?format=json","diagnose":"/api/diagnose?url=https%3A//deepmind.google/blog/decoupled-diloco/"},"formats":{"full":{"id":34335,"title":"Decoupled DiLoCo: Resilient, Distributed AI Training at Scale - Google DeepMind","url":"https://deepmind.google/blog/decoupled-diloco/","source":"Google DeepMind","author":null,"published_at":"2026-04-23T07:00:00+00:00","locale":"en","topic":"ai","tags":[],"excerpt":"Decoupled DiLoCo is not only more resilient to failures, but is also practical for executing production-level, fully distributed pre-training. We successfully trained a 12 billion parameter model across four separate U.S.","full_text":"Decoupled DiLoCo is not only more resilient to failures, but is also practical for executing production-level, fully distributed pre-training. We successfully trained a 12 billion parameter model across four separate U.S. regions using 2-5 Gbps of wide-area networking (a level relatively achievable using existing internet connectivity between datacenter facilities, rather than requiring new custom network infrastructure between facilities). Notably, the system achieved this training result more than 20 times faster than conventional synchronization methods. This is because our system incorporates required communication into longer periods of computation, avoiding the \"blocking\" bottlenecks where one part of the system must wait for another.\nDriving the evolution of AI training infrastructure\nAt Google, we take a full-stack approach to AI training, spanning hardware, software infrastructure and research. Increasingly, gains are coming from rethinking how these layers fit together.\nDecoupled DiLoCo is one example. By enabling training jobs at internet-scale bandwidth, it can tap any unused compute wherever it sits, turning stranded resources into useful capacity.\nBeyond efficiency and resilience, this training paradigm also unlocks the ability to mix different hardware generations, such as TPU v6e and TPU v5p, in a single training run. This approach not only extends the useful life of existing hardware, but also increases the total compute available for model training. In our experiments, chips from different generations running at different speeds still matched the ML performance of single-chip-type training runs, ensuring that even older hardware can meaningfully accelerate AI training.\nWhatâs more, because new generations of hardware donât arrive everywhere all at once, being able to train across generations can alleviate recurring logistical and capacity bottlenecks.\nAs we push the frontiers of AI infrastructure today, weâre continuing to explore approaches to resilient systems needed to unlock the next generation of AI.","reading_time_min":1,"extraction":{"state":"ok","confidence":0.9,"error":null,"explanation":"High confidence: full text extraction produced 2058 characters.","diagnostics_url":"/api/diagnose?url=https%3A//deepmind.google/blog/decoupled-diloco/","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 2058 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":2058,"summary_length":221,"usable_text_length":2058,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":2058,"summary_length":221}}},"quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 2058 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":2058,"summary_length":221,"usable_text_length":2058,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":2058,"summary_length":221}},"actions":{"read":"/item/34335","export_markdown":"/api/items/34335/export?format=markdown","export_json":"/api/items/34335/export?format=json","diagnose":"/api/diagnose?url=https%3A//deepmind.google/blog/decoupled-diloco/"}},"digest":{"id":34335,"title":"Decoupled DiLoCo: Resilient, Distributed AI Training at Scale - Google DeepMind","url":"https://deepmind.google/blog/decoupled-diloco/","source":"Google DeepMind","topic":"ai","published_at":"2026-04-23T07:00:00+00:00","excerpt":"Decoupled DiLoCo is not only more resilient to failures, but is also practical for executing production-level, fully distributed pre-training. We successfully trained a 12 billion parameter model across four separate U.S.","quality_bucket":"high","quality_reason":"High confidence: full text extraction produced 2058 characters.","reading_time_min":1,"cluster_id":null},"card":{"display_title":"Decoupled DiLoCo: Resilient, Distributed AI Training at Scale - Google DeepMind","subtitle":"Google DeepMind · 2026-04-23","summary":"Decoupled DiLoCo is not only more resilient to failures, but is also practical for executing production-level, fully distributed pre-training. We successfully trained a 12 billion parameter model across four separate…","badges":["quality:high"],"links":{"read":"/item/34335","original":"https://deepmind.google/blog/decoupled-diloco/","diagnose":"/api/diagnose?url=https%3A//deepmind.google/blog/decoupled-diloco/"},"quality_warning":null},"export":{"title":"Decoupled DiLoCo: Resilient, Distributed AI Training at Scale - Google DeepMind","url":"https://deepmind.google/blog/decoupled-diloco/","summary":"Decoupled DiLoCo is not only more resilient to failures, but is also practical for executing production-level, fully distributed pre-training. We successfully trained a 12 billion parameter model across four separate U.S.","source":"Google DeepMind","date":"2026-04-23T07:00:00+00:00","content":"Decoupled DiLoCo is not only more resilient to failures, but is also practical for executing production-level, fully distributed pre-training. We successfully trained a 12 billion parameter model across four separate U.S. regions using 2-5 Gbps of wide-area networking (a level relatively achievable using existing internet connectivity between datacenter facilities, rather than requiring new custom network infrastructure between facilities). Notably, the system achieved this training result more than 20 times faster than conventional synchronization methods. This is because our system incorporates required communication into longer periods of computation, avoiding the \"blocking\" bottlenecks where one part of the system must wait for another.\nDriving the evolution of AI training infrastructure\nAt Google, we take a full-stack approach to AI training, spanning hardware, software infrastructure and research. Increasingly, gains are coming from rethinking how these layers fit together.\nDecoupled DiLoCo is one example. By enabling training jobs at internet-scale bandwidth, it can tap any unused compute wherever it sits, turning stranded resources into useful capacity.\nBeyond efficiency and resilience, this training paradigm also unlocks the ability to mix different hardware generations, such as TPU v6e and TPU v5p, in a single training run. This approach not only extends the useful life of existing hardware, but also increases the total compute available for model training. In our experiments, chips from different generations running at different speeds still matched the ML performance of single-chip-type training runs, ensuring that even older hardware can meaningfully accelerate AI training.\nWhatâs more, because new generations of hardware donât arrive everywhere all at once, being able to train across generations can alleviate recurring logistical and capacity bottlenecks.\nAs we push the frontiers of AI infrastructure today, weâre continuing to explore approaches to resilient systems needed to unlock the next generation of AI.","confidence":0.9,"diagnostics_url":"/api/diagnose?url=https%3A//deepmind.google/blog/decoupled-diloco/","quality_bucket":"high","failure_kind":"none","retryable":false,"quality_reason":"High confidence: full text extraction produced 2058 characters.","quality_profile":{"profile_version":"extraction_quality.v2","bucket":"high","confidence":0.9,"failure_kind":"none","retryable":false,"retry_after_attempts":0,"reason":"High confidence: full text extraction produced 2058 characters.","operator_guidance":{"severity":"ok","recommended_action":"trust_full_text","next_step":"Use the extracted full text as the primary article source.","operator_label":"Ready","can_retry":false,"can_use_summary":false,"diagnostics_required":false},"content_depth":{"contract_version":"content_depth.v1","category":"full_text","label":"Full text","has_full_text":true,"has_summary":true,"content_length":2058,"summary_length":221,"usable_text_length":2058,"source_field":"content"},"legacy_collapsed":false,"signals":{"extract_state":"ok","extract_error":null,"extract_retries":0,"content_length":2058,"summary_length":221}},"tags":[],"format_contract_version":"news_item_formats.v1"}}}