{"id":90053,"date":"2026-06-30T00:00:09","date_gmt":"2026-06-30T00:00:09","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/90053\/"},"modified":"2026-06-30T00:00:09","modified_gmt":"2026-06-30T00:00:09","slug":"building-a-fault-tolerant-training-system-with-pytorch-monarch-on-oke","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/90053\/","title":{"rendered":"Building a fault-tolerant training system with PyTorch Monarch on OKE"},"content":{"rendered":"<p>Training large models is hard. Training them reliably on a shared GPU cluster \u2014 where a bad NIC, a flaky node, or a noisy neighbor can erase hours of progress \u2014 is harder. This post walks through a training system on Oracle Kubernetes Engine (OKE) that:<\/p>\n<p>orchestrates multi-node PyTorch jobs from a single controller,<\/p>\n<p>is aware of the cluster\u2019s RDMA topology so collective traffic stays on the fast lane,<\/p>\n<p>and survives mid-run failures without restarting from scratch.<\/p>\n<p>We get there step by step \u2014 each a small, testable change, each with real numbers from our cluster.<\/p>\n<p>Architecture Diagram<\/p>\n<p>Here\u2019s where we\u2019re headed \u2014 the full system end-to-end (the final configuration 3.5). We\u2019ll assemble it one piece at a time in the sections below.<\/p>\n<p><img fetchpriority=\"high\" decoding=\"async\" width=\"2560\" height=\"1980\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/06\/architecture_diagram-1-scaled.png\" alt=\"Architecture Diagram of Fault-Tolerant Training System with PyTorch Monarch on OKE\" class=\"wp-image-4248\"  \/>Architecture Diagram of Fault-Tolerant Training System with PyTorch Monarch on OKE<\/p>\n<p>How the layers stack: Kueue admits the job atomically and places pods to minimize RDMA switch hops. The Monarch Operator provisions the worker pods, and the Monarch Controller \u2014 one Python script \u2014 orchestrates them. Two TorchTitan trainer pods (one per A100 node, 8 GPUs each) run as Monarch actors, with the TorchFT Lighthouse coordinating per-step fault tolerance. The RDMA (RoCE) fabric (bold blue) carries both FSDP and inter-replica gradient allreduce at line rate.<\/p>\n<p>1. The building blocks<\/p>\n<p>A quick tour of the pieces we\u2019re combining.<\/p>\n<p><a href=\"https:\/\/docs.oracle.com\/en-us\/iaas\/Content\/ContEng\/home.htm\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">OKE (Oracle Kubernetes Engine)<\/a> \u2014 Oracle Cloud\u2019s managed Kubernetes service: the foundation for pods, jobs, the operator that schedules Monarch workers, and the RDMA configuration. Our cluster:<\/p>\n<p>2 \u00d7 A100 bare-metal nodes, 16 GPUs total, with RDMA-capable inter-node networking.<\/p>\n<p>Two nodes is the smallest configuration where multi-node networking, gang scheduling, and topology awareness actually matter.<\/p>\n<p><a href=\"https:\/\/github.com\/meta-pytorch\/monarch\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">PyTorch Monarch<\/a> \u2014 Open Source distributed orchestration framework. Instead of a torchrun wrapper, a Helm chart, and an operator config, you write one Python controller script describing the job \u2014 hosts, GPUs per host, and what runs on each rank \u2014 and Monarch handles the rest. It runs on Kubernetes (via the <a href=\"https:\/\/github.com\/meta-pytorch\/monarch-kubernetes\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Monarch Operator<\/a> and ports to Slurm by swapping just the pod-spec helper. And it\u2019s a single point of orchestration: failures, resizing, and restarts all flow through the same controller \u2014 exactly what we need before layering fault tolerance on top.<\/p>\n<p><a href=\"https:\/\/github.com\/pytorch\/torchtitan\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">TorchTitan<\/a> \u2014 PyTorch\u2019s reference implementation for large-model training. It bundles FSDP, tensor\/pipeline parallelism, and activation checkpointing behind a clean config, letting us focus on the system instead of the model-parallelism strategy.<\/p>\n<p><a href=\"https:\/\/github.com\/pytorch\/torchft\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">TorchFT<\/a> \u2014 PyTorch\u2019s fault-tolerance library. The headline idea is per-step fault tolerance: when a replica dies, survivors keep training and the dead replica rejoins later. It works via a coordination server (the Lighthouse) plus runtime wrappers around DDP and the optimizer. Monarch orchestrates; TorchFT handles \u201cwhat happens when something dies mid-step.\u201d<\/p>\n<p><a href=\"https:\/\/kueue.sigs.k8s.io\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Kueue<\/a> \u2014 a Kubernetes-native job queue. We use it for gang scheduling (all pods of a job start together or not at all) and RDMA topology-aware scheduling (placing pods so their NICs share the fewest switch hops).<\/p>\n<p>2. The evaluation model<\/p>\n<p>To compare each step apples-to-apples, we train the same model every time:<\/p>\n<p>Llama 3 \u2014 8B parameters, trained on the C4 dataset.<\/p>\n<p>This isn\u2019t about a state-of-the-art checkpoint \u2014 it\u2019s about validating the *system*. Llama 3 8B is large enough to be realistic, small enough to iterate quickly, and well-understood enough that the metrics tell a clear story.<\/p>\n<p>3. Construction<\/p>\n<p>We build the system incrementally \u2014 start simple, measure, then add one component at a time. Each configuration links to its own branch with a detailed write-up.<\/p>\n<p>3.1 <a href=\"https:\/\/github.com\/oci-ai-incubations\/monarch-recipe-bp\/blob\/cfg1_oke_torchtitan_torchrun\/examples\/k8s_titan_torchft_non_monarch\/README.md\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">The baseline \u2014 OKE + TorchTitan + torchrun<\/a><\/p>\n<p>3.2 <a href=\"https:\/\/github.com\/oci-ai-incubations\/monarch-recipe-bp\/blob\/cfg2_oke_torchtitan_monarch\/examples\/k8s_titan_torchft_monarch\/README.md\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Adding Monarch \u2014 one controller to orchestrate them all<\/a><\/p>\n<p>3.3 <a href=\"https:\/\/github.com\/oci-ai-incubations\/monarch-recipe-bp\/blob\/cfg3_oke_torchtitan_monarch_torchft\/examples\/k8s_titan_torchft_monarch\/README.md\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Adding TorchFT \u2014 surviving the failures that <\/a><a href=\"https:\/\/github.com\/oci-ai-incubations\/monarch-recipe-bp\/blob\/cfg3_oke_torchtitan_monarch_torchft\/examples\/k8s_titan_torchft_monarch\/README.md\" rel=\"nofollow noopener\" target=\"_blank\">will<\/a><a href=\"https:\/\/github.com\/oci-ai-incubations\/monarch-recipe-bp\/blob\/cfg3_oke_torchtitan_monarch_torchft\/examples\/k8s_titan_torchft_monarch\/README.md\" rel=\"nofollow noopener\" target=\"_blank\"> happen<\/a><\/p>\n<p>3.4 <a href=\"https:\/\/github.com\/oci-ai-incubations\/monarch-recipe-bp\/blob\/cfg4_oke_torchtitan_monarch_torchft_rdma\/examples\/k8s_titan_torchft_monarch\/README.md\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Adding RDMA \u2014 paying off the TorchFT bill<\/a><\/p>\n<p>3.5 <a href=\"https:\/\/github.com\/oci-ai-incubations\/monarch-recipe-bp\/blob\/cfg5_oke_torchtitan_monarch_torchft_rdma_kueue\/examples\/k8s_titan_torchft_monarch\/README.md\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">RDMA Topology-Aware Scheduling and Gang Scheduling<\/a><\/p>\n<p>Results<\/p>\n<p>All runs train Llama 3 8B on C4 for 1000 steps on 2 \u00d7 A100 BM (16 GPUs)<\/p>\n<p>ConfigTimeLossStatusMFUTPSTFLOPsGrad NormMemory<a href=\"https:\/\/github.com\/oci-ai-incubations\/monarch-recipe-bp\/blob\/cfg1_oke_torchtitan_torchrun\/examples\/k8s_titan_torchft_non_monarch\/README.md\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">3.1<\/a> torchrun baseline2473 s12.25 \u2192 4.64Success55.34%3355172.681.065550.26 GiB<a href=\"https:\/\/github.com\/oci-ai-incubations\/monarch-recipe-bp\/blob\/cfg2_oke_torchtitan_monarch\/examples\/k8s_titan_torchft_monarch\/README.md\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">3.2<\/a> + Monarch2505 s12.27 \u2192 4.65Success55.49%3364173.131.092350.26 GiB<a href=\"https:\/\/github.com\/oci-ai-incubations\/monarch-recipe-bp\/blob\/cfg3_oke_torchtitan_monarch_torchft\/examples\/k8s_titan_torchft_monarch\/README.md\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">3.3<\/a> + TorchFT (TCP overlay)6389 s12.25 \u2192 4.65Success21.46%130166.971.242255.12 GiB<a href=\"https:\/\/github.com\/oci-ai-incubations\/monarch-recipe-bp\/blob\/cfg4_oke_torchtitan_monarch_torchft_rdma\/examples\/k8s_titan_torchft_monarch\/README.md\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">3.4<\/a> + RDMA (RoCE)2566 s12.24 \u2192 4.31Success54.25%3288169.250.966455.12 GiB<a href=\"https:\/\/github.com\/oci-ai-incubations\/monarch-recipe-bp\/blob\/cfg5_oke_torchtitan_monarch_torchft_rdma_kueue\/examples\/k8s_titan_torchft_monarch\/README.md\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">3.5<\/a> + Kueuesame as 3.4same as 3.4Success\u2014\u2014\u2014\u2014\u2014<\/p>\n<p>Notes:<\/p>\n<p>3.3: TorchFT buys per-step fault tolerance but is ~2.5\u00d7 slower \u2014 its cross-replica gradient allreduce runs over the slow TCP\/IP pod overlay instead of RDMA.<\/p>\n<p>3.4: Moving that allreduce onto the RDMA fabric (RoCE) erases the slowdown \u2014 back to baseline throughput with fault tolerance kept.<\/p>\n<p>3.5: Kueue adds gang scheduling and RDMA topology-aware admission. It\u2019s invisible during steady-state training \u2014 it changes how jobs land on and share the cluster, not the per-step math \u2014 so its throughput matches 3.4.<\/p>\n<p>Conclusion<\/p>\n<p>The system now has:<\/p>\n<p>a single Python controller (3.2 \u2014 Monarch),<\/p>\n<p>per-step fault tolerance with no restart cost (3.3 \u2014 TorchFT),<\/p>\n<p>inter-replica gradient sync at line rate on the RDMA fabric (3.4 \u2014 RDMA),<\/p>\n<p>and atomic, topology-aware admission so jobs don\u2019t strand resources or hop across the fabric (3.5 \u2014 Kueue).<\/p>\n<p>That\u2019s the system we set out to build.<\/p>\n<p>Build it yourself<\/p>\n<p>Building a reliable, large-scale training system is genuinely hard \u2014 but it doesn\u2019t have to be built from scratch. With OKE, Monarch, and the recipes in this post, the hard parts \u2014 multi-node orchestration, RDMA-aware scheduling, and per-step fault tolerance \u2014 collapse into a handful of composable, testable steps.<\/p>\n<p>\ud83d\udc49 <a href=\"https:\/\/github.com\/oci-ai-incubations\/monarch-recipe-bp\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Explore the full repo<\/a> \u2014 every configuration links to its own branch with runnable code and write-ups, so you can reproduce these results and adapt them to your own cluster.<\/p>\n<p>And to stand up the underlying cluster even faster, check out <a href=\"https:\/\/github.com\/oracle-quickstart\/oci-ai-blueprints\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">OCI AI Blueprints<\/a> for simplified, opinionated OKE provisioning.<\/p>\n","protected":false},"excerpt":{"rendered":"Training large models is hard. Training them reliably on a shared GPU cluster \u2014 where a bad NIC,&hellip;\n","protected":false},"author":2,"featured_media":90054,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,25,11053,1085,4407,7343,7344],"class_list":["post-90053","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-artificial-intelligence","tag-best-practices","tag-data-science","tag-oracle-cloud-infrastructure","tag-oracle-cloud-infrastructure-oci","tag-technical-solutions"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/90053","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=90053"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/90053\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/90054"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=90053"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=90053"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=90053"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}