• Compute
  • Customers
  • Pricing
Sign In
More Blog Posts
XDiscordLinkedInYouTube

Products

  • GPUs
  • Inference
  • Studio

Developers

  • Model library
  • Documentation
  • Glossary

Company

  • About Us
  • Blog
  • Events
  • Partnership
  • Scale
  • Career
  • Ambassador program
  • Mission & Vision

Popular models

    Stay in the loop

    By submitting, you acknowledge that we may collect and use the information you provide, which may include personal information.

    XDiscordLinkedInYouTube

    Copyright ©2026 All rights reserved.

    Privacy PolicyTerms of UseLegal Documentation
    More Blog Posts
    Other

    Multi-ControlNet With Private Nodes: Which Workflow Platform Runs It Across GPUs in Parallel?

    September 25, 2026

    Running Multi-ControlNet with private nodes across GPUs in parallel comes down to three decisions: which kind of parallelism the graph needs, which card one copy of the stack fits on, and where the private nodes execute.

    GMI Studio covers all three on one platform: its canvas "is a ComfyUI-based visual editor", and the Studio page lists "Multi-ControlNet setups", "Custom nodes & private logic", "Parallel execution across GPUs", and "Single or multi-GPU execution" on L40, A6000, A100, H100, H200, and B200 GPUs (Workflow Canvas docs, GMI Studio).

    For production pipelines, Studio Enterprise adds "Dedicated GPU clusters" and "Full architectural customization", so the stack and your private logic run on capacity reserved for your organization.

    GMI Cloud is an AI-native inference cloud that runs GPU infrastructure, model APIs, and workflow tooling on NVIDIA GPU platforms (About).

    GMI Studio is its workflow platform for image, video, audio, and LLM pipelines, and every Studio run executes on GMI Cloud's managed GPU backend, where an execution engine "schedules and runs workflows on GPUs" (Studio Introduction).

    What does "parallel" mean for a Multi-ControlNet workflow?

    For throughput across many images, a Multi-ControlNet workflow on GMI Studio or any ComfyUI-based platform runs in parallel in one of two ways, batch-level or stage-level, and picking the wrong one is why a four-card setup can run no faster than one card.

    Batch-level parallelism puts a full copy of the graph on each GPU and gives each copy different inputs: different frames, camera angles, or level layouts.

    Stage-level parallelism splits the work by component or stage boundary, so models such as the text encoder, the diffusion model, or an upscaler sit on different cards, and in a pipelined setup different jobs occupy different stages at the same time.

    (Batch-level / Stage-level)

    • What gets split | Batch-level: Inputs (images, shots, variations) | Stage-level: Components or stages (encode, sample, decode, upscale)
    • Each GPU holds | Batch-level: The whole stack: checkpoint plus every ControlNet | Stage-level: Only the models of its stage
    • Speeds up | Batch-level: Throughput: more finished images per hour | Stage-level: Utilization when one stage is much heavier than the others
    • Breaks when | Batch-level: One copy of the stack does not fit on one card | Stage-level: Stages are unbalanced, so one card waits on another
    • Related GMI Studio capability | Batch-level: "Parallel execution across GPUs"; GMI Batch Images groups up to 10 image inputs | Stage-level: "Multi-stage production graphs"; "Cross-model pipelines"
    Batch-level (conceptual): every card runs the full stack on different inputs
    
      inputs 1-4   -> GPU 0: [preprocess] -> [checkpoint + CN1 + CN2 + CN3] -> [upscale]
      inputs 5-8   -> GPU 1: [preprocess] -> [checkpoint + CN1 + CN2 + CN3] -> [upscale]
      inputs 9-12  -> GPU 2: [preprocess] -> [checkpoint + CN1 + CN2 + CN3] -> [upscale]
    
    Stage-level (conceptual): each card owns one stage, jobs flow through like an assembly line
    
      GPU 0: [depth / pose / edge preprocess]  job N+2
      GPU 1: [checkpoint + CN1 + CN2 + CN3]    job N+1
      GPU 2: [upscale / refine]                job N
    

    One split is missing from both modes: putting each ControlNet of the same chain on its own GPU.

    In upstream ComfyUI, every Apply ControlNet node links to the ControlNet before it; during sampling, each ControlNet first calls the previous one's get_control, then, inside its own start and end range, computes its control signal and merges it with the previous output (comfy/controlnet.py).

    The chain is one conditioning path for one sampling pass, and upstream's multi-GPU nodes place the diffusion model, the CLIP text encoder, and the VAE, with no node for individual ControlNets (nodes_multigpu.py).

    Plan the diffusion model and its whole ControlNet chain as one unit per card.

    Why doesn't adding GPUs to a ComfyUI server speed up the queue by itself?

    Adding GPUs does not speed up a ComfyUI server's queue by itself, because the server works through its queue one workflow at a time.

    Upstream main.py starts a single prompt_worker thread that takes queue items one by one and executes each (main.py), and the --cuda-device flag sets which CUDA devices "this instance will use" (cli_args.py).

    Upstream now also ships multi-GPU nodes for work inside one job: "MultiGPU CFG Split" prepares a model "to have sampling accelerated via splitting work units", and "Select Model Device" pins the diffusion model to a chosen GPU (nodes_multigpu.py).

    Those nodes help one job; they do not turn one queue into several.

    For throughput across many jobs, self-hosting teams start one ComfyUI process per GPU, each bound to its own device and port, then route jobs across them.

    Community extensions such as ComfyUI_NetDist and ComfyUI-Distributed package that pattern.

    Either way, the pipeline team now owns the router, the per-GPU processes, model file sync across machines, and the node-version drift between them.

    GMI Studio removes that layer.

    The canvas is ComfyUI-based, so the graph model your TDs already know carries over, while GMI Cloud's execution engine, the backend that "schedules and runs workflows on GPUs", takes the place of a router your team maintains (Studio Introduction).

    Teams that want to run their own ComfyUI fleet anyway can do it on GMI Cloud's Container Service, which provides "Kubernetes-based GPU environments" with "Elastic scaling"; Studio is the path that takes the fleet off your hands.

    Which GPU does a Multi-ControlNet stack need?

    The weights of an SDXL stack with three ControlNets take 30% of a 48 GB L40 or A6000, while a FLUX.1-dev stack with two ControlNets takes 88%, so plan the FLUX stack on an 80 GB A100 or H100 or larger.

    For batch-level throughput, the simplest sizing target is a card that keeps the checkpoint, the text encoders, the VAE, and every ControlNet resident at once, with room left for activations.

    ComfyUI can run smaller cards by moving models out of GPU memory (by default, "models will be unloaded to CPU memory after being used" (cli_args.py)), but reloads add transfer time whenever a model has to come back onto the card.

    Adding up the published weight files gives the resident budget, assuming each model loads at its file precision. The sizes below are the safetensors files on Hugging Face, as of September 2026:

    Stack (Components (file size) / Resident weights)

    • SDXL + 3 ControlNets | Components (file size): SDXL base 1.0 checkpoint incl. text encoders and VAE (6.94 GB); canny SDXL fp16 (2.50 GB); depth SDXL fp16 (2.50 GB); openpose SDXL (2.50 GB) | Resident weights: 14.44 GB
    • FLUX.1-dev + 2 ControlNets | Components (file size): FLUX.1-dev (23.80 GB); T5-XXL fp16 (9.79 GB); CLIP-L (0.25 GB); FLUX VAE (0.34 GB); Union Pro 2.0 (4.28 GB); InstantX Canny (3.58 GB) | Resident weights: 42.04 GB
    • Same FLUX stack, T5-XXL in fp8 | Components (file size): T5-XXL fp8 e4m3fn (4.89 GB) replaces the fp16 encoder | Resident weights: 37.14 GB

    Sources: SDXL base 1.0, controlnet-canny-sdxl-1.0, controlnet-depth-sdxl-1.0, controlnet-openpose-sdxl-1.0, FLUX.1-dev, flux_text_encoders, FLUX.1-dev ControlNet Union Pro 2.0, FLUX.1-dev ControlNet Canny.

    Set those totals against the GPUs GMI Studio lists, with memory from NVIDIA's own spec pages:

    L40

    • Memory (NVIDIA spec): 48 GB GDDR6 (NVIDIA L40)
    • SDXL + 3 CN share: 30%
    • FLUX + 2 CN share: 88%
    • Headroom left for FLUX stack: 5.96 GB

    A6000

    • Memory (NVIDIA spec): 48 GB GDDR6 (NVIDIA RTX A6000)
    • SDXL + 3 CN share: 30%
    • FLUX + 2 CN share: 88%
    • Headroom left for FLUX stack: 5.96 GB

    A100

    • Memory (NVIDIA spec): 80 GB HBM2e (NVIDIA A100)
    • SDXL + 3 CN share: 18%
    • FLUX + 2 CN share: 53%
    • Headroom left for FLUX stack: 37.96 GB

    H100

    • Memory (NVIDIA spec): 80 GB (NVIDIA H100)
    • SDXL + 3 CN share: 18%
    • FLUX + 2 CN share: 53%
    • Headroom left for FLUX stack: 37.96 GB

    H200

    • Memory (NVIDIA spec): 141 GB HBM3e (NVIDIA H200)
    • SDXL + 3 CN share: 10%
    • FLUX + 2 CN share: 30%
    • Headroom left for FLUX stack: 98.96 GB

    B200

    • Memory (NVIDIA spec): 180 GB per GPU (DGX B200: 1,440 GB across 8 GPUs, NVIDIA DGX B200)
    • SDXL + 3 CN share: 8%
    • FLUX + 2 CN share: 23%
    • Headroom left for FLUX stack: 137.96 GB

    A practical sizing rule for batch-level throughput: replicate the full stack per card (batch-level) when resident weights take 60% of card memory or less, because latents, ControlNet hint images, and preprocessor models such as depth or pose estimators need the rest.

    By that rule an SDXL stack with three ControlNets fits the budget on every GPU class GMI Studio lists, including the 48 GB L40. A FLUX.1-dev stack with two ControlNets belongs on an 80 GB A100 or H100 or larger; switching T5-XXL to fp8 brings it to 46% of an 80 GB card but still 77% of a 48 GB card.

    Treat the table as the starting budget, then confirm the peak with one real run at production resolution before committing a batch.

    How do private nodes and ComfyUI nodes fit together in GMI Studio?

    In GMI Studio, native ComfyUI nodes, GMI Official nodes, and your own logic sit on the same canvas and are wired the same way. The docs separate two node families:

    (GMI Official Nodes / ComfyUI Nodes)

    • Best for | GMI Official Nodes: "Quick setup, managed inference" | ComfyUI Nodes: "Fine-grained control, custom logic"
    • Control level | GMI Official Nodes: "Simplified inputs / outputs" | ComfyUI Nodes: "Full node-level customization"
    • Performance | GMI Official Nodes: "Optimized, fully managed" | ComfyUI Nodes: "Depends on workflow configuration"

    ComfyUI nodes "are sourced directly from the upstream Comfy repo for users who need lower-level building blocks", and the Comfy Library palette opens "the full Comfy node tree" (Workflow Canvas docs).

    In the universal node search, the Comfy filter narrows results to "anything sourced from the upstream ComfyUI repo" (Library, Search & Blueprints).

    A Multi-ControlNet graph therefore keeps the ComfyUI shape a pipeline TD expects, and "Depends on workflow configuration" is the reminder that the ComfyUI part of the graph performs as well as you configure it, which is what the sizing table above is for.

    Private logic is a listed Studio capability: "Custom nodes & private logic" appears under "Run any model without refactoring your stack", next to "Multi-ControlNet setups" and "Model loading & versioning" (GMI Studio).

    For a studio whose private nodes are proprietary code, Studio Enterprise is "For organizations operating business-critical AI workflows" and adds "Full architectural customization", "Org-level workflow ownership", "Dedicated GPU clusters", and "Production workflow support".

    Bring your private node list and ControlNet checkpoints into the Studio Enterprise conversation, so the environment is set up around them before the first production batch.

    How do you prepare a Multi-ControlNet graph for parallel runs in GMI Studio?

    Prepare the graph in GMI Studio so that each copy is small enough to replicate, each branch can be tested alone, and the inputs that vary per job are the only inputs anyone touches. Step 1 is arithmetic, and steps 2 to 7 are canvas operations documented for GMI Studio:

    1. Size the stack first. Add up checkpoint, encoders, VAE, and every ControlNet as in the table above, and pick the GPU class where the total stays within the 60% planning budget.
    2. Batch the inputs. For multi-image input to a GMI generation node, the documented pattern is up to 10 GMI Image Upload nodes connected to GMI __ Utils __ Image __ GMI Batch Images, with the batched output wired to the generation node (Tutorials). This batches images inside one run, which is separate from spreading many runs across GPUs.
    3. Test one control branch at a time. The node toolbar's Run Branch runs "just this node and upstream", so a depth preprocessor can be checked without paying for a full sampling pass (Workflow Canvas docs).
    4. A/B a ControlNet without rewiring. Each node's Settings tab has a node state of Normal, Bypass, or Mute; bypass the pose branch for one run to see what it actually contributes.
    5. Collapse the control stack. Convert to Subgraph turns the chain into one node with its own inputs and outputs; saving it as a reusable Blueprint and versioning it is covered in Seedream 5.0 Pro and private image workflows.
    6. Expose only the per-job inputs. Favorite the control images and ControlNet strengths, so Workflow Overview > Parameters lists exactly the inputs that change from one job to the next.
    7. Add a Save node for every output you need. "You must add a Save node to retrieve workflow outputs", and nodes "execute in the order defined by the workflow graph, respecting dependencies" (Running a Workflow).

    The order matters for cost as well: "Runs spend credits according to the model nodes in the graph", and Run Branch executes only the selected node and what feeds it, so a preprocessor check does not have to pay for every model node downstream (Workflow Canvas docs).

    With the graph prepared, a multi-GPU setup is sized from four numbers. Send them to GMI Cloud when requesting Studio Enterprise capacity:

    Number (Where it comes from / What it decides)

    • Resident weights per copy | Where it comes from: Step 1 total (for example, 14.44 GB or 42.04 GB) | What it decides: GPU class
    • Peak memory at production resolution | Where it comes from: One real test run | What it decides: Whether the 60% rule holds for your graph
    • Jobs per day and runtime per job | Where it comes from: Your shot list and one timed run | What it decides: Number of GPUs running in parallel
    • Turnaround target | Where it comes from: Your delivery schedule | What it decides: Shared capacity or dedicated GPU clusters

    Which setup should a pipeline TD choose?

    On GMI Studio, plan stacks sized like the SDXL example on 48 GB L40 or A6000 cards, stacks sized like the FLUX example on 80 GB A100 or H100 cards or larger, and move to Studio Enterprise dedicated GPU clusters when a deadline depends on queue time. The table below turns the sizing math into a decision:

    Your situation (Parallel mode / GPU class / GMI Studio tier)

    • SDXL-class stack, 2 to 4 ControlNets, many variations per shot | Parallel mode: Batch-level, full stack per card | GPU class: L40 or A6000 (48 GB) | GMI Studio tier: Studio, "Single or multi-GPU execution"
    • FLUX-class stack with 2 or more ControlNets | Parallel mode: Batch-level on larger cards | GPU class: A100 or H100 (80 GB), H200 for high resolution | GMI Studio tier: Studio
    • One stage dominates, such as heavy preprocessing or a 4K upscale pass | Parallel mode: Stage-level, split at that boundary | GPU class: Sized by the heaviest stage | GMI Studio tier: Scope the split with GMI Cloud under Studio Enterprise
    • Delivery dates depend on queue time; proprietary nodes | Parallel mode: Either, on reserved capacity | GPU class: Sized with GMI Cloud | GMI Studio tier: Studio Enterprise, "Dedicated GPU clusters"

    On shared capacity, the Studio FAQ notes that "Queue times depend on GPU availability and overall workload demand" (Studio FAQ).

    "No shared queues" is listed under Studio's Dedicated Infrastructure, which is the tier to plan for when a render day cannot slip (GMI Studio).

    Production media teams already build on this stack: at Summer Signal '26, Utopai Studios' CTO presented agentic filmmaking with PAI 3.0 on the same night GMI Cloud walked through running frontier video models on its infrastructure.

    Not every consistency problem needs another ControlNet.

    When the goal is keeping a product or character's look rather than locking an exact layout, reference-image models can carry part of the load: in GMI Cloud's Hy Image 3.5 Preview test, two reference mugs placed on one counter "kept their look".

    Keep ControlNets for structure (depth, edges, pose) and drop the ones that only restate identity; every ControlNet model dropped from the stack takes its weights off every card.

    FAQ

    How many GPUs can a GMI Studio workflow run on in parallel?

    GMI Studio lists "Single or multi-GPU execution" and "Parallel execution across GPUs" for its workflows, and Studio Enterprise provides "Dedicated GPU clusters" whose size is scoped with GMI Cloud for your workload (GMI Studio).

    For batch-level runs, each GPU holds one full copy of the Multi-ControlNet stack, so the number of useful cards follows the number of independent inputs in the batch.

    Which GPUs can run a Multi-ControlNet stack in GMI Studio?

    GMI Studio lists L40, A6000, A100, H100, H200, and B200 GPUs (GMI Studio).

    An SDXL stack with three ControlNets has 14.44 GB of resident weights and fits a 60% budget on every class; a FLUX.1-dev stack with two ControlNets has 42.04 GB, 88% of a 48 GB card, so plan it on 80 GB A100 or H100 GPUs or larger.

    Can I use ComfyUI nodes in GMI Studio?

    Yes. The GMI Studio canvas "is a ComfyUI-based visual editor", and its ComfyUI nodes "are sourced directly from the upstream Comfy repo"; the Comfy Library palette shows the full Comfy node tree (Workflow Canvas docs).

    GMI Official nodes sit on the same canvas for managed model inference.

    Can two ControlNets in the same chain run on different GPUs?

    No, not with upstream ComfyUI's ControlNet design: each ControlNet in a chain calls the previous one and merges its output into the same conditioning path during sampling, and upstream's multi-GPU nodes place the diffusion model, CLIP, and VAE, not individual ControlNets (comfy/controlnet.py, nodes_multigpu.py).

    In GMI Studio, parallelize across copies of the whole chain (batch-level) or across graph stages such as preprocessing and upscaling (stage-level) instead.

    Where do private nodes run in GMI Studio?

    All GMI Studio execution runs on GMI Cloud's managed GPU backend, so private logic in a Studio workflow runs on GMI Cloud GPUs, not on local hardware (Studio Introduction).

    "Custom nodes & private logic" is a listed Studio capability, and Studio Enterprise adds "Dedicated GPU clusters" and "Full architectural customization" for studios whose private nodes are proprietary.

    Next step: size your stack and run one batch in GMI Studio

    Add up your checkpoint and ControlNet weights, pick the GPU class from the table, and build the graph in GMI Studio with a single test batch.

    When the pipeline carries delivery dates or proprietary nodes, contact GMI Cloud sales about Studio Enterprise and dedicated GPU clusters; the full GPU lineup is on the GMI Cloud GPU page.

    For version control and rollback of the same workflow, see updating and rolling back models in a private video workflow. We can help map your current ComfyUI graph and private nodes onto a Studio setup sized for your stack.

    Colin Mo

    Build AI Without Limits

    GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

    FAQ

    GMI Studio lists "Single or multi-GPU execution" and "Parallel execution across GPUs" for its workflows, and Studio Enterprise provides "Dedicated GPU clusters" whose size is scoped with GMI Cloud for your workload (GMI Studio). For batch-level runs, each GPU holds one full copy of the Multi-ControlNet stack, so the number of useful cards follows the number of independent inputs in the batch.

    Ready to build?

    Explore powerful AI models and launch your project in just a few clicks.

    Get Started