September 25, 2026
Running Multi-ControlNet with private nodes across GPUs in parallel comes down to three decisions: which kind of parallelism the graph needs, which card one copy of the stack fits on, and where the private nodes execute.
GMI Studio covers all three on one platform: its canvas "is a ComfyUI-based visual editor", and the Studio page lists "Multi-ControlNet setups", "Custom nodes & private logic", "Parallel execution across GPUs", and "Single or multi-GPU execution" on L40, A6000, A100, H100, H200, and B200 GPUs (Workflow Canvas docs, GMI Studio).
For production pipelines, Studio Enterprise adds "Dedicated GPU clusters" and "Full architectural customization", so the stack and your private logic run on capacity reserved for your organization.
GMI Cloud is an AI-native inference cloud that runs GPU infrastructure, model APIs, and workflow tooling on NVIDIA GPU platforms (About).
GMI Studio is its workflow platform for image, video, audio, and LLM pipelines, and every Studio run executes on GMI Cloud's managed GPU backend, where an execution engine "schedules and runs workflows on GPUs" (Studio Introduction).
For throughput across many images, a Multi-ControlNet workflow on GMI Studio or any ComfyUI-based platform runs in parallel in one of two ways, batch-level or stage-level, and picking the wrong one is why a four-card setup can run no faster than one card.
Batch-level parallelism puts a full copy of the graph on each GPU and gives each copy different inputs: different frames, camera angles, or level layouts.
Stage-level parallelism splits the work by component or stage boundary, so models such as the text encoder, the diffusion model, or an upscaler sit on different cards, and in a pipelined setup different jobs occupy different stages at the same time.
(Batch-level / Stage-level)
Batch-level (conceptual): every card runs the full stack on different inputs
inputs 1-4 -> GPU 0: [preprocess] -> [checkpoint + CN1 + CN2 + CN3] -> [upscale]
inputs 5-8 -> GPU 1: [preprocess] -> [checkpoint + CN1 + CN2 + CN3] -> [upscale]
inputs 9-12 -> GPU 2: [preprocess] -> [checkpoint + CN1 + CN2 + CN3] -> [upscale]
Stage-level (conceptual): each card owns one stage, jobs flow through like an assembly line
GPU 0: [depth / pose / edge preprocess] job N+2
GPU 1: [checkpoint + CN1 + CN2 + CN3] job N+1
GPU 2: [upscale / refine] job N
One split is missing from both modes: putting each ControlNet of the same chain on its own GPU.
In upstream ComfyUI, every Apply ControlNet node links to the ControlNet before it; during sampling, each ControlNet first calls the previous one's get_control, then, inside its own start and end range, computes its control signal and merges it with the previous output (comfy/controlnet.py).
The chain is one conditioning path for one sampling pass, and upstream's multi-GPU nodes place the diffusion model, the CLIP text encoder, and the VAE, with no node for individual ControlNets (nodes_multigpu.py).
Plan the diffusion model and its whole ControlNet chain as one unit per card.
Adding GPUs does not speed up a ComfyUI server's queue by itself, because the server works through its queue one workflow at a time.
Upstream main.py starts a single prompt_worker thread that takes queue items one by one and executes each (main.py), and the --cuda-device flag sets which CUDA devices "this instance will use" (cli_args.py).
Upstream now also ships multi-GPU nodes for work inside one job: "MultiGPU CFG Split" prepares a model "to have sampling accelerated via splitting work units", and "Select Model Device" pins the diffusion model to a chosen GPU (nodes_multigpu.py).
Those nodes help one job; they do not turn one queue into several.
For throughput across many jobs, self-hosting teams start one ComfyUI process per GPU, each bound to its own device and port, then route jobs across them.
Community extensions such as ComfyUI_NetDist and ComfyUI-Distributed package that pattern.
Either way, the pipeline team now owns the router, the per-GPU processes, model file sync across machines, and the node-version drift between them.
GMI Studio removes that layer.
The canvas is ComfyUI-based, so the graph model your TDs already know carries over, while GMI Cloud's execution engine, the backend that "schedules and runs workflows on GPUs", takes the place of a router your team maintains (Studio Introduction).
Teams that want to run their own ComfyUI fleet anyway can do it on GMI Cloud's Container Service, which provides "Kubernetes-based GPU environments" with "Elastic scaling"; Studio is the path that takes the fleet off your hands.
The weights of an SDXL stack with three ControlNets take 30% of a 48 GB L40 or A6000, while a FLUX.1-dev stack with two ControlNets takes 88%, so plan the FLUX stack on an 80 GB A100 or H100 or larger.
For batch-level throughput, the simplest sizing target is a card that keeps the checkpoint, the text encoders, the VAE, and every ControlNet resident at once, with room left for activations.
ComfyUI can run smaller cards by moving models out of GPU memory (by default, "models will be unloaded to CPU memory after being used" (cli_args.py)), but reloads add transfer time whenever a model has to come back onto the card.
Adding up the published weight files gives the resident budget, assuming each model loads at its file precision. The sizes below are the safetensors files on Hugging Face, as of September 2026:
Stack (Components (file size) / Resident weights)
Sources: SDXL base 1.0, controlnet-canny-sdxl-1.0, controlnet-depth-sdxl-1.0, controlnet-openpose-sdxl-1.0, FLUX.1-dev, flux_text_encoders, FLUX.1-dev ControlNet Union Pro 2.0, FLUX.1-dev ControlNet Canny.
Set those totals against the GPUs GMI Studio lists, with memory from NVIDIA's own spec pages:
A practical sizing rule for batch-level throughput: replicate the full stack per card (batch-level) when resident weights take 60% of card memory or less, because latents, ControlNet hint images, and preprocessor models such as depth or pose estimators need the rest.
By that rule an SDXL stack with three ControlNets fits the budget on every GPU class GMI Studio lists, including the 48 GB L40. A FLUX.1-dev stack with two ControlNets belongs on an 80 GB A100 or H100 or larger; switching T5-XXL to fp8 brings it to 46% of an 80 GB card but still 77% of a 48 GB card.
Treat the table as the starting budget, then confirm the peak with one real run at production resolution before committing a batch.
In GMI Studio, native ComfyUI nodes, GMI Official nodes, and your own logic sit on the same canvas and are wired the same way. The docs separate two node families:
(GMI Official Nodes / ComfyUI Nodes)
ComfyUI nodes "are sourced directly from the upstream Comfy repo for users who need lower-level building blocks", and the Comfy Library palette opens "the full Comfy node tree" (Workflow Canvas docs).
In the universal node search, the Comfy filter narrows results to "anything sourced from the upstream ComfyUI repo" (Library, Search & Blueprints).
A Multi-ControlNet graph therefore keeps the ComfyUI shape a pipeline TD expects, and "Depends on workflow configuration" is the reminder that the ComfyUI part of the graph performs as well as you configure it, which is what the sizing table above is for.
Private logic is a listed Studio capability: "Custom nodes & private logic" appears under "Run any model without refactoring your stack", next to "Multi-ControlNet setups" and "Model loading & versioning" (GMI Studio).
For a studio whose private nodes are proprietary code, Studio Enterprise is "For organizations operating business-critical AI workflows" and adds "Full architectural customization", "Org-level workflow ownership", "Dedicated GPU clusters", and "Production workflow support".
Bring your private node list and ControlNet checkpoints into the Studio Enterprise conversation, so the environment is set up around them before the first production batch.
Prepare the graph in GMI Studio so that each copy is small enough to replicate, each branch can be tested alone, and the inputs that vary per job are the only inputs anyone touches. Step 1 is arithmetic, and steps 2 to 7 are canvas operations documented for GMI Studio:
Normal, Bypass, or Mute; bypass the pose branch for one run to see what it actually contributes.The order matters for cost as well: "Runs spend credits according to the model nodes in the graph", and Run Branch executes only the selected node and what feeds it, so a preprocessor check does not have to pay for every model node downstream (Workflow Canvas docs).
With the graph prepared, a multi-GPU setup is sized from four numbers. Send them to GMI Cloud when requesting Studio Enterprise capacity:
Number (Where it comes from / What it decides)
On GMI Studio, plan stacks sized like the SDXL example on 48 GB L40 or A6000 cards, stacks sized like the FLUX example on 80 GB A100 or H100 cards or larger, and move to Studio Enterprise dedicated GPU clusters when a deadline depends on queue time. The table below turns the sizing math into a decision:
Your situation (Parallel mode / GPU class / GMI Studio tier)
On shared capacity, the Studio FAQ notes that "Queue times depend on GPU availability and overall workload demand" (Studio FAQ).
"No shared queues" is listed under Studio's Dedicated Infrastructure, which is the tier to plan for when a render day cannot slip (GMI Studio).
Production media teams already build on this stack: at Summer Signal '26, Utopai Studios' CTO presented agentic filmmaking with PAI 3.0 on the same night GMI Cloud walked through running frontier video models on its infrastructure.
Not every consistency problem needs another ControlNet.
When the goal is keeping a product or character's look rather than locking an exact layout, reference-image models can carry part of the load: in GMI Cloud's Hy Image 3.5 Preview test, two reference mugs placed on one counter "kept their look".
Keep ControlNets for structure (depth, edges, pose) and drop the ones that only restate identity; every ControlNet model dropped from the stack takes its weights off every card.
GMI Studio lists "Single or multi-GPU execution" and "Parallel execution across GPUs" for its workflows, and Studio Enterprise provides "Dedicated GPU clusters" whose size is scoped with GMI Cloud for your workload (GMI Studio).
For batch-level runs, each GPU holds one full copy of the Multi-ControlNet stack, so the number of useful cards follows the number of independent inputs in the batch.
GMI Studio lists L40, A6000, A100, H100, H200, and B200 GPUs (GMI Studio).
An SDXL stack with three ControlNets has 14.44 GB of resident weights and fits a 60% budget on every class; a FLUX.1-dev stack with two ControlNets has 42.04 GB, 88% of a 48 GB card, so plan it on 80 GB A100 or H100 GPUs or larger.
Yes. The GMI Studio canvas "is a ComfyUI-based visual editor", and its ComfyUI nodes "are sourced directly from the upstream Comfy repo"; the Comfy Library palette shows the full Comfy node tree (Workflow Canvas docs).
GMI Official nodes sit on the same canvas for managed model inference.
No, not with upstream ComfyUI's ControlNet design: each ControlNet in a chain calls the previous one and merges its output into the same conditioning path during sampling, and upstream's multi-GPU nodes place the diffusion model, CLIP, and VAE, not individual ControlNets (comfy/controlnet.py, nodes_multigpu.py).
In GMI Studio, parallelize across copies of the whole chain (batch-level) or across graph stages such as preprocessing and upscaling (stage-level) instead.
All GMI Studio execution runs on GMI Cloud's managed GPU backend, so private logic in a Studio workflow runs on GMI Cloud GPUs, not on local hardware (Studio Introduction).
"Custom nodes & private logic" is a listed Studio capability, and Studio Enterprise adds "Dedicated GPU clusters" and "Full architectural customization" for studios whose private nodes are proprietary.
Add up your checkpoint and ControlNet weights, pick the GPU class from the table, and build the graph in GMI Studio with a single test batch.
When the pipeline carries delivery dates or proprietary nodes, contact GMI Cloud sales about Studio Enterprise and dedicated GPU clusters; the full GPU lineup is on the GMI Cloud GPU page.
For version control and rollback of the same workflow, see updating and rolling back models in a private video workflow. We can help map your current ComfyUI graph and private nodes onto a Studio setup sized for your stack.
Colin Mo
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
