September 25, 2026
A 5-second 720p clip costs $0.30 on Luma Ray 3.2 and $0.56 on Kling 3.0 Turbo, but at 1080p the order flips: $1.20 for Luma against $0.70 for Kling (GMI Cloud list prices as of September 2026, Luma Ray 3.2 docs, Kling 3.0 Turbo docs).
To test the two models together, use GMI Cloud's Model-as-a-Service (MaaS) platform, which serves luma-ray-3-2-generate, luma-ray-3-2-edit, luma-ray-3-2-reframe, kling-3.0-turbo-t2v, and kling-3.0-turbo-i2v behind one API key and one request-queue endpoint, so both models are submitted, polled, and billed the same way.
Because the cost winner depends on resolution, a fair bake-off has to be scored at the resolution you will actually deliver.
This guide is for the creative technology lead at an agency or video app who needs to run the same shot list through both models and pick a primary one.
It covers the provider options, the resolution cost flip, a parameter map for like-for-like requests, a bake-off budget, a harness that runs both models through one queue, and a scoring rule with thresholds.
GMI Cloud MaaS is the strongest place to run the test, because both models sit behind the same key, the same submit endpoint (POST /api/v1/ie/requestqueue/apikey/requests), and the same status call (GET .../requests/{request_id}).
GMI Cloud is the AI-native inference cloud; its MaaS page describes the layer as "A Model-as-a-Service platform for LLM, image, video, and audio models, with unified APIs, discounted pricing, and enterprise-grade guarantees" (MaaS).
1. GMI Cloud MaaS (recommended). Five endpoints cover the full Luma Ray 3.2 and Kling 3.0 Turbo surface: Luma generate (text-to-video, image-to-video, extend, interpolation), Luma edit, Luma reframe, Kling text-to-video, and Kling image-to-video. For a bake-off, that setup provides:
{"model": ..., "payload": {...}} against the same endpoint, and every job returns a request_id you poll until success or failed (Video API reference). Only the payload fields differ, and the mapping table below covers them.2. fal. fal lists both models as separate endpoints.
Its Kling V3 Turbo pages price Standard at $0.112 per second and Pro at $0.14 per second, and its Luma Ray 3.2 pages price a 5-second 720p clip at $0.30 for image-to-video and $1 for text-to-video (fal Kling V3 Turbo, fal Luma Ray 3.2).
3. Luma Agents API. Luma's own API serves Luma's models only, so Kling needs a second account, key, and integration. Its SDR create tier lists $0.15, $0.30, and $1.20 per 5 seconds at 540p, 720p, and 1080p (Luma pricing).
4. Kling Developer Platform. Kling's own platform serves Kling models only, which puts Luma on a separate integration and bill (Kling AI API).
5.
Replicate. Replicate lists Luma Ray 3.2 as luma/ray-3.2, and its Kling 3.0 listings are Kling Video 3.0 (kwaivgi/kling-v3-video) and Kling Video 3.0 Omni; check its catalog for a 3.0 Turbo listing before planning a Turbo test there (Replicate Luma Ray 3.2, Replicate Kling models).
Luma Ray 3.2 and Kling 3.0 Turbo use different billing units, and Luma's 1080p rate is four times its 720p rate while Kling's rises by only 25%. On GMI Cloud, Kling 3.0 Turbo costs "$0.112 per second ($0.14 per second at 1080p)" and defaults to 720p.
For Luma Ray 3.2 generate, the docs state "Billed at the create-tier per 5-second block," with SDR blocks at $0.15 (540p), $0.30 (720p), and $1.20 (1080p), and "a 10s video bills as two 5s blocks" (Kling docs, Luma docs).
Clip (SDR, GMI Cloud list prices, Sept 2026) (Kling 3.0 Turbo / Luma Ray 3.2 generate / Cheaper model)
The flip decides the test plan. A screening round at 720p makes Luma look cheap and a delivery round at 1080p makes Kling look cheap, so compare cost at the delivery resolution only.
Each 1080p render is also a new take rather than an upscale of the 720p one, so the finals get scored again instead of inheriting screening scores.
HDR stays out of the head-to-head: Luma HDR is priced separately at $1.20 (720p) and $4.80 (1080p) per 5-second block on GMI Cloud (Luma docs), and the Kling 3.0 Turbo quickstart lists no HDR option.
Hold duration, resolution, aspect ratio, and prompt constant across both models, and restrict each one to the values the other also accepts. The two quickstarts define overlapping but different ranges:
Setting (kling-3.0-turbo-t2v / -i2v / luma-ray-3-2-generate / Bake-off value)
Note the duration format: Kling takes "5", Luma takes "5s". A harness that sends one format to both models will be rejected by one of them, so build the payload per model, as the harness below does.
A two-stage bake-off of 20 prompts, 3 takes each, costs about $80 at GMI Cloud list prices: screen everything at 720p, then re-render the top 5 prompts at 1080p. Running every clip at 1080p from the start costs $114 for the same shot list.
Switching the test to 10-second clips doubles every figure, since Kling bills per second and Luma bills two 5-second blocks.
Plan the submission rate before the run. GMI Cloud rate-limits video models in requests per hour (RPH), enforced at the organization level, with quotas set per model and per usage tier.
Tiers upgrade automatically within 24 hours of cumulative credit purchases of $50, $500, and $1,000, and vouchers do not count (rate limits).
The screening stage above sends 60 requests to each model; if the tier's hourly quota is lower, move up a tier before the run, since credit purchases upgrade the organization within 24 hours, or pace the batch across hours.
Submit Luma Ray 3.2 and Kling 3.0 Turbo takes interleaved from one script to the GMI Cloud request queue, so queue conditions at any moment affect both equally, and poll every request_id against one time budget.
The harness below does that on GMI Cloud MaaS, records the list-price cost of each clip, and writes a CSV with empty scoring columns for the review session.
import csv
import json
import os
import time
import requests
URL = "https://console.gmicloud.ai/api/v1/ie/requestqueue/apikey/requests"
HEADERS = {
"Authorization": f"Bearer {os.environ['GMI_API_KEY']}",
"Content-Type": "application/json",
}
MODELS = ["kling-3.0-turbo-t2v", "luma-ray-3-2-generate"]
# GMI Cloud list prices from the model quickstarts, September 2026. Re-check before each run.
KLING_PER_SECOND = {"720p": 0.112, "1080p": 0.14}
LUMA_PER_5S_BLOCK = {"720p": 0.30, "1080p": 1.20} # SDR create tier
RESOLUTION = "720p" # "720p" or "1080p": the only values both models accept
SECONDS = 5 # 5 or 10: the only durations both models accept
ASPECT = "16:9"
TAKES = 3
BUDGET_SECONDS = 1800 # whole run: submitting and polling
POLL_INTERVAL = 10
PROMPTS = [
"Slow push-in on a ceramic mug on a sunlit kitchen counter, steam rising",
"Handheld tracking shot of a cyclist crossing a rainy city intersection at night",
]
assert RESOLUTION in KLING_PER_SECOND and SECONDS in (5, 10)
def build_payload(model, prompt):
if model.startswith("kling"):
duration = str(SECONDS) # Kling expects "5"
else:
duration = f"{SECONDS}s" # Luma expects "5s"
return {"prompt": prompt, "resolution": RESOLUTION,
"aspect_ratio": ASPECT, "duration": duration}
def list_cost(model):
if model.startswith("kling"):
return KLING_PER_SECOND[RESOLUTION] * SECONDS
return LUMA_PER_5S_BLOCK[RESOLUTION] * (SECONDS // 5)
def result_url(outcome):
if outcome.get("video_url"): # Luma
return outcome["video_url"]
urls = outcome.get("media_urls") or [] # Kling
if isinstance(urls, list) and urls:
first = urls[0]
return first.get("url", "") if isinstance(first, dict) else first
return ""
def submit(model, prompt):
r = requests.post(URL, headers=HEADERS, timeout=30,
json={"model": model, "payload": build_payload(model, prompt)})
if r.status_code == 429:
raise RuntimeError(f"{model}: hourly request quota reached")
r.raise_for_status()
return r.json()["request_id"]
def poll_once(pending, deadline):
for request_id in list(pending):
remaining = deadline - time.monotonic()
if remaining <= 0:
return
try:
r = requests.get(f"{URL}/{request_id}", headers=HEADERS,
timeout=min(30, remaining))
r.raise_for_status()
body = r.json()
except (requests.RequestException, ValueError):
continue # transient error: retry on the next pass
status = body.get("status")
if status not in ("success", "failed"):
continue
job = pending.pop(request_id)
job["observed_seconds"] = round(time.monotonic() - job["submitted_at"], 1)
job["url"] = result_url(body.get("outcome") or {}) if status == "success" else ""
job["status"] = status if (status == "failed" or job["url"]) else "success_no_url"
jobs, pending = [], {}
deadline = time.monotonic() + BUDGET_SECONDS
with open("submitted.jsonl", "a") as log: # accepted request IDs survive a crash
try:
for prompt_id, prompt in enumerate(PROMPTS):
for take in range(TAKES):
n = prompt_id * TAKES + take
order = MODELS if n % 2 == 0 else MODELS[::-1] # alternate who goes first
for model in order:
if time.monotonic() >= deadline:
raise RuntimeError("time budget used up during submission")
job = {"prompt_id": prompt_id, "take": take, "model": model,
"resolution": RESOLUTION, "seconds": SECONDS,
"list_cost_usd": round(list_cost(model), 3),
"request_id": submit(model, prompt),
"submitted_at": time.monotonic()}
jobs.append(job)
pending[job["request_id"]] = job
log.write(json.dumps({k: job[k] for k in ("prompt_id", "take", "model", "request_id")}) + "\n")
log.flush()
poll_once(pending, deadline) # keep polling while submitting
except (RuntimeError, requests.RequestException, KeyError, ValueError) as err:
print(f"Submission stopped ({err}); polling the {len(pending)} accepted jobs")
while pending and time.monotonic() < deadline:
time.sleep(min(POLL_INTERVAL, max(0.0, deadline - time.monotonic())))
poll_once(pending, deadline)
for job in pending.values():
job.update(status="timeout", url="", observed_seconds="") # may still finish: check before re-submitting
fields = ["prompt_id", "take", "model", "resolution", "seconds", "status", "url",
"observed_seconds", "list_cost_usd", "request_id",
"adherence_1to5", "motion_1to5", "consistency_1to5", "usable_yes_no"]
with open("bakeoff.csv", "w", newline="") as f:
writer = csv.DictWriter(f, fieldnames=fields, extrasaction="ignore")
writer.writeheader()
writer.writerows(jobs)
The harness covers four failure points:
BUDGET_SECONDS is used up, the harness stops sending new submissions and new status calls, and each status call's timeout shrinks to the time left. GMI Cloud's Luma docs note that "a 5s/720p clip is usually well under two minutes," while 10-second, 1080p, or HDR jobs "can run several times longer," so raise the budget for the finals.request_id is written to submitted.jsonl the moment it comes back. A rejected submission (such as a 429 once the hourly quota is used) stops new submissions, and the accepted jobs are still polled and written to the CSV.observed_seconds is the time until the harness sees a finished job. Polling runs during submission and the model order alternates on every take, so neither model always goes first; treat it as a rough relative signal, not render time.list_cost_usd is a list-price estimate; the charged amount is what appears in the GMI Cloud console.Score the Luma Ray 3.2 and Kling 3.0 Turbo clips blind, then pick the model with the lower cost per usable clip at your delivery resolution.
Rename each downloaded clip to a random ID (for example p03-t2-7f3a.mp4) and keep the ID-to-request_id map in a separate file the reviewers never see; after scoring, join the scores back to bakeoff.csv on request_id to reveal the model.
Criterion (Weight / How to score)
Consistency earns a fifth of the weight because it is where production video breaks.
The GMI Cloud write-up of a four-hour film sprint puts it plainly: "Almost every team hit it: the same character, the same product, the same look across dozens of shots" (What 23 films made in four hours taught us).
The pass-rate thresholds follow directly from the price ratios in the cost table:
Score Luma Ray 3.2 edit and reframe as a separate Luma-only round, because they transform an existing clip rather than generate one, and Kling 3.0 Turbo's two GMI Cloud endpoints are text-to-video and image-to-video.
Both take the source as source_generation_id, "the id of a prior completed video generation owned by the same API key," or inline as base64 source_data, and the source "must be 30 seconds or shorter" (Luma edit docs, Luma reframe docs).
Endpoint (What it does / SDR price per 5-second block on GMI Cloud (540p / 720p / 1080p))
For an agency shipping one hero shot in 16:9, 9:16, and 1:1, reframe is the relevant test: two extra aspect ratios of a 5-second 720p master cost $0.60 and keep the approved content, whereas regenerating each ratio produces a new take that has to be reviewed again.
After the bake-off, keep the winning model and the runner-up on the same GMI Cloud MaaS key, so a fallback or a re-test reuses the same endpoint and the payload adapter from the harness, and build the production pipeline from there:
Kling 3.0 Turbo is billed per second of output: $0.112 per second at 720p and $0.14 per second at 1080p, for clips of 3 to 15 seconds. Luma Ray 3.2 generate is billed per 5-second block by resolution and dynamic range: $0.15, $0.30, and $1.20 at 540p, 720p, and 1080p in SDR, with HDR priced separately.
Both are GMI Cloud list prices as of September 2026 (Kling docs, Luma docs) and appear on the same invoice.
Yes. Add https://mcp.gmicloud.ai/mcp as a custom connector in Claude and sign in with a GMI Cloud account; no API key is required.
The GMI MCP Server FAQ lists Kling and Luma Ray among the supported video model families and suggests asking the agent to search for currently available models, and the agent can request a free cost estimate before each generation; charges are based on the model, generation settings, and usage.
Kling 3.0 Turbo is cheaper at 1080p on GMI Cloud: a 5-second clip costs $0.70, against $1.20 for Luma Ray 3.2 generate in SDR. At 720p the order reverses, $0.30 for Luma against $0.56 for Kling. Compare the two at the resolution you will deliver, not the one you screen at.
Yes, Kling 3.0 Turbo and Luma Ray 3.2 can run the same image-to-video test on GMI Cloud for 5-second clips.
Kling uses the kling-3.0-turbo-i2v model with a first_frame image, and Luma uses the same luma-ray-3-2-generate model with start_frame_url; Luma does not support a start frame with 10-second duration.
Use the same source image and prompt for both, and note that Kling's output framing follows the source image's dimensions.
GMI Cloud limits video models in requests per hour, per model, at the organization level, and the quota depends on the usage tier.
New organizations start at Tier 1 and move up within 24 hours of cumulative credit purchases of $50, $500, and $1,000, and [email protected] handles manual tier upgrades (rate limits).
Before a large screening batch, ask [email protected] for your tier's hourly quota on each video model and size the batch to it.
Create an account in the GMI Cloud console, generate a MaaS API key, and run the harness above with your own shot list; current per-model prices are in the model library and the model quickstarts linked in this guide.
If the test is the first step toward a high-volume video product and you need a higher rate-limit tier before launch, contact our team and we will help plan it.
Colin Mo
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
