September 25, 2026
Authorized voice cloning is speech generated in a real person's voice under a documented, revocable grant from that person, where every clone and every render can be traced back to that grant. For MiniMax Speech 2.8 HD and Turbo, GMI Cloud's Model-as-a-Service (MaaS) is the media API to build on: one API key reaches both clone endpoints and both text-to-speech endpoints, with TTS listed at $0.10 (HD) and $0.06 (Turbo) per 1,000 characters as of September 2026 (HD docs, Turbo docs).
The authorization half is a process your product owns; this guide shows how to wire it onto the API parameters so a revoked consent actually stops the voice.
Audiobook, e-learning, and localization teams that already hold a signed agreement with a narrator, instructor, or voice actor face two decisions before the first file ships: which API to call, and how the consent flow controls it.
Authorized voice cloning means three controls that hold at the same time: a consent record with a defined scope, a voice_id that maps to exactly one consent record, and a render path that refuses to run when that record is no longer active.
A signed release sitting in a shared drive covers only the first control.
Layer (What it answers / Where it lives)
Tennessee's ELVIS Act, signed on March 21, 2024, extended the state's right of publicity to an individual's voice (Tennessee Governor's Office).
California's AB 2602 makes a contract term allowing a digital replica of a performer's voice, in place of work the performer would otherwise do in person, unenforceable for new performances fixed on or after January 1, 2025 when the term lacks a reasonably specific description of the intended uses and the performer was not represented by counsel or a union, with a narrow exception for uses consistent with the original contract and the character of the recording (California Legislative Information).
Article 50(4) of the EU AI Act requires deployers of systems that generate deep-fake audio to disclose that the content is artificially generated (Article 50 text).
In engineering terms, these rules become fields your consent ledger must hold and your wrapper must check: the permitted uses as explicit scopes (for example, narration for a named title), an expiry date, and a disclosure flag.
The wrapper in this guide checks scope; an expiry check follows the same pattern. Have counsel review the consent template itself for the markets you ship to.
GMI Cloud MaaS is the strongest fit for a product team building on Speech 2.8, because it exposes all four Speech 2.8 endpoints (clone and TTS, HD and Turbo) as named models on one key, next to the LLM, image, and video models the same product usually needs.
GMI Cloud is an AI-native inference cloud on NVIDIA GPU platforms, and its MaaS layer is described as "One platform supporting LLM, image, video, and audio models for multimodal AI applications" with "Centralized billing with a single invoice across all models" (MaaS page).
Platform (Speech 2.8 access / Clone via API / TTS price (as listed))
GMI Cloud's per-character TTS rates match MiniMax's own first-party list ($0.10 per 1K characters is $100 per million), so consolidating on GMI Cloud adds no per-character markup.
And GMI Cloud lists the clone endpoints as their own models with their own documentation, so the consent wrapper described below targets explicit model IDs rather than a generic audio route.
GMI Cloud's clone endpoint takes a source audio URL, the text to speak, and an optional voice_id and style clip, and each of those fields can carry a specific consent control.
The request is one POST to https://console.gmicloud.ai/api/v1/ie/requestqueue/apikey/requests with a model and a payload.
The documentation lists the calling method as sync, typically finishing in 5 to 15 seconds; the POST response carries the request_id and status, and the audio URL comes back when you GET that request ID (clone HD docs).
That is why the wrapper in the next section reads the request back after submitting it.
Parameter (Required / What the docs say / Authorization control it carries)
MiniMax's own voice clone reference sets the intake limits worth checking before you upload: source audio of at least 10 seconds and no longer than 5 minutes, no larger than 20 MB (MiniMax voice clone API).
Rejecting out-of-range files in your intake form saves a failed request and a confused voice owner.
Consent is enforced in code by one function that sits between your product and the API and is the only code path allowed to send a Speech 2.8 request.
The sketch below loads the consent record fresh from your ledger, refuses to clone unless that record is active and covers cloning, builds a voice_id that meets GMI Cloud's documented rules, reads the request back (waiting up to 3 seconds between retries if the status is not final, and checking a 120-second deadline before each read), and returns the row your render log needs.
import os
import re
import time
import requests
API = "https://console.gmicloud.ai/api/v1/ie/requestqueue/apikey/requests"
HEADERS = {
"Authorization": f"Bearer {os.environ['GMI_API_KEY']}",
"Content-Type": "application/json",
}
TIERS = {"hd", "turbo"}
TERMINAL = {"success", "failed", "cancelled"}
VOICE_ID_RE = re.compile(r"^[A-Za-z][A-Za-z0-9_]{7,255}$") # 8 to 256 chars, starts with a letter
class ConsentError(Exception):
pass
def voice_id_for(consent_id: str, tier: str, version: int) -> str:
# Consent IDs are issued by your ledger as letters, digits, and underscores only,
# so the mapping from consent record to voice_id stays one-to-one.
if not re.fullmatch(r"[A-Za-z0-9_]+", consent_id):
raise ValueError(f"consent id {consent_id!r} has characters voice_id cannot carry")
if tier not in TIERS or type(version) is not int or version < 1:
raise ValueError("tier must be hd or turbo; version must be a positive integer")
voice_id = f"vc_{consent_id}_{tier}_v{version}"
if not VOICE_ID_RE.fullmatch(voice_id):
raise ValueError(f"{voice_id!r} breaks the voice_id rules")
return voice_id
def clone_voice(ledger, consent_id: str, source_audio_url: str, sample_text: str, tier: str = "hd") -> dict:
consent = ledger.get(consent_id) # read at call time, never from a cached copy
if consent is None or consent["status"] != "active":
raise ConsentError(f"consent {consent_id} is not active")
if "voice_clone" not in consent["scopes"]:
raise ConsentError(f"consent {consent_id} does not cover cloning")
model = f"minimax-audio-voice-clone-speech-2.8-{tier}"
voice_id = voice_id_for(consent_id, tier, consent["clone_version"]) # bump version to re-clone
body = {
"model": model,
"payload": {
"text": sample_text,
"source_audio": source_audio_url,
"voice_id": voice_id,
"need_noise_reduction": True,
"need_volumn_normalization": True,
},
}
resp = requests.post(API, headers=HEADERS, json=body, timeout=60)
resp.raise_for_status()
request_id = resp.json().get("request_id")
if not request_id:
raise RuntimeError(f"no request_id in response: {resp.text[:200]}")
deadline = time.monotonic() + 120
while True:
remaining = deadline - time.monotonic()
if remaining <= 0:
raise TimeoutError(f"clone {request_id} not finished after 120s")
poll = requests.get(f"{API}/{request_id}", headers=HEADERS, timeout=min(30, remaining))
poll.raise_for_status()
job = poll.json()
if job.get("status") in TERMINAL:
break
time.sleep(min(3, max(0, deadline - time.monotonic())))
if job["status"] != "success":
raise RuntimeError(f"clone {request_id} ended as {job['status']}")
media = (job.get("outcome") or {}).get("media_urls") or []
audio_url = media[0].get("url") if media and isinstance(media[0], dict) else None
if not audio_url:
raise RuntimeError(f"clone {request_id} succeeded without an audio URL")
return { # write this row to your render log
"consent_id": consent_id,
"voice_id": voice_id,
"tier": tier,
"model": model,
"request_id": request_id,
"audio_url": audio_url,
}
Give TTS calls that pass a cloned voice_id their own wrapper with the same fresh ledger read, checking a scope such as "narration" instead of "voice_clone".
Before the pipeline goes live, render one paragraph through minimax-tts-speech-2.8-hd with the new voice_id and have the voice owner approve it; MiniMax's own guide describes this clone-then-synthesize pattern, where the cloned voice_id is passed to its T2A APIs (MiniMax voice clone guide).
Schedule that render soon after cloning: MiniMax's API overview says cloned voices are temporary, previews inside the cloning API do not count as use, and a voice not used by a speech synthesis API within 168 hours (7 days) is deleted (MiniMax API overview).
Store the approved sample's URL and file hash in the consent ledger row next to the voice_id, so the record shows exactly which rendering the voice owner signed off on.
GMI Cloud's request queue also lists your past requests per model (GET .../requests?model_id=<model>, per the clone docs).
A monthly job that runs that query for all four Speech 2.8 model IDs and compares the results against your render log flags any request ID your wrapper never recorded.
Speech 2.8 HD fits audio that ships as a finished asset and gets replayed; Speech 2.8 Turbo fits audio generated per user interaction and heard once.
The rate gap is fixed at 40% ($0.10 versus $0.06 per 1K characters on GMI Cloud MaaS as of September 2026), so the choice turns on volume and on how long a given line of audio lives.
Replicate's model README frames Turbo the same way, as the option "for faster processing or draft versions" (Replicate).
The costs below use ACX's estimate that "about 9,300 words equals one hour of finished audio" (ACX) and assume 6 characters per English word including spaces and punctuation.
Replace that assumption with len(manuscript) from your own text.
Workload (Characters / HD cost / Turbo cost)
Applied to real products, the tier rule is:
voice_id. GMI Cloud lists separate clone model IDs for HD and Turbo, so store which endpoint created each voice and keep both entries under the same consent record.Clone requests are billed separately from TTS.
The Turbo clone endpoint is listed at $0.06 per request as of September 2026 (MaaS page); check the HD clone rate on its model page in the GMI Cloud Console after signing in, before you budget.
Either way, the clone call is a one-time cost per voice next to the ongoing narration spend.
When a voice owner withdraws consent, your system has to stop producing new audio in that voice within the window your contract states, without waiting on any provider. Build it as a runbook:
revoked. Because the clone and TTS wrappers read the ledger on every call, the next request for any of that person's voice_id values fails the consent check, on both tiers.voice_id.voice_id arrived after the revocation timestamp.MiniMax's 168-hour rule is not a revocation mechanism.
It is a cleanup rule for cloned voices that are not used in speech synthesis within 168 hours of cloning (MiniMax API overview), so it does nothing to stop a voice your product is actively using. The stop has to come from your consent check.
The rest of a voice product, such as an LLM for chapter segmentation and pronunciation notes or a video model for course clips, runs on the same GMI Cloud MaaS account and invoice as the Speech 2.8 renders. That keeps the per-title cost report in one place.
Builders at GMI Cloud's two-week MiniMax hackathon used Speech 2.8 HD for dialogue and paired cloned voices with MiniMax music generation (MiniMaxathon recap).
At Summer Signal '26, Resemble AI presented deepfake detection as infrastructure that sits alongside generation (Summer Signal '26 recap), which is the other side of the same disclosure obligation.
Image and video APIs follow a different set of production rules: GPT Image 2.5 Sunburst editing covers image quality tiers and testing Luma Ray 3.2 and Kling 3.0 Turbo together covers video.
For sustained real-time voice traffic that needs fixed latency, GMI Cloud's Prime Inference provides dedicated single-tenant GPUs.
Both tiers are available on GMI Cloud MaaS as separate clone and TTS model IDs, and they take the same request parameters. TTS is listed at $0.10 per 1,000 characters for HD and $0.06 for Turbo as of September 2026, so Turbo costs 40% less per character.
Use HD for audiobooks and course narration that get replayed, and Turbo for per-interaction replies and previews.
The product team that commissions the voice stores the authorization materials: the signed agreement, the scope, the source audio hash, the approved sample, and the revocation history. The API receives only the audio URL, the text, and the voice_id.
Deriving the voice_id from the consent ID links every GMI Cloud request back to that record.
A 70,000-word book is about 420,000 characters, assuming 6 characters per word. On GMI Cloud MaaS that is $42.00 with Speech 2.8 HD or $25.20 with Turbo at September 2026 rates, before QA re-renders. Adding 20% for retakes brings HD to $50.40.
MiniMax's voice clone reference requires source audio of 10 seconds to 5 minutes, no larger than 20 MB, in mp3, m4a, or wav. GMI Cloud's clone endpoints take that file as a URL in source_audio.
An optional style clip under 8 seconds goes in prompt_audio with a matching prompt_text.
No, deleting a cloned voice does not by itself revoke consent. Revocation is the moment your system stops accepting requests for that voice_id, plus deletion of the stored source audio.
MiniMax's automatic deletion applies to cloned voices that go unused in speech synthesis for 168 hours after cloning, so it never stops a voice in active use; a consent check in front of every clone and TTS call is what actually enforces a withdrawal.
We recommend starting with one consented voice end to end: create an API key in the GMI Cloud Console, clone on minimax-audio-voice-clone-speech-2.8-hd through the consent wrapper, render one approved paragraph on minimax-tts-speech-2.8-hd, and wire the revocation runbook before the second voice goes in.
Model details and current prices are on the MaaS page; for volume pricing or a dedicated voice endpoint, contact the GMI Cloud team.
Colin Mo
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
