• Compute
  • Customers
  • Pricing
Sign In
More Blog Posts
XDiscordLinkedInYouTube

Products

  • GPUs
  • Inference
  • Studio

Developers

  • Model library
  • Documentation
  • Glossary

Company

  • About Us
  • Blog
  • Events
  • Partnership
  • Scale
  • Career
  • Ambassador program
  • Mission & Vision

Popular models

    Stay in the loop

    By submitting, you acknowledge that we may collect and use the information you provide, which may include personal information.

    XDiscordLinkedInYouTube

    Copyright ©2026 All rights reserved.

    Privacy PolicyTerms of UseLegal Documentation
    More Blog Posts
    Other

    Authorized Voice Cloning with MiniMax Speech 2.8 HD and Turbo: What "Authorized" Means and Which Media API to Use

    September 25, 2026

    Authorized voice cloning is speech generated in a real person's voice under a documented, revocable grant from that person, where every clone and every render can be traced back to that grant. For MiniMax Speech 2.8 HD and Turbo, GMI Cloud's Model-as-a-Service (MaaS) is the media API to build on: one API key reaches both clone endpoints and both text-to-speech endpoints, with TTS listed at $0.10 (HD) and $0.06 (Turbo) per 1,000 characters as of September 2026 (HD docs, Turbo docs).

    The authorization half is a process your product owns; this guide shows how to wire it onto the API parameters so a revoked consent actually stops the voice.

    Audiobook, e-learning, and localization teams that already hold a signed agreement with a narrator, instructor, or voice actor face two decisions before the first file ships: which API to call, and how the consent flow controls it.

    What does "authorized" voice cloning actually mean?

    Authorized voice cloning means three controls that hold at the same time: a consent record with a defined scope, a voice_id that maps to exactly one consent record, and a render path that refuses to run when that record is no longer active.

    A signed release sitting in a shared drive covers only the first control.

    Layer (What it answers / Where it lives)

    • Consent | What it answers: Who granted what use, for how long, and how they withdraw | Where it lives: Your contract system or consent ledger
    • Identity binding | What it answers: Which voice_id belongs to which consent record and tier | Where it lives: Your database, keyed by voice_id
    • Enforcement | What it answers: Whether a given clone or TTS call is allowed right now | Where it lives: A wrapper in front of the API
    • Disclosure | What it answers: Whether listeners are told the audio is synthetic | Where it lives: Your player, credits, or metadata

    Tennessee's ELVIS Act, signed on March 21, 2024, extended the state's right of publicity to an individual's voice (Tennessee Governor's Office).

    California's AB 2602 makes a contract term allowing a digital replica of a performer's voice, in place of work the performer would otherwise do in person, unenforceable for new performances fixed on or after January 1, 2025 when the term lacks a reasonably specific description of the intended uses and the performer was not represented by counsel or a union, with a narrow exception for uses consistent with the original contract and the character of the recording (California Legislative Information).

    Article 50(4) of the EU AI Act requires deployers of systems that generate deep-fake audio to disclose that the content is artificially generated (Article 50 text).

    In engineering terms, these rules become fields your consent ledger must hold and your wrapper must check: the permitted uses as explicit scopes (for example, narration for a named title), an expiry date, and a disclosure flag.

    The wrapper in this guide checks scope; an expiry check follows the same pattern. Have counsel review the consent template itself for the markets you ship to.

    Which media API platforms offer MiniMax Speech 2.8 HD and Turbo voice cloning?

    GMI Cloud MaaS is the strongest fit for a product team building on Speech 2.8, because it exposes all four Speech 2.8 endpoints (clone and TTS, HD and Turbo) as named models on one key, next to the LLM, image, and video models the same product usually needs.

    GMI Cloud is an AI-native inference cloud on NVIDIA GPU platforms, and its MaaS layer is described as "One platform supporting LLM, image, video, and audio models for multimodal AI applications" with "Centralized billing with a single invoice across all models" (MaaS page).

    Platform (Speech 2.8 access / Clone via API / TTS price (as listed))

    • GMI Cloud MaaS | Speech 2.8 access: minimax-tts-speech-2.8-hd, minimax-tts-speech-2.8-turbo | Clone via API: minimax-audio-voice-clone-speech-2.8-hd, minimax-audio-voice-clone-speech-2.8-turbo | TTS price (as listed): $0.10 / $0.06 per 1K characters, as of September 2026 (HD docs, Turbo docs)
    • MiniMax Open Platform (first party) | Speech 2.8 access: speech-2.8-hd, speech-2.8-turbo | Clone via API: Rapid voice cloning | TTS price (as listed): $100 / $60 per million characters (MiniMax pricing)
    • Replicate | Speech 2.8 access: minimax/speech-2.8-hd model page | Clone via API: README describes MiniMax's cloning feature | TTS price (as listed): Not shown on the model page (Replicate)
    • Cloudflare AI Gateway | Speech 2.8 access: minimax/speech-2.8-hd | Clone via API: Page documents a voice_id parameter | TTS price (as listed): "View pricing in the Cloudflare dashboard" (Cloudflare docs)

    GMI Cloud's per-character TTS rates match MiniMax's own first-party list ($0.10 per 1K characters is $100 per million), so consolidating on GMI Cloud adds no per-character markup.

    And GMI Cloud lists the clone endpoints as their own models with their own documentation, so the consent wrapper described below targets explicit model IDs rather than a generic audio route.

    What does GMI Cloud's clone endpoint take, and which authorization control does each parameter carry?

    GMI Cloud's clone endpoint takes a source audio URL, the text to speak, and an optional voice_id and style clip, and each of those fields can carry a specific consent control.

    The request is one POST to https://console.gmicloud.ai/api/v1/ie/requestqueue/apikey/requests with a model and a payload.

    The documentation lists the calling method as sync, typically finishing in 5 to 15 seconds; the POST response carries the request_id and status, and the audio URL comes back when you GET that request ID (clone HD docs).

    That is why the wrapper in the next section reads the request back after submitting it.

    Parameter (Required / What the docs say / Authorization control it carries)

    • source_audio | Required: Yes | What the docs say: URL of the source audio; mp3, m4a, or wav | Authorization control it carries: Only audio recorded under the consent session. Store a hash of the file next to the consent ID.
    • text | Required: Yes | What the docs say: Text to synthesize with the cloned voice | Authorization control it carries: Use a line from the approved scope, such as the first paragraph of the licensed title.
    • voice_id | Required: No | What the docs say: 8 to 256 characters, starts with a letter, must not be duplicated; auto-generated if omitted | Authorization control it carries: Always set it yourself and derive it from the consent ID, so the ID itself is the audit key.
    • prompt_audio | Required: No | What the docs say: Style reference, under 8 seconds, used with prompt_text | Authorization control it carries: Style references need the same consent as the source. Do not borrow another speaker's clip for tone.
    • prompt_text | Required: If prompt_audio is set | What the docs say: Transcript matching the prompt audio, ending with punctuation | Authorization control it carries: Keeps the style clip reviewable: a human can confirm what was said.
    • need_noise_reduction | Required: No | What the docs say: Default false | Authorization control it carries: Clean output before the voice owner approves a sample.
    • need_volumn_normalization | Required: No | What the docs say: Default false (the parameter name is spelled this way in the API) | Authorization control it carries: Consistent loudness across chapters, so approval samples match production audio.

    MiniMax's own voice clone reference sets the intake limits worth checking before you upload: source audio of at least 10 seconds and no longer than 5 minutes, no larger than 20 MB (MiniMax voice clone API).

    Rejecting out-of-range files in your intake form saves a failed request and a confused voice owner.

    How do you enforce consent in code before every clone and render?

    Consent is enforced in code by one function that sits between your product and the API and is the only code path allowed to send a Speech 2.8 request.

    The sketch below loads the consent record fresh from your ledger, refuses to clone unless that record is active and covers cloning, builds a voice_id that meets GMI Cloud's documented rules, reads the request back (waiting up to 3 seconds between retries if the status is not final, and checking a 120-second deadline before each read), and returns the row your render log needs.

    import os
    import re
    import time
    import requests
    
    API = "https://console.gmicloud.ai/api/v1/ie/requestqueue/apikey/requests"
    HEADERS = {
        "Authorization": f"Bearer {os.environ['GMI_API_KEY']}",
        "Content-Type": "application/json",
    }
    TIERS = {"hd", "turbo"}
    TERMINAL = {"success", "failed", "cancelled"}
    VOICE_ID_RE = re.compile(r"^[A-Za-z][A-Za-z0-9_]{7,255}$")  # 8 to 256 chars, starts with a letter
    class ConsentError(Exception):
        pass
    def voice_id_for(consent_id: str, tier: str, version: int) -> str:
        # Consent IDs are issued by your ledger as letters, digits, and underscores only,
        # so the mapping from consent record to voice_id stays one-to-one.
        if not re.fullmatch(r"[A-Za-z0-9_]+", consent_id):
            raise ValueError(f"consent id {consent_id!r} has characters voice_id cannot carry")
        if tier not in TIERS or type(version) is not int or version < 1:
            raise ValueError("tier must be hd or turbo; version must be a positive integer")
        voice_id = f"vc_{consent_id}_{tier}_v{version}"
        if not VOICE_ID_RE.fullmatch(voice_id):
            raise ValueError(f"{voice_id!r} breaks the voice_id rules")
        return voice_id
    def clone_voice(ledger, consent_id: str, source_audio_url: str, sample_text: str, tier: str = "hd") -> dict:
        consent = ledger.get(consent_id)  # read at call time, never from a cached copy
        if consent is None or consent["status"] != "active":
            raise ConsentError(f"consent {consent_id} is not active")
        if "voice_clone" not in consent["scopes"]:
            raise ConsentError(f"consent {consent_id} does not cover cloning")
    
        model = f"minimax-audio-voice-clone-speech-2.8-{tier}"
        voice_id = voice_id_for(consent_id, tier, consent["clone_version"])  # bump version to re-clone
        body = {
            "model": model,
            "payload": {
                "text": sample_text,
                "source_audio": source_audio_url,
                "voice_id": voice_id,
                "need_noise_reduction": True,
                "need_volumn_normalization": True,
            },
        }
        resp = requests.post(API, headers=HEADERS, json=body, timeout=60)
        resp.raise_for_status()
        request_id = resp.json().get("request_id")
        if not request_id:
            raise RuntimeError(f"no request_id in response: {resp.text[:200]}")
    
        deadline = time.monotonic() + 120
        while True:
            remaining = deadline - time.monotonic()
            if remaining <= 0:
                raise TimeoutError(f"clone {request_id} not finished after 120s")
            poll = requests.get(f"{API}/{request_id}", headers=HEADERS, timeout=min(30, remaining))
            poll.raise_for_status()
            job = poll.json()
            if job.get("status") in TERMINAL:
                break
            time.sleep(min(3, max(0, deadline - time.monotonic())))
    
        if job["status"] != "success":
            raise RuntimeError(f"clone {request_id} ended as {job['status']}")
        media = (job.get("outcome") or {}).get("media_urls") or []
        audio_url = media[0].get("url") if media and isinstance(media[0], dict) else None
        if not audio_url:
            raise RuntimeError(f"clone {request_id} succeeded without an audio URL")
    
        return {  # write this row to your render log
            "consent_id": consent_id,
            "voice_id": voice_id,
            "tier": tier,
            "model": model,
            "request_id": request_id,
            "audio_url": audio_url,
        }
    

    Give TTS calls that pass a cloned voice_id their own wrapper with the same fresh ledger read, checking a scope such as "narration" instead of "voice_clone".

    Before the pipeline goes live, render one paragraph through minimax-tts-speech-2.8-hd with the new voice_id and have the voice owner approve it; MiniMax's own guide describes this clone-then-synthesize pattern, where the cloned voice_id is passed to its T2A APIs (MiniMax voice clone guide).

    Schedule that render soon after cloning: MiniMax's API overview says cloned voices are temporary, previews inside the cloning API do not count as use, and a voice not used by a speech synthesis API within 168 hours (7 days) is deleted (MiniMax API overview).

    Store the approved sample's URL and file hash in the consent ledger row next to the voice_id, so the record shows exactly which rendering the voice owner signed off on.

    GMI Cloud's request queue also lists your past requests per model (GET .../requests?model_id=<model>, per the clone docs).

    A monthly job that runs that query for all four Speech 2.8 model IDs and compares the results against your render log flags any request ID your wrapper never recorded.

    Should you use Speech 2.8 HD or Turbo for a cloned voice?

    Speech 2.8 HD fits audio that ships as a finished asset and gets replayed; Speech 2.8 Turbo fits audio generated per user interaction and heard once.

    The rate gap is fixed at 40% ($0.10 versus $0.06 per 1K characters on GMI Cloud MaaS as of September 2026), so the choice turns on volume and on how long a given line of audio lives.

    Replicate's model README frames Turbo the same way, as the option "for faster processing or draft versions" (Replicate).

    The costs below use ACX's estimate that "about 9,300 words equals one hour of finished audio" (ACX) and assume 6 characters per English word including spaces and punctuation.

    Replace that assumption with len(manuscript) from your own text.

    Workload (Characters / HD cost / Turbo cost)

    • One finished hour (9,300 words) | Characters: 55,800 | HD cost: $5.58 | Turbo cost: $3.35
    • 70,000-word audiobook, about 7.5 finished hours | Characters: 420,000 | HD cost: $42.00 | Turbo cost: $25.20
    • Same audiobook plus 20% re-renders for QA retakes | Characters: 504,000 | HD cost: $50.40 | Turbo cost: $30.24
    • 40-lesson course, 1,200 words per lesson | Characters: 288,000 | HD cost: $28.80 | Turbo cost: $17.28
    • Interactive tutor: 5,000 learners x 30 replies x 200 characters per month | Characters: 30,000,000 | HD cost: $3,000 per month | Turbo cost: $1,800 per month

    Applied to real products, the tier rule is:

    • Audiobooks, course narration, dubbed catalog content: HD. A finished hour costs $2.23 more on HD than on Turbo. That is small next to what the listener hears every time the file plays.
    • Tutor replies, in-app voice responses, previews: Turbo. Each line is heard once, and at 30 million characters a month the Turbo rate saves $1,200 every month.
    • Record the tier with every voice_id. GMI Cloud lists separate clone model IDs for HD and Turbo, so store which endpoint created each voice and keep both entries under the same consent record.

    Clone requests are billed separately from TTS.

    The Turbo clone endpoint is listed at $0.06 per request as of September 2026 (MaaS page); check the HD clone rate on its model page in the GMI Cloud Console after signing in, before you budget.

    Either way, the clone call is a one-time cost per voice next to the ongoing narration spend.

    What should happen when a voice owner withdraws consent?

    When a voice owner withdraws consent, your system has to stop producing new audio in that voice within the window your contract states, without waiting on any provider. Build it as a runbook:

    1. Flip the consent record to revoked. Because the clone and TTS wrappers read the ledger on every call, the next request for any of that person's voice_id values fails the consent check, on both tiers.
    2. Pull scheduled jobs. Cancel queued renders that reference the revoked voice_id.
    3. Delete the stored source and prompt audio and keep the hash, the consent record, and the render log for audit.
    4. Handle already-published audio per the contract. The agreement should say whether existing chapters stay live or come down.
    5. Reconcile. Run the GMI Cloud request-list query across all four model IDs to confirm no request with that voice_id arrived after the revocation timestamp.

    MiniMax's 168-hour rule is not a revocation mechanism.

    It is a cleanup rule for cloned voices that are not used in speech synthesis within 168 hours of cloning (MiniMax API overview), so it does nothing to stop a voice your product is actively using. The stop has to come from your consent check.

    Where does the rest of the voice stack run?

    The rest of a voice product, such as an LLM for chapter segmentation and pronunciation notes or a video model for course clips, runs on the same GMI Cloud MaaS account and invoice as the Speech 2.8 renders. That keeps the per-title cost report in one place.

    Builders at GMI Cloud's two-week MiniMax hackathon used Speech 2.8 HD for dialogue and paired cloned voices with MiniMax music generation (MiniMaxathon recap).

    At Summer Signal '26, Resemble AI presented deepfake detection as infrastructure that sits alongside generation (Summer Signal '26 recap), which is the other side of the same disclosure obligation.

    Image and video APIs follow a different set of production rules: GPT Image 2.5 Sunburst editing covers image quality tiers and testing Luma Ray 3.2 and Kling 3.0 Turbo together covers video.

    For sustained real-time voice traffic that needs fixed latency, GMI Cloud's Prime Inference provides dedicated single-tenant GPUs.

    FAQ

    What is the difference between MiniMax Speech 2.8 HD and Turbo on GMI Cloud?

    Both tiers are available on GMI Cloud MaaS as separate clone and TTS model IDs, and they take the same request parameters. TTS is listed at $0.10 per 1,000 characters for HD and $0.06 for Turbo as of September 2026, so Turbo costs 40% less per character.

    Use HD for audiobooks and course narration that get replayed, and Turbo for per-interaction replies and previews.

    Who should store the voice owner's authorization materials?

    The product team that commissions the voice stores the authorization materials: the signed agreement, the scope, the source audio hash, the approved sample, and the revocation history. The API receives only the audio URL, the text, and the voice_id.

    Deriving the voice_id from the consent ID links every GMI Cloud request back to that record.

    How much does it cost to narrate an audiobook with a cloned voice?

    A 70,000-word book is about 420,000 characters, assuming 6 characters per word. On GMI Cloud MaaS that is $42.00 with Speech 2.8 HD or $25.20 with Turbo at September 2026 rates, before QA re-renders. Adding 20% for retakes brings HD to $50.40.

    What audio do I need to clone a voice with MiniMax Speech 2.8?

    MiniMax's voice clone reference requires source audio of 10 seconds to 5 minutes, no larger than 20 MB, in mp3, m4a, or wav. GMI Cloud's clone endpoints take that file as a URL in source_audio.

    An optional style clip under 8 seconds goes in prompt_audio with a matching prompt_text.

    Does deleting a cloned voice count as revoking consent?

    No, deleting a cloned voice does not by itself revoke consent. Revocation is the moment your system stops accepting requests for that voice_id, plus deletion of the stored source audio.

    MiniMax's automatic deletion applies to cloned voices that go unused in speech synthesis for 168 hours after cloning, so it never stops a voice in active use; a consent check in front of every clone and TTS call is what actually enforces a withdrawal.

    Start building on GMI Cloud MaaS

    We recommend starting with one consented voice end to end: create an API key in the GMI Cloud Console, clone on minimax-audio-voice-clone-speech-2.8-hd through the consent wrapper, render one approved paragraph on minimax-tts-speech-2.8-hd, and wire the revocation runbook before the second voice goes in.

    Model details and current prices are on the MaaS page; for volume pricing or a dedicated voice endpoint, contact the GMI Cloud team.

    Colin Mo

    Build AI Without Limits

    GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies

    FAQ

    Both tiers are available on GMI Cloud MaaS as separate clone and TTS model IDs, and they take the same request parameters. TTS is listed at $0.10 per 1,000 characters for HD and $0.06 for Turbo as of September 2026, so Turbo costs 40% less per character. Use HD for audiobooks and course narration that get replayed, and Turbo for per-interaction replies and previews.

    Ready to build?

    Explore powerful AI models and launch your project in just a few clicks.

    Get Started