GMI MCP links your GMI Cloud account to Claude, Cursor, ChatGPT, Codex, and any AI tool that speaks MCP. Ask for an image, video, voiceover, or an answer from GMI's docs, and get it right where you're already working.
September 15, 2026
.png)
A landing page needs a hero image. A highlight reel needs a track under the b-roll. A keynote needs a ten-second title card. A tutorial script needs a voice reading it. A launch post needs five product shots in different settings. A support reply needs the exact name of a parameter buried three pages into a doc.
In every case, the media or the answer is one step inside something bigger: the project. The GMI MCP Server puts that step inside the tool where the bigger thing is already happening. Connect it to Claude, Cursor, ChatGPT, or Codex, and ask for what you need in the same sentence you'd use to describe it to a coworker.
You're working on | You ask for | What comes back |
A product landing page in Cursor | Three hero image options in the same style as the mockups already in the repo | Images sized for the hero section, dropped into your working directory |
A highlight reel in your editor | A 30 second instrumental track that matches the energy of the opening clip | An audio file sized to the clip, ready to drop under the b-roll |
A conference talk | A ten second animated title card with your talk name on screen | A short video clip you open your deck with |
A tutorial you're scripting | The script read back in a clear, neutral voice | A voiceover file to sync under your screen recording |
A blog post announcing a launch | Five variations of the same product shot in different settings | A set of images to place through the post |
A social clip for a release | A 15 second vertical video from a single product photo | A clip sized for a feed, with a slow camera move built in |
A support answer you're drafting | A check of GMI's own docs for the exact parameter name and its accepted values | The answer, with a link to the source page, in the same tab |
A monthly report to your team | This month's spend and which models it went to | A breakdown by model, pulled straight from your account |
Say you're editing a product demo and the last shot needs a five second clip of a still photo coming to life with a slow push in. You type the request into Claude Code the way you'd describe it to an editor: turn this photo into a five second video with a slow camera push in, and tell me what it costs before you run it.
Your assistant reads the model's parameters instead of guessing at them, since video models each name settings their own way. It checks the exact resolution and duration you asked for, quotes a price, and waits. You approve. Because a video job takes minutes rather than seconds, you get a tracking reference right away and a note confirming the job is actively running. Ask again in a few minutes and the finished clip comes back, ready to cut into the sequence you're already working in.
That five-step shape, describe it, read the parameters, price it, approve it, collect it, is the same whether you're asking for a video, an image, a voiceover, or a track. Only the parameters and the pricing unit change.
MCP, or Model Context Protocol, is a standard that lets AI assistants like Claude, ChatGPT, or Cursor connect directly to outside tools and services. Instead of an AI only knowing what you type, MCP gives it a way to reach out, check real data, and take real actions, like generating an image, checking an account balance, or looking something up, all through plain conversation instead of custom code.
We ran the same request, a man in a park with a dog, two ways: once by hand against GMI's raw API, once through the GMI MCP Server. The image that came back was equally strong on both paths. The difference showed up entirely in how many steps it took to get there.
Dimension | Raw API, by hand | GMI MCP Server |
Best caller | A backend or pipeline you control | Any assistant, in any chat |
Integration cost | Custom code written for each model's request shape | Handled by the server, ready on connect |
Model selection | You supply the exact model ID | Your assistant matches a model to the goal you describe |
Cost preview | Available once the job finishes | Quoted upfront, before anything runs |
Polling | A loop you write and maintain | Handled internally, you just ask again later |
Output shape | Varies by model, some return a URL, some return inline base64 data | One normalized shape, every time |
Discoverability | Found by searching documentation directly | Described in plain language, on request |
Both paths produce usable results. The raw API gives a backend full control over every parameter and retry. The GMI MCP Server gives an assistant the same access with the setup already done, which is the difference that matters most in the middle of building something else.
A different shape, same pattern. You've finished a five minute tutorial script and need it read aloud for the screen recording you're about to cut together. You ask your assistant what voice options exist, pick one that sounds clear and neutral, and ask for the script read back at that voice. Since text-to-speech is priced per character rather than per second, your assistant quotes the exact cost for the script length you gave it. You approve, the audio comes back, and you drop it under the recording, no separate narration tool or manual export step involved.
The pattern holds for a team checking in on spend, too. Someone on your team asks what the account spent on generation over the past week and which models it went to. The answer comes back broken down by model, over the exact window asked for, so a quick question during a standup gets answered on the spot.
Generating a piece of media usually means leaving the page you're writing, the deck you're building, or the edit you're cutting, opening a different tool, describing the same thing again in that tool's own interface, and bringing the result back yourself. The GMI MCP Server removes the middle step. You describe what you need once, where you already are, and your assistant handles picking a model, shaping the request, and returning a finished file or a link to one.
Step | Separate tool | Same chat, GMI MCP Server |
Describe what you need | Once in your project, again in the generation tool | Once |
Pick the right model for the job | You compare options yourself | Your assistant matches the request to a model that fits |
Check the price before it runs | Shown only after the job has already run | Quoted upfront, you approve before anything generates |
Track a longer job | You keep a tab open and wait | You ask again later, the job keeps running in the background |
Bring the result back into your project | Manual download and move | Lands where you're working |
The server covers four kinds of requests, and your assistant picks between them based on what you ask for:
Category | What you can get | Example ask |
Models | Browse the catalog by task, modality, or provider, then check a model's parameters and pricing before you commit | “Which image-to-video models can do 1080p, and what do they cost per second?” |
Images | Product shots, hero images, illustrations, variations on a style | "Give me three versions of this logo on a dark background" |
Video | Clips from a text description, or from a still photo brought to life | "Animate this photo with a slow zoom, five seconds, 720p" |
Audio | Voiceovers, narration, voice cloning, music from lyrics and a style prompt | "Read this script in a calm, neutral voice" |
Account and docs | Your balance, recent spend by model, answers pulled from GMI's documentation | "What did we spend on generation last week, broken down by model" |
Support | File a ticket with GMI's support team without leaving the conversation | “My last video generation failed, open a ticket with the request ID” |
Each generation request is priced and shown to you before it runs, so approving a job means you already know what it costs. Generating and charging both happen only after you approve.
More tools are on the way, expanding beyond images, video, and audio into new parts of the GMI Cloud platform. Each one will follow the same pattern: ask in plain language, see what it involves, and approve before anything runs.
The server lives at one address, and every supported client points at the same one: Claude, Claude Code, Codex, ChatGPT, and Cursor all connect the same way, through a one-time browser sign-in. Sign-in happens once through the browser, and one address covers setup for every client. Whichever tool you're already writing in, the same set of requests works the same way inside it.
This fits the team you already have, generation as a side task rather than a dedicated job. A two person team shipping a landing page, a solo developer recording a course, a marketer cutting a highlight reel between meetings, all of these are the same shape of request: something is almost done, and one piece of media or one quick answer is what's left. The GMI MCP Server is built around that moment rather than around a production workflow with its own tooling and its own login.
That also means the same server scales down cleanly. A single request costs a single request. The first image or clip comes back on the same login and plan you already have.
Point your client (Claude, Claude Code, Codex, ChatGPT, or Cursor) at GMI MCP, sign in once through your browser, and you're connected.
In Claude Code that's one line in the terminal; in Cursor, Claude Desktop, or any JSON-configured client, it's one block in the MCP settings file.
claude mcp add --transport http gmi https://mcp.gmicloud.ai/mcp{
"mcpServers": {
"gmi": {
"url": "https://mcp.gmicloud.ai/mcp"
}
}
}Full setup steps for each client are in the GMI MCP Server docs.
Once it's connected, ask for the next piece your project needs.
Roan Weigert
DevRel Lead @ GMI Cloud
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
