She's not dubbed. @MiniMax_AI H3 generated her voice and her face in the same pass, from a single prompt.
And the number she says is real. Five dollars of free credit is worth almost four minutes of finished video with sound, and it lands in your account just for signing up.
H3 is live on deAPI.
GPT Image 1.5 and GPT Image 2 are now on deAPI. Same endpoint you already call for FLUX or Z-Image, so switching is a one-string edit.
Both are Experimental and need a paid account.
Features get cut when someone opens the invoice.
That's how voiceover sits behind a "coming soon" badge for eight months, and how transcription quietly turns into a Pro-tier upsell.
A 1024×1024 image from FLUX.1 Schnell costs $0.00268. An hour of Whisper transcription runs $0.052. Ten thousand images a month is $27, and nobody calls a roadmap meeting about $27.
$5 in credits if you want to check the math.
New on the blog: GPT Image 1.5 and GPT Image 2 are live on deAPI.
The post covers what silently stops working when you point existing code at them, what a single image costs at each quality setting, and which of the two fits the resolution you need.
https://t.co/N11uFTLObA
New on the blog: a prompting guide for MiniMax H3 image-to-video.
It covers the line your prompt has to open with, why describing your input image is the fastest way to get a clip that doesn't move, and three prompts you can paste straight into the playground.
https://t.co/T6X3ITKVOu
A talking clip used to take three calls: a video model, a TTS model, and a sync step that never quite landed.
@MiniMax_AI H3 returns both signals from one denoising pass, so the lip movement is already correct in the file you get back. It's live on deAPI.
Learn more at https://t.co/VAnVUh54LK
An interview transcript that doesn't say who was talking is half a transcript.
Whisper Large V3 CT2 is live on deAPI. It labels the speakers and times every word, and returns the whole thing as JSON.
https://t.co/HQ4596SSHR
A soap bubble is two hard problems at once. The surface has to throw iridescence without turning into an oil slick, and the film has to read as something a needle would destroy on contact.
We gave that prompt to FLUX.1 Schnell, FLUX.2 Klein and Z-Image-Turbo. Each one decided for itself how close the bubble gets to the points.
All three sit side by side in Playground - feel free to test before you ship.
Write "cinematic and emotional" in a MiniMax H3 prompt and the model guesses. Write "a sustained low cello note held under the dialogue" and it builds exactly that.
H3 was trained on structured documents with named fields and shot markers. Mood words do nothing. Observable details do everything.
Full prompting guide on our blog.
https://t.co/8YpSVAmqFV
Same prompt - a magnifying glass on fabric. Three models, three takes on how light bends through glass.
FLUX.2 Klein, Z-Image-Turbo, FLUX.1 schnell. Try each one in Playground, zero code.
https://t.co/w8RsilCdwW
Every image editing tool assumes you know where the object is. Draw a mask, set coordinates, paint over the area.
Instruct-edit models flip that. You say what to change, the model figures out where it is.
Qwen Image Edit Plus is the best open-source one we've tested. We just published a prompting guide - the mental model is different enough from regular image gen that it's worth reading before your first API call:
https://t.co/BPlHaXkHao
Jensen Huang's first post on X is a letter signed by 50 companies - Nvidia, Microsoft, Meta, AMD, Google among them. The ask: make open weights the industry standard.
Two years ago, open-source AI was "good enough for hobbyists." Now it's a policy position backed by companies worth $10 trillion combined.
deAPI runs entirely on open-source models. Seeing the industry's biggest players rally behind the same principle is a good sign we picked the right foundation.
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
In blind listening tests, 63.75% of evaluators preferred Chatterbox over ElevenLabs. It's open-source, MIT licensed, and runs 22 languages.
We just published a prompting guide for it - how to shape delivery when text is your only control surface: https://t.co/4LS62eOhkA