Describe an idea, the assistant sharpens the prompt, and 0G renders the images and video.
The assistant asks what you left out and explains why each answer changes the result. Then it renders through 0G Compute, over a route you can verify on-chain: stills on Z-Image-Turbo, motion on MiniMax H3 or Seedance 2.5.
DeepSeek V4 Flash runs a short interview instead of taking a prompt. One question at a time, always with the reason it matters, always with real options you can ignore. Five answers later you have a brief worth rendering, whether that is one still or a whole video.
A matte ceramic mug for a product page. Four angles.
What is it sitting on, and against?
WhyBackgrounds are where generated images fall apart. Naming the surface and the backdrop removes most of the mush.
Worn woodWarm, tactile, lived-in
MarbleCool, clean, premium
Seamless paperOne flat colour, no context
In its real placeA kitchen, a desk, a street
Worn oak. Warm, a bit beaten up.
How should it feel?
WhyMood ties the grade, the palette, and the styling together so a set of images reads as one set.
WarmInviting, lived-in, unhurried
ClinicalPrecise, cool, technical
PremiumQuiet, dark, expensive
PlayfulSaturated, bold, direct
Brief complete. Nothing renders until you approve.
A matte ceramic mug for a product page. Four angles.
What is it sitting on, and against? Backgrounds are where generated images fall apart. Naming the surface and the backdrop removes most of the mush. Answer: Worn wood.
How should it feel? Mood ties the grade, the palette, and the styling together so a set of images reads as one set. Answer: Warm.
Free-type anything instead. The options are shortcuts, not a menu.
Brief so far
subject
A product
framing
Product, centred
light
Window, from the left
surface
Worn oak
mood
Warm
Prompt for Z-Image-TurboReady
Matte ceramic mug, centred with room to breathe, soft morning window light from the left, long falloff across the body, visible glaze texture, 4:5. On worn oak against a neutral backdrop.Warm and unhurried.
Subject, framing, light, surface, mood. The answers compose into a single prompt, and the whole set renders from that one prompt instead of five unrelated ones.
01What is the subject, exactly?
02How close are we?
03Where is the light coming from?
04What is it sitting on, and against?
05How should it feel?
Four images from one brief · 624 credits · ~8s each
Generating a video
Five questions about a sequence
Subject, hook, setting, tone, payoff. The answers become a shot list. Every beat carries a frame prompt for z-image and a motion direction for whichever video model you picked.
01Who or what is on camera?
02What happens in the first second?
03Where are we?
04How should it feel?
05What do you want them to do at the end?
3 shots · 12s · 1,808 credits · MiniMax-H3
Nothing renders untilYou approve the briefYou see the credit costYou can still edit every line
Features
Built for a set.Not one lucky render.
A brief you can read, images and video that match each other, and a line item for every request that ran.
One reference, read into every prompt
z-image takes no image input, so a reference is read rather than pasted in. Qwen3-VL writes a precise description of it and folds that into the prompt behind every image in the brief.
01
02
03
04
Described, not copied. Subject, palette, framing and mood carry over. The pixels do not.
Guided prompting
Every question the assistant asks comes with the reason it matters, so the next brief you write is better than this one.
How close are we?
Framing decides what the image is about. A macro of one detail and a wide of the whole thing are two different arguments.
MacroProductPortrait
Approve the frame, then move it
A still is a finished deliverable on its own. When you want motion, that exact image becomes the locked first frame. You know what is moving before you pay for it.
z-image · 6 crbecomes the first frame of
MiniMax H3 · 6s · 180 cr
Brand kit
Brand name, palette, a default look and a default aspect ratio, saved to your profile once so you are not retyping them into every brief. Small on purpose: fonts, logo files and caption styling are not stored yet.
Queue variants of one brief and 0G Compute routes them upstream in parallel.
Ceramic mug · 4 anglesdone
Skincare range · flat lay64%
Shoe drop · 20s reveal22%
Cancel a run mid-queue. A request that fails is released, never charged.
A receipt for every request
0G's broker serves each request over a hardware-attested route. Utsuro writes the model, the work, the cost and that route to 0G Chain as a receipt you can open in the explorer.
Recent on-chain inference receipts
Tx
Model
Work
Cost
Route
0x9f3a…c41b
z-image-turbo
4 images · 4:5
24 cr
Attested
0x1d70…8ae2
minimax-h3
8s motion · 24fps
240 cr
Attested
0x4c92…07f5
deepseek-v4-flash
10k tokens · interview
2 cr
Attested
Settled per request. What a receipt proves is the cost and the route that served it. The model itself runs at the upstream, under that provider’s policy.
The stack
Four models.One decentralized network.
Utsuro does not host GPUs. Every model is reached through 0G Compute on Verified Routing: the route from 0G’s broker to the upstream model is hardware-attested and provable on-chain, and every request settles per call.
Image generation
Z-Image-Turbo
https://router-api.0g.ai/v1/z-image-turbo
Photoreal stills at speed. Generate finished images on their own, or use one as the locked first frame of a video so you know exactly what will move.
Cost
6 cr / image
Latency
~8s
Photoreal
Text-to-image
1024 × 1024
Two per request
Verified route
Verified Routing. The route to this model is attested on-chain
Video generation
MiniMax-H3
https://router-api.0g.ai/v1/minimax-h3
Motion with intent. Takes a prompt, or an image plus a direction, and renders footage with coherent camera work instead of drifting slideshow.
The first video model on the 0G Private Computer. Weights went public the same morning it landed here; Artificial Analysis ranked it #2 text-to-video and #3 image-to-video on day one.
Cost
30 cr / second
Latency
minutes
Text-to-video
Image-to-video
4–15s
2K
Verified route
Verified Routing. The route to this model is attested on-chain
The one that arrives with sound. Renders synchronised audio alongside the picture, holds a scene together for longer, and takes a still as its first frame like the other one does.
Billed by the vendor on tokens rather than seconds, so a clip's cost is not knowable until after it renders. Sold here per second at the 720p tier instead, which means the price on the button is the price, and a quote exists before you commit to it.
Cost
35 cr / second
Latency
minutes
Text-to-video
Image-to-video
4–15s
Renders audio
720p
Verified route
Verified Routing. The route to this model is attested on-chain
Creative assistant
DeepSeek V4 Flash
https://router-api.0g.ai/v1/deepseek-v4-flash
The part that makes the other two worth using. It asks what you left out, explains why it matters, and turns a half-formed idea into a prompt worth rendering.
The 0731 release is now the official one, swapped in under the same deepseek-v4-flash id. Terminal Bench 2.1 went from 61.8 to 82.7, on a 13B-active MoE.
Cost
2 cr / 10k tokens
Latency
~600ms
Guided prompting
Shot direction
Composition notes
Rewrites and variants
Streaming
Verified Routing. The route to this model is attested on-chain
Every model here runs on Verified Routing: the path from 0G's broker to the upstream is hardware-attested and provable on-chain. The inference itself runs at the upstream, under that provider's policy. What you can check is which route served your request.
0G Storage
Outputs you actually hold
Images and video are pinned and content-addressed, so an output has an address of its own rather than living in a bucket you cannot point at.
0G Chain
Costs you can audit
Each request settles on-chain at its listed rate. Every credit you spend has a receipt with a hash behind it, and no month-end number you have to take on trust.
0G Mainnet · chain 16661. These landed on 0G within a day of shipping, and you do not have to take that on faith: the created timestamps are public at router-api.0g.ai/v1/models. Credits settle on-chain behind the scenes. You never need a wallet to generate anything.
How it works
Four steps.One of them is talking.
One line is enough
~10 seconds
No format, no keyword soup, no negative prompts. Say what you want and whether you want it still or moving. If you do not say, the assistant asks.
Your idea
A matte ceramic mug for a product page. Four angles.
StillVideo
What you get
Every frame herewas made with Utsuro.
MiniMax H3
z-image
z-image
z-image
MiniMax H3
MiniMax H3
MiniMax H3
Stills and video from the models above. Three of the tiles below are the video files themselves, playing as they came out of MiniMax H3, not frames pulled from them. Each tile carries the model that made it and what it cost in credits.
Ceramic mug · three-quarterz-image6 cr
Mug to camera · counter light · 5sMiniMax H3150 cr
Costs are the listed rate: 6 credits an image, 30 credits a second of video. Every one of these settled on 0G Chain with a receipt attached.
Pricing
Every render is pricedbefore it runs.
Everyone starts on the same footing: 200 credits, every model, and the whole assistant. Credits are spent per request at the rates below, and the cost is on screen before you press generate.
Free
No card
200credits, once
33 images, or 6 seconds of video, or any mix of the two
The whole productEvery model and the full assistant. Nothing here is held back for a paid tier.
Granted once at signupNot a monthly pool. No card, no wallet, no trial clock counting down.
No watermark, commercial use includedWhat comes out is yours to publish, on the free grant like everywhere else.
Images and video. Product stills, flat lays, brand portraits, hero frames: z-image renders those on their own, and plenty of people never generate anything else. When you do want motion, two models take a prompt, or a still you already approved, and move it with real camera direction: MiniMax H3, or Seedance 2.5, which renders its own audio. All of it comes out of the same interview.
Do images and video cost the same?
No, and the gap is large. An image is 6 credits. A second of video is 30, so ten seconds of motion is 300. The assistant session costs a credit or two either way. A four-image product set lands around 26 credits total. A fifteen-second, three-shot video, first frames included, lands near 468.
Is there a subscription?
Not yet. No payment provider is wired up, so nothing in Utsuro can be bought today: no plans, no top-ups, no card field anywhere. Everyone signs up with the same 200 credits and spends them at the published rates. We would rather say that plainly than take pre-orders for a checkout that does not exist.
Can I keep the same face across a whole set?
Not reliably, and we are not going to pretend otherwise. Z-Image-Turbo takes no image input, so a reference cannot condition a still: it is read by a vision model and described into the prompt, which carries subject, palette, framing and mood across a set but not an identity. Saved characters with a locked face are not built. Where a reference does condition the render is video, because both video models accept a real first frame.
Who owns what I generate?
You do. Utsuro claims no rights over your briefs, your images, or your finished video. We do not train on your projects, and nothing you make appears anywhere public unless you put it there yourself.
Can I use the output commercially?
Yes. Commercial use is included, there is no separate licence tier, and there is no watermark on anything you generate. The usual limits still apply: no real person's likeness without their permission, no trademarks you do not own.
Why is a blockchain involved at all?
Three concrete reasons. 0G Compute reaches every model on Verified Routing: the route from 0G's broker to the upstream is hardware-attested and provable on-chain, so you can check which route served a given request rather than taking a vendor's word for it. What that does not mean: the inference itself runs at the upstream provider, under that provider's policy. The tier attests the path, not the enclave the model ran in, and we are not going to claim otherwise. 0G Storage pins your outputs so they are content-addressed and stay addressable rather than living in a bucket you cannot point at. 0G Chain settles each request at its listed rate, so your bill is a list of receipts instead of a number you have to trust. You never need a wallet.
What is actually new here?
There are two video models. MiniMax H3 was the first on the 0G Private Computer: text-to-video, or image-to-video from a first frame you have already approved, at 2K, 4 to 15 seconds, billed per generated second. MiniMax's weights went public the same morning it landed, and Artificial Analysis ranked it #2 text-to-video and #3 image-to-video on day one. Seedance 2.5 sits beside it now and renders synchronised audio with the picture, which MiniMax does not, at a slightly higher rate a second. Alongside it, DeepSeek V4 Flash 0731 is now the official release, swapped in under the same deepseek-v4-flash id, with Terminal Bench 2.1 up from 61.8 to 82.7 on a 13B-active MoE. 0G took all three of that morning's launches the same day; the created timestamps are public at router-api.0g.ai/v1/models.
What if the result is wrong?
Most of the time you catch it before it costs anything, because you approve the brief first. Rewrite the prompt, change the light, cut a beat, or send the assistant back with a note, and nothing is reserved until you say go. A render already in flight can be cancelled, and a request that fails releases its credits instead of spending them. You only pay for what actually ran.
Start free
Bring the idea.It will ask for the rest.
One sentence is enough to start. Ninety seconds later you have a brief you would have been proud to write, and images or video to match.