画面訳写・鏡語訳写・声意訳写の使い方とよくある質問
最終更新:2026年8月29日
Upload a reference image to unpack light, framing, and style into model-ready prompts.
JPG / PNG / WEBP / GIF, max 2MB. Prefer 512–4096px; longest side ≤8192px. Uploads are deleted after processing.
General descriptions, structured terms, and native Flux / Midjourney / Stable Diffusion formats you can paste.
It softens plastic AI look toward more photographic language. Some features require sign-in.
Upload a short clip, auto-cut shots, then reverse image and video prompts for kept frames.
Common formats (MP4 / MOV / WEBM / MKV), ≤1 minute, ≤30MB. Standard / Fine / Coarse modes all cap at 10 shots—Standard balances cuts, Fine is more sensitive, Coarse keeps only strong cuts. Uploads are deleted after processing.
First cut and caption shots, then generate dual prompts for remaining shots; delete unused frames anytime.
Turn text into speech via emotion shaping, cloning, or presets.
Clean speech without BGM. 3–20s is best; longer audio is auto-trimmed. Max 10MB. Uploads are deleted after the job.
Emotion shaping layers vectors on a clone; clone uses your sample; presets play from the library with no upload.
MP3 / FLAC / OPUS with pace about 0.5x–2.0x.
Sign-in, quotas, and data handling.
Core capabilities are free for a limited time after you register. Quotas follow on-page rules; paid plans apply when the promotion ends.
Browsing may be open, but submitting jobs and history usually need sign-in. Free quotas follow the plan pages.
History is short-lived (e.g. about 3 days—see in-app notice). Copy or download important results promptly.
By default we do not use your uploads to train unrelated models. Originals are deleted after processing. See the Privacy Policy.