MULTIMODAL AI WORKFLOWS: CHAINING TEXT, IMAGE, VOICE AND VIDEO INTO ONE PIPELI

Handoff contracts across the multimodal pipeline

Text→imageJPEG/PNG/WebP exact pxMatch video hop's resolutionN/ASend exact width/heightYes
Image→videoMP4, 16–20s (max 120s)sora-2/-pro, ≤6 extensions1 hour (download URL)Sora ends Sep 24, 2026Yes
Text→speechPCM/WAV 44.1–48kHzPro tier+ (Creator+ for mp3)N/AFree tier: lossy MP3 128kbpsNo
Audio→transcriptmp3/wav/m4a ≤25MBchunking_strategy=autoN/AAssemblyAI: 5GB/10h filesYes
Transcript→clipsSRT/VTT + timestampswhisper-1 onlyN/AKeep whisper-1 for captionsYes
Render→publishMP4 moov-front, -16 LKFSYouTube & Apple specsN/APrecondition loudness firstYes

Expiry is N/A except Sora, whose download URLs lapse in 1 hour; 'Free tier?' flags contracts a no-cost plan cannot satisfy.