You are going to produce a complete, ready-to-upload YouTube video end-to-end. Work in the current directory. Create a folder structure: /script, /audio, /visuals, /final. VIDEO BRIEF - Topic: [things hotel staff wont tell you] - Target length: 4 minutes (~6 words of script) - Style: Educational, conversational — hook, 3-act structure, rehook every 60-90 seconds - Aspect ratio: 16:9, 1080p - Tone: Confident, direct, no fluff, contractions OK STEP 1 — SCRIPT Write the full script as plain text in /script/script.txt. Use this structure: - Hook (15 seconds, statement or pattern interrupt) - Context preview (3 things theyll learn) - Body: 5 main points, each with rehook → story/example → payout - Outro with relink suggestion Also output /script/shot_list.txt that breaks the script into 25-40 numbered b-roll shots, each with: shot number, timestamp range, the script line it covers, and a visual description suitable for an AI video generator. Aim for 8-12 second clips. STEP 2 — VOICEOVER (ElevenLabs MCP) Read /script/script.txt. Split it into chunks of ~250 words. For each chunk, call the ElevenLabs MCP to generate audio using voice_id yr43K8H5LoTp6S1QFSGg. Save chunks as /audio/vo_001.mp3, /audio/vo_002.mp3, etc. Then concatenate them into /audio/voiceover_full.mp3 using FFmpeg. Report the total duration in seconds. STEP 3 — VISUALS (Higgsfield MCP) Read /script/shot_list.txt. For each shot: - If it's a static concept (charts, text overlays, title cards), use Higgsfield's image generation with Nano Banana Pro (best for text) or GPT Image 2, 1920x1080, save to /visuals/shot_XX.png - If it's motion footage and for the introduction which is the most important part of the video (people, objects, environments, action), use Seedance 2.0 via Higgsfield, 1920x1080, duration matching the shot length (8s default), save to /visuals/shot_XX.mp4 Run these in parallel where possible. Poll the job status until complete. Maintain a /visuals/manifest.json mapping shot number → file path → duration. STEP 4 — THUMBNAIL Generate a YouTube thumbnail using Higgsfield GPT Image 2 (it's the best at rendering text). 1280x720, high contrast, 3-5 word headline pulled from the hook. Save to /final/thumbnail.png. Generate 3 variations. STEP 5 — STITCH AND RENDER (FFmpeg) Use FFmpeg in bash to: 1. Build a concat list of all visual clips in shot order, padding any static images to their assigned duration 2. Overlay the /audio/voiceover_full.mp3 as the audio track 3. If total visual duration < audio duration, loop or extend the last shot to match 4. If total visual duration > audio duration, trim 5. Add a 0.5s crossfade between shots 6. Render to /final/video.mp4 at 1920x1080, H.264, AAC audio, YouTube-ready settings (recommended: -crf 18, -preset slow, -pix_fmt yuv420p) 7. Verify the output file plays and report duration + file size STEP 6 — METADATA Write /final/title_options.txt with 5 title options (curiosity-driven, under 60 chars). Write /final/description.txt with a 200-word description, keywords, and placeholder for affiliate links. Write /final/tags.txt with 20 relevant YouTube tags. CRITICAL RULES - Do NOT use any third-party brands, logos, or copyrighted characters in visuals - Polling: if a Higgsfield job is still processing, wait 15 seconds and re-check, up to 10 times before moving on - If any step fails, report the error and continue with the next step — do not silently skip - At the end, give me a final report: word count, voiceover duration, number of visual shots generated, final video duration, and the path to /final/video.mp4 Begin.