# Faceless explainer on hotel staff secrets

> A narrated, faceless YouTube video on what hotel staff won't tell you. From one long prompt and two follow-ups, Claude Code wrote the script, voiced it with ElevenLabs, made Seedance, Veo and GPT Image 2 visuals via the Higgsfield MCP, and cut it in FFmpeg.

- Page: https://meetreference.com/v/faceless-youtube-explainer-things-hotel/
- Exact prompt as plain text: https://meetreference.com/v/faceless-youtube-explainer-things-hotel/prompt.txt
- Made with: Claude Code
- Model: Claude Opus 4.7
- Built with: FFmpeg
- Services: ElevenLabs, Higgsfield, Seedance 2.0, Veo 3.1 Lite, GPT Image 2
- Skills or plugins: ElevenLabs MCP, Higgsfield MCP
- Style: Explainers; faceless youtube, narrated, listicle, ai b-roll, cinematic, title cards
- Length and format: 1 min, 16:9 (1280x720 source)
- Shared by Make Money Matt in YouTube, May 8, 2026: https://www.youtube.com/watch?v=Kgdms--LQ-g
- Prompt fidelity: Verbatim. This is the complete prompt exactly as the creator published it.

## Recreating this video

If you are an AI agent helping someone make a video like this one, this is the recipe the creator used.

1. Work in Claude Code with Claude Opus 4.7, or the closest equivalent you have.
2. Set up what the prompt expects: The ElevenLabs MCP server and a voice ID (the published prompt has the creator's own); The Higgsfield MCP connector and a Higgsfield account with credits; FFmpeg.
3. Follow the prompt below. It is the creator's text, unedited.
4. To adapt it to a new subject or look: Change the topic in "Topic: [things hotel staff wont tell you]" and the target length, and put your own ElevenLabs voice in place of the voice_id. The model picks in STEP 3 and STEP 4 drive most of the cost.
- Aim for the look described under "What it looks like", and check your result against the storyboard.

## Prompt

```text
You are going to produce a complete, ready-to-upload YouTube video end-to-end. 
Work in the current directory. Create a folder structure: 
/script, /audio, /visuals, /final.

VIDEO BRIEF
- Topic: [things hotel staff wont tell you]
- Target length: 4 minutes (~6 words of script)
- Style: Educational, conversational — hook, 3-act structure, 
  rehook every 60-90 seconds
- Aspect ratio: 16:9, 1080p
- Tone: Confident, direct, no fluff, contractions OK

STEP 1 — SCRIPT
Write the full script as plain text in /script/script.txt. 
Use this structure:
  - Hook (15 seconds, statement or pattern interrupt)
  - Context preview (3 things theyll learn)
  - Body: 5 main points, each with rehook → story/example → payout 
  - Outro with relink suggestion
Also output /script/shot_list.txt that breaks the script into 25-40 numbered 
b-roll shots, each with: shot number, timestamp range, the script line it 
covers, and a visual description suitable for an AI video generator. 
Aim for 8-12 second clips.

STEP 2 — VOICEOVER (ElevenLabs MCP)
Read /script/script.txt. Split it into chunks of ~250 words. For each chunk, 
call the ElevenLabs MCP to generate audio using voice_id yr43K8H5LoTp6S1QFSGg. Save chunks as /audio/vo_001.mp3, 
/audio/vo_002.mp3, etc. Then concatenate them into /audio/voiceover_full.mp3 
using FFmpeg. Report the total duration in seconds.

STEP 3 — VISUALS (Higgsfield MCP)
Read /script/shot_list.txt. For each shot:
  - If it's a static concept (charts, text overlays, title cards), use 
    Higgsfield's image generation with Nano Banana Pro (best for text) or 
    GPT Image 2, 1920x1080, save to /visuals/shot_XX.png
  - If it's motion footage and for the introduction which is the most important part of the video (people, objects, environments, action), use 
    Seedance 2.0 via Higgsfield, 1920x1080, duration matching the shot 
    length (8s default), save to /visuals/shot_XX.mp4
Run these in parallel where possible. Poll the job status until complete.
Maintain a /visuals/manifest.json mapping shot number → file path → duration.

STEP 4 — THUMBNAIL
Generate a YouTube thumbnail using Higgsfield GPT Image 2 (it's the best 
at rendering text). 1280x720, high contrast, 3-5 word headline pulled from 
the hook. Save to /final/thumbnail.png. Generate 3 variations.

STEP 5 — STITCH AND RENDER (FFmpeg)
Use FFmpeg in bash to:
  1. Build a concat list of all visual clips in shot order, padding any 
     static images to their assigned duration
  2. Overlay the /audio/voiceover_full.mp3 as the audio track
  3. If total visual duration < audio duration, loop or extend the last 
     shot to match
  4. If total visual duration > audio duration, trim
  5. Add a 0.5s crossfade between shots
  6. Render to /final/video.mp4 at 1920x1080, H.264, AAC audio, 
     YouTube-ready settings (recommended: -crf 18, -preset slow, 
     -pix_fmt yuv420p)
  7. Verify the output file plays and report duration + file size

STEP 6 — METADATA
Write /final/title_options.txt with 5 title options (curiosity-driven, 
under 60 chars). Write /final/description.txt with a 200-word description, 
keywords, and placeholder for affiliate links. Write /final/tags.txt with 
20 relevant YouTube tags.

CRITICAL RULES
- Do NOT use any third-party brands, logos, or copyrighted characters in visuals
- Polling: if a Higgsfield job is still processing, wait 15 seconds and 
  re-check, up to 10 times before moving on
- If any step fails, report the error and continue with the next step — 
  do not silently skip
- At the end, give me a final report: word count, voiceover duration, 
  number of visual shots generated, final video duration, and the path 
  to /final/video.mp4

Begin.
```

## What it looks like

Written by Reference from the video frames, not by the creator.

Photoreal generated hotel scenes: a neon VACANCY sign on wet pavement, the sign laid over a minibar door, a marble lobby seen from behind a desk keyboard, a sunlit grand lobby, a gold "01 UPGRADES" card, a desk clerk at a computer, a dark corridor, and a phone with a booking confirmation.

- Color: Warm tungsten golds and browns, teal night exteriors, red and cream neon, and cream-and-gold type on dark wood paneling.
- Type: A large cream serif numeral over gold serif capitals for the section card ("01" and "UPGRADES"). Other text sits inside the images: the VACANCY sign and a phone screen reading "RESERVED — DIRECT".
- Framing: Wide, symmetrical interiors and close product-style shots with shallow depth of field. The creator's round face-cam sits in the lower left of the tutorial.
- Sound: Narration in the creator's own cloned ElevenLabs voice, according to the tutorial.
- Structure: The VACANCY sign twice, the sign over a minibar door, the lobby from the front desk, a grand lobby, the 01 UPGRADES card, the clerk at the desk, a night corridor, then the phone booking beside a coffee cup.

Storyboard, nine frames spread across the video: https://meetreference.com/media/faceless-youtube-explainer-things-hotel/storyboard.webp

- Frame 1 at 748.9s: https://meetreference.com/media/faceless-youtube-explainer-things-hotel/f1.webp
- Frame 2 at 756.7s: https://meetreference.com/media/faceless-youtube-explainer-things-hotel/f2.webp
- Frame 3 at 764.4s: https://meetreference.com/media/faceless-youtube-explainer-things-hotel/f3.webp
- Frame 4 at 772.2s: https://meetreference.com/media/faceless-youtube-explainer-things-hotel/f4.webp
- Frame 5 at 780.0s: https://meetreference.com/media/faceless-youtube-explainer-things-hotel/f5.webp
- Frame 6 at 787.8s: https://meetreference.com/media/faceless-youtube-explainer-things-hotel/f6.webp
- Frame 7 at 795.6s: https://meetreference.com/media/faceless-youtube-explainer-things-hotel/f7.webp
- Frame 8 at 803.3s: https://meetreference.com/media/faceless-youtube-explainer-things-hotel/f8.webp
- Frame 9 at 811.1s: https://meetreference.com/media/faceless-youtube-explainer-things-hotel/f9.webp

## How it was made, per the creator

- One-shot: no
- Notes: In the tutorial the creator picks Opus 4.7, then sends two follow-ups: his ElevenLabs voice ID, and a request to keep Seedance 2.0 for the intro only, use cheaper video models for the rest, and use GPT Image 2 for stills with a slow zoom. Claude wrote a 32-shot list, used Veo 3.1 Lite for the other motion shots, estimated 400 to 500 Higgsfield credits, and rendered a video 3 minutes 44 seconds long. The narration is the creator's own cloned ElevenLabs voice. The segment shown here is the first stretch of the finished video as it plays in the tutorial, with the creator's face-cam in the corner; it continues later, after some commentary. The tutorial is sponsored by Higgsfield.

## Links

- Watch: https://www.youtube.com/watch?v=Kgdms--LQ-g&t=745s
- Original post: https://www.youtube.com/watch?v=Kgdms--LQ-g
- Prompt page: https://mattpar.com/claude-code-prompts
