The Content Factory That Runs Without You On Camera
One AI employee, three tools, no camera. The exact stack — Authority Jacker script, ElevenLabs voice clone, HeyGen avatar — with every prompt to build a faceless video content factory.
On day 14 I stopped touching the camera completely. My AI employee now makes the videos end to end — it writes them, voices them, and puts a version of me on screen — and I never shoot a thing. No camera, no clips, no editing.
It is one AI employee wired to three tools: a script engine, a voice, and a face. Put them together and you have a content factory that runs without you ever being on camera. Here is exactly how it's built, with every prompt you need to run it yourself.
The three tools, wired together
The whole machine is three moving parts, in order:
- The script engine — Claude writes the script using a framework that positions you as the authority in your niche.
- The voice — ElevenLabs turns that script into spoken audio in your own cloned voice.
- The face — HeyGen takes that audio plus a photo of you and produces an avatar of you delivering it to camera, in any look you want.
Script in, finished talking-head video out. Nobody sets up a light. Let's build each part.
Step 1 — The script engine (the Authority Jacker)
Most "value" content teaches something and stops there. People learn, then forget who taught them. The Authority Jacker prompt is built so every section does double duty: it delivers real value and plants a specific, ownable claim to expertise (a named method, a specific number, or a narrow niche) that's hard to copy and easy to remember. A family-law attorney doesn't "give tips," they become the custody-schedule person in their city.
It runs on the 7-part Yap Framework (Calm Hook → Visual Proof → Benefit Re-Hook → Retention Statement → Named Steps → Save Bait → CTA), with an authority layer on every beat.
Paste this into Claude:
You are an expert short-form copywriter who specializes in "Authority Jacker"
scripts — high-retention Instagram Reels that position the speaker as the
definitive, go-to authority in their specific niche, using the 7-part Yap
Framework.
Your job is to interview me for the required inputs, then write a complete
Instagram Reel script using the exact structure below. Do not begin writing
until you have collected every required input.
---------------------------------------
STEP 1: INTERVIEW ME
---------------------------------------
Ask me these one at a time if I haven't already answered them:
1. What is my niche, and — critically — what is the NARROW slice of it I want to
own? (Not "renovation contractor" — "the guy who fixes botched kitchen
renovations other contractors walked away from.")
2. What result/transformation/proof am I claiming (a number, a before/after, a
specific client outcome)?
3. Who is this reel for — describe the exact person watching.
4. What are two currently-trending topics I can combine in the hook?
5. What proof do I have? (screenshots, revenue, before/after photos,
testimonials, review counts, completed job count, years in the niche)
6. What are 3-5 benefits of listening to me specifically vs. a generic
competitor?
7. What pain points does my expertise solve that a generalist can't?
8. How many steps/frameworks am I teaching (this becomes the Named Steps)? For
each step give me: a unique, ownable NAME for the method (the single most
important input — a named framework is what makes the authority claim stick),
what it does, the main lesson, and an open loop into the next step.
9. What should appear on the Save Bait screen — an information-dense
checklist/framework people will want to screenshot?
10. What keyword should people comment to get the full breakdown?
11. What exactly are they getting when they comment?
12. Any urgency, pricing, or bonus to mention?
13. What tone — calm, confident, educational, authoritative?
---------------------------------------
STEP 2: WRITE THE REEL
---------------------------------------
Use this exact structure. Format ON SCREEN / SPOKEN / CAMERA NOTES for every
section like a production script.
# SECTION 1 — Calm Hook (4-7s)
Combine the two trending topics with a specific, credible result. Say it like a
casual observation, not hype. The bigger the claim, the calmer the delivery.
# SECTION 2 — Visual Proof (3-5s)
Immediately show and name the proof. "In fact, here's [the exact proof]." One
sentence.
# SECTION 3 — Benefit Re-Hook (5-8s)
"Now the cool thing is..." — list 2-3 benefits of this specific expertise and
remove the most common objection to hiring a specialist over a generalist.
# SECTION 4 — Retention Statement (2-4s)
State how many steps are coming. Frame the LAST one as the most important, to
hold retention through the value section.
# SECTION 5 — Named Steps (30-50s)
For EACH step: tease the value -> introduce the named framework (a specific,
ownable name, e.g. "The Reveal-Day Protocol," not "my process") -> teach the real
lesson clearly -> open loop into the next step. Every named step reinforces
"nobody else frames it this way" — that IS the authority claim.
# SECTION 6 — Save Bait Interrupt (3-5s)
An on-screen, information-dense summary people can't fully read in one viewing —
dense enough that they rewatch and hit save. Write the full screen text exactly.
# SECTION 7 — CTA (8-12s)
"Now if you want [the full system]..." — state exactly what they get, mention any
pricing/urgency naturally, end with "Comment '[KEYWORD]' and I'll DM it to you."
Never sound pushy.
---------------------------------------
STYLE RULES
---------------------------------------
- Write exactly as if spoken aloud — short, conversational sentences.
- Every section must plant or reinforce the specific authority claim from input #1.
- Avoid sounding like sales copy. Open and close loops between every section.
- The Named Steps must use unique, specific names — no "step 1, step 2."
- Optimize for a 60–90 second reel that maximizes retention, saves, and comments.
You end up with a script that reads like the most confident person in your niche talking, which is exactly what the next two tools turn into a video.
Step 2 — The voice (ElevenLabs)
This is where the script becomes spoken audio in a voice that sounds like you, without you reading every video out loud.
Set it up once. Create an ElevenLabs account. Go to Voices → Add a new voice → Instant Voice Clone, and upload one to two minutes of clean audio of yourself talking: no music, no background noise, just you. That's the whole setup. From then on, every script you paste comes back in your voice.
Then, per video: open Text to Speech, pick your cloned voice, paste the Authority Jacker script (spoken lines only, strip the on-screen and camera notes), generate, and download the mp3. One tip: the model paces deliberately, so a script you think is 30 seconds often lands closer to 45. That's fine. The avatar lip-syncs to whatever length the audio is. Tighten the script if you need it shorter.
If you'd rather not clone your own voice yet, pick any natural stock voice to start. But cloning it once is what makes the whole factory feel like you, at scale.
Step 3 — The face (HeyGen avatar)
Now we put a version of you on screen. This is two moves: first create a clean source photo of yourself in the look you want, then let HeyGen turn it into a talking avatar driven by your ElevenLabs audio.
First, make the avatar source stills. Feed one clear photo of yourself to an image-edit model (Nano Banana / Gemini or GPT Image both work) with the prompts below. Each one keeps your exact face but places you in a different setting, so your videos don't all look identical. Here are three looks to start: a podcast studio, relaxed on the couch, and a walk-and-talk. The one rule that matters most in every prompt is preserve the exact likeness. That's the instruction that breaks the output if you drop it.
Podcast studio look:
Edit this reference photo into a podcast-studio avatar still, for use as a HeyGen
talking-head avatar source image.
SUBJECT & LIKENESS PRESERVATION:
Preserve this exact person's facial likeness from the reference photo — same face
shape, skin tone, eye color, hair color and style, and approximate age. Do not
beautify, slim, de-age, or change bone structure. Do not face-swap. This must read
as an edited photo of the SAME person.
WARDROBE: A solid, well-fitted quarter-zip or structured collared shirt in a muted,
confident color — navy, charcoal, or forest green.
SETTING: A modern podcast recording studio. Softly blurred background: acoustic
foam panels or slatted wood paneling, a warm bookshelf out of focus, one microphone
on a boom arm angled toward the subject but not blocking the face. "This person has
a show" energy — not a home office.
LIGHTING: Soft key light from camera-left, gentle fill from camera-right, warm
color temperature (~3200-4000K), subtle rim/hair light. Even, broadcast-quality
light on the face — no harsh shadows, no blown highlights.
LENS & CAMERA: 50-85mm portrait lens feel, shallow depth of field, eye-level,
subject looking directly into the lens.
FRAMING: Front-facing, chest-up, centered with even headroom, shoulders square to
camera. Nothing overlapping the mouth or jaw. Clean negative space above the head.
EXPRESSION: Calm, confident, mid-conversation — the look of someone mid-sentence on
a podcast, not a stiff headshot smile.
HEYGEN AVATAR REQUIREMENTS: Single subject, front-facing, upper-body framing, face
evenly lit with no motion blur, mouth and eyes fully unobstructed, clean background
separation. Output a single high-resolution still, portrait or square.
Relaxed-on-couch look: same prompt, but swap the wardrobe to a soft casual crewneck or quarter-zip in a neutral tone (heather grey, sage, stone), the setting to a comfortable living-room couch, softly blurred, a throw blanket on the armrest, warm lamp light out of focus, "talking to a friend" energy, the lighting to soft warm ambient light as if from a window or lamp (~3000-3500K), and the expression to warm, casual, mid-thought talking to a friend. Keep the likeness-preservation and HeyGen blocks identical.
Walk-and-talk look: same again, but wardrobe a smart-casual jacket or overshirt over a plain tee (camel, navy, charcoal), setting a downtown city street softly blurred to bokeh, daylight urban energy, framed and stabilized like a talking-head shot, not an action shot, lighting natural soft daylight or golden-hour warmth, no harsh midday sun, expression confident, energetic, explaining something mid-walk. Likeness and HeyGen blocks unchanged.
Then, in HeyGen: upload your source still and create a Photo Avatar (also called a Talking Photo). Choose the audio-driven option and upload the mp3 you made in ElevenLabs. HeyGen lip-syncs the avatar to your voice. Render, and you have a finished talking-head clip of you delivering the script, in whichever look you chose.
Putting it together — the factory
That's the machine: Authority Jacker script → ElevenLabs voice → HeyGen avatar. Once it's wired, a new video is just a new script fed through it.
Point it at the pain your customers actually have and it becomes a lead engine, not just content. That's the other half of this system, mining real pain points from reviews and writing them into scripts, and it's a full breakdown of its own: the faceless video system. Run the two together and you can put out fifty videos in your voice and face without ever setting up a camera.
The catch most people miss
Two things decide whether this looks real or generic. First, the voice clone is only as good as the audio you feed it. Record your sample clean, or every video inherits the noise. Second, the Authority Jacker only works if the named framework in Step 8 is genuinely yours. A real, ownable method is what makes people remember you; a generic "3 tips" script runs through the same machine and comes out forgettable. The tools are the easy part. The claim to a specific throne is the work.
If you'd rather we build and run this whole content factory for you, with the scripts, the voice, and the avatar on autopilot, that's what we do.
Frequently Asked Questions
Do I need my own voice for this, or can I use an AI voice?
Either. ElevenLabs lets you clone your own voice once from a short clean sample so every video sounds like you, or you can pick a natural stock voice to start. Cloning your own is what makes the factory feel personal at scale.
Do I have to appear on camera at all?
No. HeyGen builds a talking avatar from a single photo of you, driven by the audio. You never set up a camera, shoot a clip, or edit — the AI assembles the video end to end.
What does each tool cost to get started?
All three have entry tiers: Claude for the scripts, ElevenLabs for the voice, and HeyGen for the avatar. Voice generation is cents per script; avatar renders use HeyGen credits (roughly a minute of render each), so batch your videos before a run.
Why use the Authority Jacker framework instead of just writing a script?
Generic value content teaches and gets forgotten. The Authority Jacker plants a specific, ownable claim to expertise in every section — a named method, a specific number, a narrow niche — so viewers remember who taught them. That memory is what turns views into leads.
Can one avatar have different looks?
Yes. You generate separate source stills — a podcast studio, relaxed on the couch, a walk-and-talk — each preserving your exact likeness, then create a HeyGen avatar from each. That variety keeps your feed from looking like the same shot every time.