Create a consistent AI presenter, virtual model, or talking character from text or an authorized image. Choose text-to-video, image-to-video, or talking-video based on the source you already have.
Step 1: Open a public model and choose the video task
The problem: APOB includes several video modes, and choosing by feature name alone can open the wrong input workflow.
From the APOB home page, choose a model you are authorized to use. For this guide, we selected Miranda, a public Community model—not a private personal model. Open Video and select Text-to-Video, Image-to-Video, or Talking Video according to the source you already have.
Realistic example: A faceless creator without a source portrait can begin with Text-to-Video. A marketer who already has an approved character image should choose Image-to-Video.


Step 2: Use Text-to-Video when you do not have a source image
The problem: A short prompt such as “woman presenting a product” leaves identity, action, setting, and camera behavior unresolved.
The inspected Text-to-Video workspace kept Miranda attached to the Description field and exposed Native audio, Voice model, Art style, Camera movement, Clothing, Architecture, Landscape, Weather, Element, Template, and Add shot controls. Describe the character, one action, the setting, framing, lighting, and what must remain stable. Check Native audio rather than assuming the output will be silent.
Realistic example: A science creator can ask the recurring host to face the camera, make one small introductory gesture, and pause in a compact studio. Add a second shot only after the first identity and framing test succeeds.
Step 3: Use Image-to-Video when the person is already approved
The problem: A cropped, obstructed, or unauthorized reference can weaken motion and introduce rights or privacy risks.
Choose Select content to reuse authorized APOB media or Upload image for a source you have the right to use. A clear face, visible hands when relevant, and room around the body give the model more motion context. Direct the movement rather than redescribing every detail already visible.
Tested prompt: The subject takes one small step toward the camera, relaxes her shoulders, and gives a natural smile. Keep the face, hair, outfit, and background consistent. Use a slow, subtle camera push-in with stable hands and no sudden motion.
Realistic example: For a LinkedIn thought-leadership intro, request a small head turn and a steady medium close-up instead of a walk across the office.


Step 4: Choose Talking Avatar or Lip Sync for speech
The problem: A still image and an existing video require different speech workflows.
Talking Avatar starts from one reference image and either generated or uploaded audio. Lip Sync starts from an existing reference video plus generated or uploaded audio. In the inspected Talking Avatar interface, Create audio provided a Voice model and a script field of up to 2,000 characters.
APOB’s official guide documents a 1-second to 3-minute audio range for Talking Avatar: https://help.apob.ai/en/articles/15699750-talking-video . The Lip Sync guide documents a 1-second to 15-minute range for reference video and audio: https://help.apob.ai/en/articles/15699765-lip-sync .
Realistic example: A SaaS team creates a 12-second fictional trainer introduction, then cuts to a real screen recording so the instruction remains accurate.
Step 5: Check duration, quality, visibility, and credits
The problem: Teams can waste credits or expose unreleased assets when they assume every mode has the same cost or privacy state.
In our dated Image-to-Video test, the selected configuration showed five seconds, 720P, Fast/Ultra/UltraS choices, 200 credits per second, and a 1,000-credit estimate. These are observations from one account and session, not fixed pricing or plan entitlements. Check Native audio, the selected voice, and the visibility state before generating.
Realistic example: A fashion seller tests one restrained five-second turn before producing a multi-outfit lookbook. This confirms framing and identity before a larger credit commitment.

Step 6: Generate, confirm visibility, and review every frame
The problem: A clip can look convincing at normal speed while containing a distorted hand, shifting product label, inconsistent face, or misleading demonstration.
Our five-second Image-to-Video test completed in approximately two minutes and appeared in Generation with a Public label. Time varies by model, quality, and queue. Review face and hair continuity, eye direction, teeth, shoulders, hands, object contact, background stability, camera behavior, and the first and last frames.
Realistic example: If a presenter appears to press a skincare pump but the dispenser changes shape, keep the presenter as the spoken hook and cut to accurate product footage for the application shot. Do not use synthetic interaction as product proof.
Expert Prompts for AI Human Videos
Ecommerce presenter — purpose: a vertical product hook without changing the supplied product. Prompt: A fictional home-organization presenter stands waist-up in a bright apartment kitchen and holds the supplied compact drawer organizer at chest height. She raises it once, points to the adjustable divider with one measured gesture, and returns both hands to a stable position. Soft daylight, realistic hands, locked 9:16 camera with a subtle push-in. Preserve face, hair, cream knit top, product shape, label, color, and divider position. No extra fingers, label changes, or fast motion. Expected output: a focused five-second TikTok or Reels opening.
Recurring science host — purpose: a repeatable host shot for Shorts. Prompt: The same fictional male science host, late 20s, short dark curls, round glasses, navy overshirt, chest-up in a compact studio with a softly lit molecule display. He looks into the camera, raises one hand slightly, then returns to neutral. Locked 9:16 camera. Preserve face, glasses, hair, wardrobe, lighting, and background. One gesture only. Expected output: a recognizable hook or bridge shot.
Talking onboarding presenter — purpose: turn an approved trainer portrait into a spoken introduction. Script: Welcome to the dashboard. In this lesson, we will create one campaign, assign its owner, and check the approval status before publishing. Keep this window open, and follow the highlighted steps on screen. Expected output: a short presenter introduction that cuts to an accurate screen recording.
Fashion lookbook turn — purpose: show silhouette with restrained motion. Prompt: Full-body fictional fashion model in the supplied charcoal jacket and black straight-leg trousers, neutral gallery background, soft directional light. She takes one slow step, turns 30 degrees left, and pauses with both hands visible. Static medium-wide 9:16 camera. Preserve garment color, seams, buttons, proportions, face, and hair. Expected output: a controlled concept lookbook clip; generate a separate rear view instead of a full rotation.
Keep one recognizable person across related content. A recurring travel host can appear in airport tips, hotel explainers, and destination clips without redefining the identity in every prompt. Consistency still depends on reference quality, angle, lighting, model, and motion complexity.
Move from identity to motion and speech in connected workflows. Approve the fictional person before spending credits on animation, talking delivery, or campaign variations.
Test one creative variable without repeating a physical shoot. Keep the presenter and offer stable while comparing a problem-led, product-first, or demonstration-led hook. Do not use the workflow to invent customer experiences or endorsements.
See the generation estimate before committing. The tested interface displayed a per-second rate and total estimate beside Generate; always rely on the current app estimate.
Work within visible limitations. Motion, quality, camera, audio, and extra-shot controls do not remove the need to review hands, object contact, identity, rights, disclosure, and product accuracy.
Before publishing a realistic AI person, verify consent and disclosure. YouTube requires disclosure for realistic content that is meaningfully altered or synthetic; review the YouTube altered-content disclosure. Disclosure is not permission to impersonate someone, so also check the YouTube impersonation policy.
Test three TikTok Shop hooks with one fictional presenter. Keep the presenter, offer, and product footage stable while changing only the first line or gesture. HubSpot’s 2025 research identifies short-form video as marketers’ most-used content format and the format most often associated with highest ROI: https://blog.hubspot.com/news-trends/content-trends-global-preferences
Introduce ecommerce products without fabricating product proof. Use the AI person for the hook and transition, then cut to accurate product imagery or footage for color, dimensions, fit, assembly, and performance.
Build a recurring fictional creator series. Maintain a persona guide for face, hair, wardrobe, tone, disclosure, and restricted claims. IAB projected US creator-economy ad spend at $37 billion in 2025 and reported that three in four brands were using or planning AI for creator-marketing tasks: https://www.iab.com/insights/2025-creator-economy-ad-spend-strategy-report/
Add a host to faceless YouTube Shorts. Use a fictional host for the opening and conclusion, with licensed B-roll, diagrams, or real demonstrations between lines. YouTube reports that Shorts average more than 200 billion daily views: https://blog.youtube/press/
Create a controlled fashion concept lookbook. Animate one step or partial turn per clip and review garment edges, hands, silhouette, and color. Do not replace accurate product imagery when fit or material affects a purchase.
Localize product onboarding with a fictional trainer. Keep the approved screen recording unchanged and create short trainer introductions for different languages. Use a qualified reviewer for technical, regulated, or safety-related wording.
Fewer shoot dependencies — useful for concepts, recurring introductions, hooks, and creative variations that do not require factual physical proof.
Reusable identity — one approved fictional persona can connect a series across scenes and formats.
Controlled testing — change one hook, gesture, background, or framing choice at a time.
Platform-aware composition — plan vertical, square, portrait-feed, or landscape shots before generation.
Connected speech options — Talking Avatar starts from a still image; Lip Sync works with existing footage and new audio.
Human details can fail — hands, teeth, eye direction, hair, clothing, and object contact need close review.
Consistency is not guaranteed — new angles, lighting, models, and complex actions can alter the person.
Long performances are harder to control — scripts and actions usually work better as separate short shots.
Credits and limits change — cost, duration, quality, and controls depend on the current account and mode.
The publisher remains responsible for consent, claims, copyright, disclosure, privacy, and platform compliance.
No Credit Card Required
Workflow, quality, credits, privacy, commercial use, and platform compatibility
Can AI make a video of a person?
Yes. AI can generate a fictional person from a description or animate an authorized image. The result can be a silent lifestyle clip, virtual presenter, talking avatar, or recurring digital character. Review every output before publishing.
How do I make a human video from a photo?
Upload a clear image that you own or have permission to use, choose Image-to-Video, and describe one limited action plus one camera behavior. A front-facing portrait works for facial motion; a full-body image provides more context for a walk or outfit turn.
Can I start without a photo?
Yes. Create a fictional APOB persona or use a text-led video workflow. Define the person’s role, appearance, setting, action, framing, and format, then save an identity reference before producing a series.
Can I make the AI person talk?
Yes. Use Talking Avatar for one face image or Lip Sync for an existing video. The inspected Talking Avatar interface provided a Voice model and a script field of up to 2,000 characters, plus generated or uploaded audio.
What is the difference between Talking Avatar and Lip Sync?
Talking Avatar animates a still image from audio. Lip Sync keeps an existing video and resynchronizes the mouth to new audio. See https://help.apob.ai/en/articles/15699750-talking-video and https://help.apob.ai/en/articles/15699765-lip-sync .
Can Text-to-Video or Image-to-Video include audio?
The public-model workspaces we inspected displayed a Native audio toggle, and Text-to-Video also exposed a Voice model control. Availability depends on the selected model and setup. Check the toggle before generating.
What is the best free AI human video generator?
Best depends on your source asset, identity consistency, speech workflow, output rights, visibility, resolution, and total credit cost. Test the exact task with an authorized asset instead of comparing only homepage claims.
Is APOB a 100% free AI human video generator?
APOB uses credits for generation and provides ways to start without buying a full production package. It should not be described as unlimited or permanently free. Check https://app.apob.ai/subscription and the live estimate beside Generate.
Can ChatGPT make an AI human video?
ChatGPT can help draft a concept, script, shot list, or motion prompt depending on the product and plan. APOB provides the identity, image, motion, talking, and lip-sync workflow described on this page. Verify current capabilities in each product.
Can I use the same person in multiple videos?
Yes. Reuse the same APOB persona or authorized reference and retain a stable identity block in your prompts. Results can still vary with angle, lighting, model choice, and motion complexity, so inspect every clip.
How are video credits calculated?
The app displays a per-second rate and total estimate before generation. Talking Avatar and Lip Sync are billed according to audio duration; video-generation rates depend on the selected mode, duration, quality, account, and configuration.
Why is the Generate button unavailable?
Generate stays unavailable until the mode has its required inputs. Image-to-Video needs a reference image and motion description. Talking Avatar needs an image plus script or audio. Lip Sync needs reference video plus audio. Failed or incomplete uploads can also block generation.
Can I use APOB-generated videos commercially?
Commercial use depends on APOB’s current terms and the rights attached to every input and depicted element. Review likeness, voice, music, trademark, product, and source-media rights at https://public.apob.ai/tos/index.html .
Are my human videos private?
Do not assume so. Check the visibility label for the selected model and completed output. Our public Community-model test displayed a Public label. Review https://public.apob.ai/privacy/index.html before uploading sensitive material.
How do I make the result look more realistic?
Use a clear reference, request one action, keep camera motion restrained, describe what must remain stable, and review frame by frame. Human-object interaction and fast full-body movement are harder to control than a simple presenter shot.
Which aspect ratio should I choose?
Use 9:16 for TikTok, Reels, and Shorts; 16:9 for standard YouTube, websites, and training; and 1:1 or 4:5 for supported feed placements. Confirm the destination specification before generating.
Does APOB work for TikTok, Instagram, and YouTube?
Yes, the workflow can create source clips for those platforms. Export and edit according to each platform’s current aspect ratio, duration, caption, advertising, music, and synthetic-media requirements.
Do I need to disclose that the person is AI-generated?
Disclosure depends on realism, subject matter, platform, jurisdiction, and commercial context. Follow current YouTube, TikTok, and Meta rules, use platform disclosure controls, and never present a fictional character as a real customer, employee, or expert.
