How to make a viral video with Higgsfield + ChatGPT
Build a 15-second video from idea to final edit, with copyable ChatGPT prompts, Higgsfield shot instructions, and a complete worked example.
Build your video, step by step
Plan the story, generate the shots, edit, and test.
Make a short video people want to finish and share. Use ChatGPT to develop the idea, script, and shot prompts; use Higgsfield to generate the visuals; then assemble the clips in your editor.
Your first project: a 15-second vertical video with one visual surprise, four shots, and a clear payoff. This guide uses “What if your morning coffee came with its own weather?” as a worked example. Nobody can guarantee virality. Build a strong first version, measure how people respond, and improve the next one.
1. Pick one idea worth watching
Start with an audience and a reaction. For example: “People who love coffee and cinematic visuals should feel surprised enough to send this to a friend.” A useful concept fits into one sentence and can be understood on a phone.
Try an ordinary object with an unexpected rule: a coffee cup containing a storm, sneakers leaving miniature forests behind, or a desk becoming a tiny city. Choose something you can communicate visually without a long explanation.
Copy into ChatGPT: find the concept
Help me develop ideas for a 15-second vertical video. Audience: [who it is for]. Niche: [topic]. Desired reaction: [surprise, laughter, recognition, curiosity]. Tools: ChatGPT for planning, Higgsfield for visuals, and an editor for assembly. Available assets: [my photos, products, footage, or none]. Suggest five original concepts. For each, give: - A one-sentence premise. - The first image the viewer sees. - The visual payoff by the end. - Why this audience might share it. - The hardest shot to generate and a simpler alternative. Keep each concept feasible in four short shots. Avoid invented statistics, fake endorsements, and claims of guaranteed virality. Recommend one idea based on clarity, originality, and production difficulty.
Pick the concept yourself. A flashy effect needs a reason to exist in the story. For the example, the rule is simple: the weather inside the cup changes with the drinker's mood.
2. Write the hook and four-shot story
Give the viewer something interesting immediately, then answer the curiosity you create. Skip the logo intro. Make the first frame show the premise, and let each following shot add information.
A practical starting structure is surprise → discovery → escalation → payoff. The timings below are editing targets; they are not Higgsfield generation settings.
Worked example: “My coffee has a weather forecast”
- 0–3 seconds — surprise: a miniature storm swirls inside a white coffee cup. On-screen hook: “My coffee has a weather forecast.”
- 3–6 seconds — discovery: a wider view shows the storm hovering above the coffee on a kitchen table. Caption: “Before the first sip…”
- 6–10 seconds — escalation: a closer angle reveals a tiny flash illuminating the cloud. Caption: “…100% chance of chaos.”
- 10–15 seconds — payoff: match cut to the same cup with a golden glow and a miniature rainbow. Caption: “After coffee: clear skies.” End with “Who needs this cup?”
Use a cut between storm and sunshine instead of asking the model to perform a complicated transformation. The visuals are fictional AI scenes; keep that clear in the post.
Copy into ChatGPT: turn the idea into a script
Turn this concept into a 15-second vertical video: [chosen concept]. Audience: [audience]. Tone: [tone]. Create exactly four shots. For each, give the edit time range, visual action, on-screen text, optional narration, sound cue, and what new information it adds. The time ranges must cover 0–15 seconds without gaps. Open on the strongest visual. Deliver the promised payoff before the ending. Keep captions short enough to read on a phone. If there is narration, keep the total under 35 words and tell me to time a spoken read. Suggest three distinct opening hooks that still match the same story. Flag difficult animation and replace it with simpler action or a match cut. Describe a specific share-worthy moment without predicting view counts.
Read the script aloud and time it. Cut words until it fits comfortably. Choose either brief narration or captions alone for your first attempt.
3. Create reference images before animation
You will need ChatGPT, access to Higgsfield, and an editor such as CapCut, Canva, or DaVinci Resolve. Have a small generation budget ready. Check the credits shown for your selected settings before each run.
Create or upload a still image that establishes the look of each shot. In Higgsfield, open its image tools and make the first reference. If you already have an image you own, start there. Keep the same object, colors, materials, location, and lighting across the sequence.
For the coffee example, reuse these anchors: plain white ceramic cup, round handle on the right, light oak table, cream kitchen background, soft morning window light from the left, no logos. Save the best image and reuse it as a reference wherever your selected model supports that input. A prompt requests consistency; inspect the result yourself.
Copy into ChatGPT: prepare the image prompts
Create four reference-image prompts for this approved shot list: [paste the shot list]. Visual anchors that must remain consistent: [subject identity, colors, materials, setting, lighting]. Style: [photorealistic, illustration, or another specific style]. Format: vertical 9:16 with the important action near the center. For each shot, write one standalone prompt specifying subject, setting, framing, lighting, and the story detail visible in that frame. Repeat the visual anchors in every prompt. Leave space for captions and request no generated text or logos. These are still images: describe a visible moment, not a sequence of actions. Identify which shots can reuse the same reference image and which need a new angle. Do not invent tool buttons or model features.
Example reference-image prompt
Vertical 9:16 photorealistic close-up of a plain white ceramic coffee cup, round handle on the right, on a light oak kitchen table. Soft morning window light from the left; cream kitchen background. Inside the cup, coffee surrounds a miniature dark storm cloud hovering just above the liquid. The tiny cloud is clearly visible and contained within the cup's width. Believable ceramic and liquid textures, whimsical fictional scene, clean centered composition with room above for a caption. No people, lettering, logos, or extra cups.
Check every reference before animating: the cup must have one handle, the cloud must read clearly at phone size, and the caption area must be usable. Correct a bad still first.
4. Generate each shot in Higgsfield
Higgsfield supports workflows that start from text or from an image. For this project, use image-to-video to give each shot a visual starting point. See Higgsfield's beginner workflow in Sources below.
- Open Higgsfield's video tools and select a model that supports image-to-video in your account.
- Upload the approved reference image as the starting image.
- Paste the motion prompt for that shot. Specify one main action and one camera move.
- Choose vertical 9:16 when supported. If unavailable, compose for a vertical crop and check that the subject survives it.
- Choose a short supported duration that covers the shot's edit time. Start with economical preview settings when available. Check the displayed credit cost.
- Generate, watch the entire result, and download the usable clip. Repeat for the other shots.
Controls, duration, sound, and reference support depend on the model. If you select a camera preset, keep the written movement consistent with it. A slow push-in and a conflicting orbit instruction give mixed direction. Higgsfield's camera guide and Prompt Bank provide movement examples.
Copy into ChatGPT: write the motion prompts
Write one image-to-video motion prompt for each approved shot: [paste the shot list and visual anchors]. Each prompt should describe one main subject action, one camera movement or a locked camera, the intended pace, and details to preserve from the reference. Keep it concise and physically clear. Avoid combining multiple locations or a complex transformation in one clip. Assume I upload a reference image for each shot. Put caption text, transitions, and sound cues in separate editor notes, not in the visual-generation prompt. If motion is difficult, give a simpler fallback. Generation settings will be chosen in Higgsfield; do not promise unsupported settings or exact timing from prompt text.
Example storm-shot motion prompt
Animate the supplied coffee-cup reference. The miniature storm cloud rotates gently above the coffee while small ripples move across the liquid. Camera makes one slow push-in toward the cloud. Keep the cup shape, handle position, oak table, cream background, and left-side morning light consistent with the reference. The cloud stays small and above the liquid. No new objects, lettering, scene changes, or cup deformation.
Generate the opening first. If the core surprise does not work, simplify it before spending credits on all four shots. For the payoff, create a separate sunny reference and animate subtle light shimmer with a locked camera.
5. Fix weak shots without burning the budget
Watch from beginning to end. Look for disappearing handles, drifting objects, unnatural liquid motion, unwanted text, and an action that never actually happens.
- The object changes shape: shorten the action, use a clearer reference, and reduce camera movement.
- The visual surprise is hard to see: enlarge it in the reference and move to a closer composition.
- The camera wanders: request a locked camera or select a matching static control when available.
- The scene is too busy: remove background activity and keep one action.
- The clip looks good but feels slow: trim it in the editor and cut when the viewer has understood the beat.
Change one variable at a time and record what you changed. Set a stopping rule before generating, such as two attempts per shot. If a shot still fails, use the fallback instead of repeatedly paying for the same difficult action.
6. Assemble a video people can follow
Import your clips into a 1080 × 1920 vertical timeline. Arrange the four beats and trim to about 15 seconds. If a generated clip is longer, use only its useful part. Add the captions in your editor so you control spelling, placement, and timing.
Keep important text away from screen edges and the platform's buttons. Preview with the posting app's interface visible before publishing. Choose readable, high-contrast text, keep the same caption style, and let captions remain long enough to read.
For the example, use a quiet rumble under the storm, a small sound accent for the flash, and a brighter sound at the sunny reveal. If you use narration, lower the music until every word is clear. Use audio you have permission to use.
Watch once with sound off and once with sound on. Both versions should communicate the premise and payoff. Export an MP4 at 1080 × 1920 if your editor supports it, then check the exported file on your phone.
Copy into ChatGPT: review the edit
Help me review this short-video edit. Premise: [premise]. Audience: [audience]. Actual shot order and timing: [list]. On-screen text: [captions]. Narration, if any: [script]. Evidence supplied: [screenshots, notes, transcript, or supported video input]. First state what you can inspect. Do not claim to have watched motion or heard audio if I supplied only still images or text. Check whether the opening communicates the premise, each shot adds information, captions can be read in their allotted time, and the ending delivers the promised payoff. Suggest the three highest-impact changes, with specific replacements. Separate observed issues from things I need to check in the actual exported video.
7. Post, measure, and improve the hook
Make three opening versions using the hooks from section 2. Keep the rest of the story the same so you have a useful comparison. Start by showing them to a few people in your intended audience; ask what made them curious and where they lost interest.
Publish your chosen version with a caption that explains the premise. For the example: “My morning mood, turned into a tiny weather system. Made with AI using Higgsfield + ChatGPT. Who needs this cup?” Use the platform's AI disclosure setting where required. Check current posting requirements in the app.
Compare posts after similar observation windows, such as 48 hours. Record video length, views, average watch time, completion rate if available, shares, and saves. Look at shares and saves relative to views rather than only raw counts. Metrics differ by platform, and organic posts are not a controlled experiment.
- People leave early: test a clearer first frame or a shorter hook.
- People watch but rarely share: make the payoff more surprising or more relevant to the audience.
- People save it: try a follow-up that teaches the process.
- One version performs better: reuse the useful structure with a new original idea; check it over several posts.
Copy into ChatGPT: learn from the results
Analyze these short-video results without claiming causation. Platform: [platform]. Audience: [audience]. For each version: [hook, video length, publish time, observation window, views, average watch time, completion rate if available, shares, saves]. My goal: [shares, saves, profile visits, or another result]. Compare only the metrics supplied. Calculate shares/views and saves/views as percentages when views are available and above zero. Flag mismatched observation windows, missing metrics, and small samples. Treat the findings as hypotheses. Recommend one change to test in the next video and explain which metric would tell me whether it helped. Do not invent universal viral thresholds or predict a view count.
Before you hit publish
- The first frame shows the visual idea immediately.
- Every shot adds something and the payoff matches the hook.
- The exported video works on your phone with sound off.
- Captions are readable and clear of interface overlays.
- Objects remain consistent throughout each generated clip.
- Music, references, and any identifiable people are cleared for use.
- Fictional AI imagery is presented honestly.
- You saved the prompts and settings so the next attempt is easier.
Your first goal is one finished video and one useful lesson from its performance. Keep the idea clear, spend credits deliberately, and improve one thing at a time.