Text to 3D motion turns one plain-language sentence into skeletal motion data: you describe an action, an AI motion model generates the movement on a rigged default character, and you export an animated FBX that opens in Blender, Unity, or Unreal. No motion capture suit, no reference footage, no keyframing, and no character upload are involved. The prompt is the only input.
That makes it the fastest route from an idea to a moving character. Where mocap needs equipment and preset libraries need luck, text to 3D motion needs a clear sentence. This guide explains what the technology actually produces, how the pipeline works, how to write prompts that generate good motion, and how to carry the result into your own projects.
Key Takeaways
- Text to 3D motion generates skeletal motion from a written description and plays it on a default character that is already rigged and skinned; the export is an animated FBX, not a rendered video.
- A reliable motion prompt follows the pattern character, then action, then detail, within a 128-character limit. Simple English or Chinese both work.
- Durations are fixed at 3, 5, 8, or 12 seconds, or Auto, which picks a length that matches the action. Each generation costs 20 credits regardless of duration.
- Prompt enhancement is on by default: the AI rewrites your sentence with extra detail before generating. Turn it off when you need literal precision, such as exact punch counts.
- To move the motion onto your own character, retarget the FBX onto a rigged model, or use the complete animation workflow that rigs your character and applies motion in one run.
What is text to 3D motion?
Text to 3D motion is a generation technique that converts a natural-language description of an action into 3D character motion. You write what should happen, for example "a person walks forward and waves with the right hand", and the system produces a skeletal animation that performs the described action. The output is motion data, the same kind of asset a motion capture session or an animator's keyframes would produce.
That last point matters more than it sounds, because search results for "text to animation" mix two different products. Some tools generate rendered video clips with talking avatars, aimed at social content. Text to 3D motion generates an animated FBX file with a skinned character, aimed at game engines and 3D software. If your goal is a moving character inside Blender, Unity, or Unreal, you want the second kind. If your goal is a finished video with zero 3D work, the first kind fits better.
Consider Tomás, a level designer blocking a stealth mission. He needed a guard patrol loop for a graybox scene: walk, stop, look around, keep walking. He had no mocap suit, no footage, and no time to keyframe a placeholder. He typed the action into a text to 3D motion tool, chose a 5-second duration, and previewed the generated clip on the default character within minutes. The clip was not final animation, and it did not need to be. It made the scene reviewable, and the note "guard motion feels too slow" became possible two weeks earlier than it otherwise would have.
How does text to 3D motion work?
The pipeline has three stages. You supply the sentence; the system supplies everything else, including the character.
1. Write the motion prompt
Describe the action in plain language, within a 128-character limit. A structure that works consistently: name the character, state the action, then add detail. "A person walks forward and waves with the right hand" follows it exactly. Writing in simple English produces the most predictable results, and Chinese is supported too. One connected action chain beats a list of unrelated verbs, a rule the 128-character limit enforces helpfully.
2. Generate the motion with enhancement
The motion model reads the prompt and generates matching movement. Before generation, prompt enhancement rewrites your sentence with additional detail: a two-word prompt arrives at the model as a fuller description, which usually produces more natural motion. Enhancement is on by default and included in the price. When literal precision matters, exact action counts, exact directions, precise ordering, you can turn it off in the advanced settings so the model follows your wording as written.
You also pick a duration at this stage: 3, 5, 8, or 12 seconds, or Auto, which chooses a length that fits the described action. Each generation costs 20 credits, whatever duration and enhancement setting you choose. Under the hood, the workflow runs on Tencent Hunyuan 3D motion generation.
3. Preview and export the animated FBX
Play the result in the browser before committing. If the motion reads wrong, the fix is usually a prompt edit, not a new pipeline run. When the preview satisfies you, export the animated FBX. The file includes the skinned default character, so it drops into standard 3D tools without assembly: no skeleton to connect, no weights to bind.
If this pipeline matches your work, the text to 3D motion tool runs it in the browser.
Text-to-motion vs mocap vs preset libraries vs keyframing
Every route to character motion solves the same problem differently. The right question is not which tool is best but which input you already have.
| Motion source | What you provide | What you get | Best for | Watch out for |
|---|---|---|---|---|
| Mocap suit session | Hardware, a performer, a take | High-fidelity motion data | Hero-quality realistic movement | Cost and scheduling |
| Video motion capture | Footage of the action | Mocap extracted from video | Movement you can film | You need the footage first |
| Preset library | A rigged character | A clip from a fixed list | Standard loops: walk, run, idle | Your exact action may not exist |
| Text prompt | One sentence | Generated motion on demand | Specific actions, fast iteration | Vague prompts produce vague motion |
| Manual keyframing | Skill and time | Exactly what you pose | Final polish, signature shots | Slowest route by far |
Preset libraries are the route most creators already know. Adobe's Mixamo made it the default: auto-rig a humanoid, browse a mocap library, export FBX for free. For a standard walk or run cycle it remains a solid answer. The limitation is selection. When the guard needs to walk, stop, and look around as one continuous action, three separate clips need cutting and blending, if all three exist at all.
Text to 3D motion inverts the deal: the motion is generated for your sentence instead of searched for in a catalog. Iteration happens by editing words, which is faster than re-recording footage or scrubbing a timeline. The trade is control. You are steering a generative model with language, so the result matches your intent, not your frame count.
Manual keyframing stays the route for final quality. Tools like Cascadeur add AI-assisted posing and physics correction to a desktop animation suite, and hero shots deserve that level of attention. The practical split many teams land on: block and iterate with generated motion, then keyframe or clean the shots that matter most.
How to write a motion prompt that works
A motion prompt is a stage direction, not a search query. The model follows what the words specify, so the words need to specify something.
Three parts make a prompt reliable:
- The character: who moves. "A person" is enough; the tool animates its default character.
- The action chain: one connected sequence, not a wish list. "Walks forward, stops, and looks around" is a chain. "Dances and fights and jumps" is three prompts pretending to be one.
- The details: direction, handedness, speed, and ending. "Waves with the right hand, slowly" tells the model things it cannot guess.
Compare weak prompts with stronger versions:
| Weak prompt | Stronger prompt | Why it works |
|---|---|---|
| "dance" | "claps twice, spins left, and bows" | Countable actions, direction, ending |
| "fighting" | "two quick jabs followed by a roundhouse kick" | Order and speed are explicit |
| "walks" | "walks forward slowly, stops, and looks around" | A paced chain, not a single verb |
| "exercise" | "does three jumping jacks then stretches both arms up" | Specific, countable, and endable |
The 128-character limit is a design feature. It is roughly the space of one well-specified action chain, and it forces you to decide what this clip actually is. When your idea does not fit, the idea is probably two clips.
Simple actions generate most reliably: walking, running, jumping, waving, short dances. Gymnastics-grade combinations can produce minor mesh intersections where limbs cross, so treat extreme actions as experiments and check them in the preview.
Duration deserves a deliberate choice, because the same prompt reads differently at different lengths:
| Duration | Fits | Example |
|---|---|---|
| 3 s | One gesture or single motion | Wave hello, turn around |
| 5 s | One action chain with an ending | Walk forward, stop, look around |
| 8 s | A combination or short sequence | Run and stop, punch combo |
| 12 s | A small performance | A short dance |
| Auto | When unsure; matches length to action | Any of the above |
Aisha, a previz artist on a short film, learned the limit's lesson in one afternoon. Her first prompt tried to schedule four actions inside 128 characters: approach, sit, stand, wave. The generated motion was mush, each action sampled and none completed. Splitting it into two prompts, "walks to the chair and sits down" and "stands up and waves", produced two clean clips that cut together exactly as her storyboard needed. The 128 characters were not the constraint. The single-clip framing was.
Prompt enhancement: when to keep it on, when to turn it off
Prompt enhancement is a quiet feature that changes results more than its name suggests. With enhancement on, the AI rewrites your prompt before generation, filling in physical detail your sentence left out: weight shifts, follow-through, a natural starting pose. A prompt like "a person dances" arrives at the motion model as a much fuller description, and fuller descriptions usually generate fuller motion.
Because it runs automatically and costs nothing extra, the default is worth keeping for most work:
| Situation | Enhancement | Reason |
|---|---|---|
| Short or generic prompt | Keep on | The rewrite adds the detail you skipped |
| Natural, expressive motion wanted | Keep on | Follow-through and weight read better |
| Exact action counts matter | Turn off | "Two punches" should stay two punches |
| Precise directions and ordering matter | Turn off | Left stays left; the sequence stays yours |
The off switch lives in the advanced settings. The rule of thumb: enhancement interprets, so use it when you want interpretation and disable it when you want obedience.
From the exported FBX into your pipeline
The export is one animated FBX that contains both the motion and the skinned default character performing it. For Blender, Unity, Unreal, and most other 3D tools, import is a standard FBX workflow: check scale and up-axis on the way in, and confirm the animation data arrived with the mesh. Blender's FBX import documentation covers the settings that matter when an imported character shows up too large, rotated, or without its animation.
From there, two paths lead to your own characters.
The first path is retargeting. Motion data moves between skeletons well, and every major DCC and engine has retargeting tools for exactly this job. The requirement is a rigged target: your character needs a skeleton before it can receive the motion. If your model is still a static mesh, an AI auto rigging tool can generate the skeleton and skin weights first, and then the retarget has something to land on.
The second path skips the assembly. The AI 3D character animation generator runs auto rigging and motion generation on your uploaded FBX or GLB in one workflow, so the motion lands on your character, not the default one. It is the right tool when the destination character already exists and the motion matters more than the drafting speed.
Ken, a solo developer, took the first path with a twist. He generated a punch combo on the default character, then tried to retarget it onto his game's low-poly fighter and found nothing to retarget onto: the model had no skeleton. One auto-rigging pass later, the retarget worked, and the whole detour cost less time than keyframing the combo would have. The lesson is not that retargeting is hard. It is that the skeleton is the prerequisite nobody mentions.
Common mistakes to avoid
Cramming several actions into one prompt. The 128-character limit is not a quota to spend. One chain per clip, cut together in your editor, beats one crowded clip every time.
Writing single-word prompts. "Dance" gives the model a category, not an instruction. Countable actions with a direction and an ending generate motion you can use.
Expecting frame-exact choreography. Text to 3D motion matches intent, not keyframe numbers. When timing must be exact to the frame, that shot belongs in keyframing, possibly with enhancement off and a second cleanup pass.
Skipping the preview. The browser preview exists so that clipping, folded wrists, and sliding feet get caught before an import round-trip. Two minutes of watching saves an hour of wondering where the artifact came from.
Treating Auto duration as a gamble. Auto reads the described action and picks a matching length. When you genuinely do not know whether a chain needs 5 or 8 seconds, Auto is a reasoned default, not a coin flip.
FAQ
What is text to 3D motion?
Text to 3D motion is a technique that turns a written description of an action into 3D character motion data. You write a sentence like "a person walks forward and waves", and an AI motion model generates the movement on a rigged default character, exported as an animated FBX for 3D software and game engines.
How long can text-generated motion be?
Durations of 3, 5, 8, or 12 seconds are available, plus an Auto option that picks the length matching the described action. For longer sequences, generate the action in parts and edit the clips together.
Do I need a rigged character to use text to 3D motion?
No. The tool animates its own default character, which is already rigged and skinned. You only write the prompt; there is no upload step.
Can I apply the motion to my own character?
Yes, in two ways. Retarget the exported FBX onto your character, which requires your character to have a skeleton first, or use the complete animation workflow that rigs your uploaded model and generates motion on it directly.
What does the export include?
One animated FBX file containing the motion and the skinned default character performing it. It imports into Blender, Unity, Unreal, and other tools that read FBX animation, without extra assembly.
How much does text to 3D motion cost?
Each generation costs 20 credits, regardless of duration or the prompt enhancement setting.
From a sentence to a moving character
Text to 3D motion collapses the distance between describing an action and seeing it performed. Write one clear sentence with a character, an action chain, and a detail or two, choose a duration or let Auto decide, preview with a critical eye, and export an animated FBX that works wherever FBX works. Keep enhancement on while exploring, turn it off when precision matters, and remember that the skeleton is the ticket when you want the motion on your own character.
When the next motion idea is still just a sentence in your head, generate 3D motion from text and watch it move. Start with the signup credits, keep the prompts specific, and let the first clip start the conversation.

