EXEPERTAI LAB
EXEPERT / DIRECTORY
← Blog

AI4AnimationPy: Character Animation Research Moves From Unity to Python

Paul and Sebastian Starke's AI4AnimationPy brings the AI4Animation research line into Python: neural locomotion, contacts, IK and training visualisation in one PyTorch loop. The demo, the browser Space, the repository and the papers behind it.

Akmal Alif · 9 October 2026 MYT

A glowing wireframe figure strides across a dark grid floor, its hips trailing a curved trajectory line, red markers where each foot touches down and a phase dial turning beside it.

Body is the part of our Neuromorphic coverage about bodies: how artificial agents move, balance, touch the world and learn from doing it. Neuromorphic asks how brains compute. Body asks what that computation has to drive, starting with the most visible case, a character that has to walk convincingly.

The first project in the section is AI4AnimationPy, an open-source Python framework for neural character animation from Paul Starke and Sebastian Starke, published under the facebookresearch organisation on GitHub (Starke & Starke, 2026). A 100-second demo video shows what it does (Starke, 2026). There is no narration, so the timestamps below follow the on-screen labels.

AI4AnimationPy: Deep Learning for Character Animation in Python (Paul Starke)The YouTube player loads in privacy-enhanced mode when you press play.Watch on YouTube ↗

Where it comes from

AI4Animation is Sebastian Starke's long-running research framework for animating characters with neural networks. It sits behind a line of SIGGRAPH papers on the problem of making a character respond to a controller in real time without hand-authoring every transition:

  • Phase-functioned networks. Holden, Komura and Saito (2017) used phase-functioned neural networks for biped control.

  • Quadrupeds. Zhang et al. (2018) extended the approach to a four-legged dog with mode-adaptive networks.

  • Scene interaction. Starke et al. (2019) added a neural state machine for interactions with the scene: sitting, carrying, opening doors.

  • Contacts and phase. Later work handled multi-contact movement with local motion phases (Starke et al., 2020) and learned phase manifolds automatically (Starke et al., 2022).

  • Codebook matching. The most recent entry, categorical codebook matching for embodied character controllers, is co-authored by both Starkes (Starke et al., 2024).

That framework was built around Unity. The README of the new project explains the cost. Training happens in Python, while visualisation and runtime inference lived in a game engine, so the two had to talk through exported ONNX models or data streaming. AI4AnimationPy moves the whole loop into Python on NumPy and PyTorch, so training, inference and rendering share one backend (Starke & Starke, 2026).

What the demo shows

The video walks through the framework one feature at a time:

  • Locomotion. It opens on a biped controlled by a gamepad (watch from 0:00), then moves to three-point tracking, where a full body follows a small number of tracked points (watch from 0:08).

  • Engine plumbing. An entity-component-system scene shows the game-engine structure underneath (watch from 0:18).

  • Training features. Three modules visualise the features a network learns from: the root trajectory (watch from 0:28), joint trajectories (watch from 0:40) and foot-contact labels (watch from 0:46).

  • Inverse kinematics. A real-time solver using FABRIK, which adjusts a chain of joints so a hand or foot reaches a target (watch from 0:56).

  • Training examples. Side-by-side views of models learning to anticipate future motion while you watch (watch from 1:06).

  • Runtime modes. The video closes on its three ways to run: Standalone, Headless and Manual (watch from 1:30).

The three modes matter more than they look. Standalone runs with the built-in renderer. Headless runs the same code on a server without a window. Manual hands the update loop to the caller, which decides when each step runs.

Try it in the browser

The authors host web demos on a Hugging Face Space: a humanoid with style control and a quadruped with gait control. Each session is time-limited (it was three minutes when we checked), and the Space can take a minute to wake up if nobody has used it recently. The embed below loads it only when you ask.

Why Python changes the workflow

The README compares the new framework with the Unity version (Starke & Starke, 2026):

Task

AI4AnimationPy

AI4Animation (Unity)

Generating training data from 20 hours of motion capture

Under 5 minutes

Over 4 hours

Setting up a new experiment

About 10 minutes

Over 4 hours

Visualising inputs and outputs during training

Built in

Requires streaming

Backpropagating through inference

Supported

Not possible

Quantisation

Full PyTorch

Limited to ONNX

These are the authors' figures, not our benchmark, but the direction is easy to believe. When the controller, the training loop and the renderer live in one process, you can watch a model's predictions change while it trains instead of exporting and re-importing. You can also differentiate through the runtime itself. That is what "backpropagating through inference" means here: the code that drives the character at runtime can be part of what gets optimised.

The rest of the feature list reads like a small game engine:

  • Engine structure. A component system, an update loop with Update, Draw and GUI callbacks, and vectorised maths for kinematics and quaternions.

  • Models. Multilayer perceptrons, autoencoders and codebook matching.

  • Rendering. A renderer with shadows, ambient occlusion and bloom, plus GPU skinned meshes.

  • Motion data. Importers for GLB, FBX and BVH motion capture, stored internally as joint positions and rotations per frame.

  • Datasets. It works with public datasets, including 100STYLE, Ubisoft's LaFAN1 and the original NSM and MANN captures.

  • Coming next. Physics, path planning and audio are listed as coming next.

The community thread, and the source

The project also travelled through community channels. One example is a video post titled "Open-Source AI Character Animation Framework" in Reddit's r/TopologyAI on 21 April 2026 (Delicious-Shower8401, 2026). The post and its discussion are below, followed by the repository itself. GitHub does not allow its pages to be framed, so it appears as a link card.

One practical detail before you build on it: the code is released under Creative Commons Attribution-NonCommercial 4.0. That suits research, teaching and experiments. A commercial game or product needs a separate conversation with the rights holders.

Why this belongs under Body

Movement is where intelligence becomes visible. A model can describe a staircase, but a character controller has to place a foot on the right step, at the right time, with the right weight transfer, sixty times a second. The research line behind AI4Animation is about learning that kind of control from data: phases, contacts and trajectories rather than hand-written state machines.

The same questions run through robotics and through neuroscience's interest in motor control. Bringing the tools into the language most machine-learning researchers already use lowers the cost of asking them. That is why it opens this section.

References

Delicious-Shower8401. (2026, April 21). Open-source AI character animation framework [Online forum post]. Reddit. https://www.reddit.com/r/TopologyAI/comments/1srueru/opensource_ai_character_animation_framework/

Holden, D., Komura, T., & Saito, J. (2017). Phase-functioned neural networks for character control. ACM Transactions on Graphics, 36(4), 1–13. https://doi.org/10.1145/3072959.3073663

Starke, P. (2026, March 17). AI4AnimationPy: Deep learning for character animation in Python [Video]. YouTube. https://www.youtube.com/watch?v=LKl7MzFENUs

Starke, P., & Starke, S. (2026). AI4AnimationPy [Computer software]. GitHub. https://github.com/facebookresearch/ai4animationpy

Starke, S., Mason, I., & Komura, T. (2022). DeepPhase: Periodic autoencoders for learning motion phase manifolds. ACM Transactions on Graphics, 41(4), 1–13. https://doi.org/10.1145/3528223.3530178

Starke, S., Starke, P., He, N., Komura, T., & Ye, Y. (2024). Categorical codebook matching for embodied character controllers. ACM Transactions on Graphics, 43(4), 1–14. https://doi.org/10.1145/3658209

Starke, S., Zhang, H., Komura, T., & Saito, J. (2019). Neural state machine for character-scene interactions. ACM Transactions on Graphics, 38(6), 1–14. https://doi.org/10.1145/3355089.3356505

Starke, S., Zhao, Y., Komura, T., & Zaman, K. (2020). Local motion phases for learning multi-contact character movements. ACM Transactions on Graphics, 39(4). https://doi.org/10.1145/3386569.3392450

Zhang, H., Starke, S., Komura, T., & Saito, J. (2018). Mode-adaptive neural networks for quadruped motion control. ACM Transactions on Graphics, 37(4), 1–11. https://doi.org/10.1145/3197517.3201366