comfyUI
Wan 3.0 vs MiniMax H3 in ComfyUI: Micro-Expressions, Cost and the AI-Influencer Question
A side-by-side ComfyUI test of Wan 3.0, paid per clip, and MiniMax H3, run locally on an RTX 5090: car-show, podcast and reaction clips judged on micro-expressions and lip sync, the price of quality, and the disclosure rules for AI influencers.
Akmal Alif · 8 October 2026 MYT

comfyUI is where we look at node-based workflows for AI image and video generation: what the models can do, what they cost, and how to use them responsibly. The first post is a live head-to-head.
In an August 2026 video, the creator behind the smallzero channel runs two video models side by side in ComfyUI with matching prompts: Wan 3.0, used through paid cloud generation, and MiniMax H3, run locally on an RTX 5090 (smallzero, 2026). The tests are aimed at AI-influencer and short-form social content, judged on micro-expressions, lip sync and realism. The video is embedded below; the timestamps jump to the matching moment. (The auto-captions mishear Wan as “Waifu Diffusion” and “Juan”; this post uses the model names from the video's description.)
The setup and the price gap
The presenter has been working mostly with MiniMax H3, calling it the best local model they have used. One of their ComfyUI workflows stitches clips into seamless two-minute sequences (watch from 0:00; 0:40). Wan 3.0 is the challenger. Early results shared in their community looked strong, but each generation costs real money in ComfyUI credits (watch from 1:02).
The presenter says up front that they are not paid by Wan, though Wan's account invited them on social media to make content about the model (watch from 1:21). For the tests, a 10-second, 720p, vertical 9:16 clip from Wan 3.0 cost between three and seven dollars, about four dollars in practice (watch from 3:26; 6:06).
Prompting with an agent
A practical tip comes before any generation. The presenter fed each model's official prompting documentation to a chat assistant, so a lazy one-line idea could be expanded into a prompt structured the way each model expects (watch from 2:02). Asking for footage that is “candid, amateur” and has some digital artifacts is a trick the presenter uses to make clips look more like real phone video (watch from 4:06).
Three tests
Text to video, a car-show interview. An influencer asks a shy car owner what year the car is, and the owner laughs. Wan 3.0's audio, micro-reactions and overall quality impress the presenter, despite loose adherence to where the character enters and leaves the frame. MiniMax H3 holds up well for a model running on a home GPU, but in the presenter's judgement it is not as good (watch from 4:46; 5:26).
Text to video, monkeys on a podcast. A non-human dialogue test about bananas. Wan 3.0 delivers convincing lip sync and expressions even on monkeys (watch from 6:46). MiniMax H3, at about 24 seconds per iteration on the 5090, produces a strikingly similar layout, which suggests similar training data, but follows the scripted dialogue less faithfully. The presenter says it usually takes three or four attempts to get a usable MiniMax scene (watch from 7:27; 8:07).
Image to video, a candid reaction. Starting from a still image, the subject realises they are being filmed rather than photographed and laughs. Wan 3.0's expressions are the high point of the video; the presenter calls it a possible new leader for AI-influencer content (watch from 10:08). MiniMax H3's version, once rerun at the right length, is closer than expected: image to video is competitive between the two (watch from 10:48).
The verdict: quality versus cost
The conclusion is a budget decision. With unlimited money the presenter would use Wan 3.0 exclusively; without it, MiniMax H3 stays the everyday model, and they hope Wan releases a local, distilled version (watch from 11:29). That is the pattern across generative video in 2026: the best results often come from paid cloud models, while local models trade some quality and more retries for no per-clip cost and full control.
The AI-influencer question
The video's use case deserves a plain note. Clips designed to look like candid, amateur phone footage of people who do not exist are exactly the kind of content audiences can mistake for real. Two rules already apply to anyone publishing them:
Platform disclosure. YouTube requires creators to disclose realistic altered or synthetic content that viewers could mistake for real people, places or events, and labels it accordingly (YouTube, 2024).
Advertising law. In the United States, the Federal Trade Commission's rule on consumer reviews and testimonials bans fake reviews and testimonials, explicitly including AI-generated ones attributed to people who do not exist (Federal Trade Commission, 2024).
Used openly, as fiction or with clear labels, these tools are remarkable for storyboarding, advertising concepts and entertainment. Used as fake customers or fake “real people” endorsing products, they cross a line that both platforms and regulators have already drawn.
References
Federal Trade Commission. (2024, August 14). Federal Trade Commission announces final rule banning fake reviews and testimonials [Press release]. https://www.ftc.gov/news-events/news/press-releases/2024/08/federal-trade-commission-announces-final-rule-banning-fake-reviews-testimonials
smallzero. (2026, August 25). Wan 3.0 AI test: Unbelievable micro-expressions for UGC [Video]. YouTube. https://www.youtube.com/watch?v=nA7LTAYXmlg
YouTube. (2024, March 18). How we're helping creators disclose altered or synthetic content. Google Blog. https://blog.google/intl/en-in/products/platforms/how-were-helping-creators-disclose-altered-or-synthetic-content/