Written in Chinese, the skill orchestrates ByteDance's agentkit-samples multimedia tools to build a nine-by-sixteen vertical video of a fixed shopping-guide character recommending a product. It generates the character's portrait and a scene background with AI image generation at a resolution of at least 1080 by 1920, cropped to the vertical ratio, synthesizes the sales script as sixteen-kilohertz mono speech and a matching background music track, then feeds the character image, scene image, voice file, music file and prompts into a video generator that produces five five-second scenes stitched into one 25-second clip.
Each step can also run alone when only the avatar, the background, the voice or the music is needed. The skill depends on Pillow, requests and numpy, and needs agentkit-samples installed plus credentials for any third-party services it calls. Reference files hold the character's fixed appearance, scene templates and tool usage notes, and the character's look must stay consistent across every generated asset.