Multimodal AI agent framework by ByteDance that processes text, images, and videos to perform complex tasks through autonomous interaction.