Core
Inspect local files and media, extract video frames, and crop or annotate images.
Explore multimodal skills, tools, and real examples. Installation guide
Inspect local files and media, extract video frames, and crop or annotate images.
Understand images, audio, and video through model APIs, including OCR, object localization, and speech transcription.
Search the web, read pages, and identify objects or places with reverse-image search.
Build searchable, hierarchical memory of long videos to summarize content and locate events, text, and dialogue.
Build and query audio-visual memory to track speakers, dialogue, sounds, and events across videos.
Edit existing footage into finished videos with pacing, sound, subtitles, and visual effects.
Create, refine, and render 3D scenes and assets in Blender.
Create and edit parametric CAD models, technical drawings, and model exports in FreeCAD.
Create narrated Mandarin math and science tutorial videos or interactive explainers from problem statements and images.
Convert a local tutorial video into an illustrated PDF using Omni audio-video understanding, screenshot selection, and PDF review.
Omni ChatCut video creation with Music-to-MV, movie commentary, and speaker-preserving video translation.
Turns a demonstration video into a reusable Agent Skill. Needs a DashScope key and ffmpeg.
Operate physical hardware through Model Hardware Standard adapters: discover devices, read sensors and camera frames, send commands with host-side safety-limit enforcement, check health, and emergency-stop — as MCP tools. Adapters are run by the hardware's owner.