Video-Spatio — spatial reasoning
Use this Skill when a question needs spatial reasoning: distance, size, facing, left/right/front/behind,
camera versus object motion, cross-frame counting, or viewpoint-dependent visibility. Answer appearance
questions directly from the images. Read the qwen-mm-plugins-video-spatio tool schemas for arguments.
The host model supplies object boxes, distance estimates, and camera-motion estimates. Geometry tools calculate from those inputs; they do not independently measure the scene. No GPU perception server is required. VLM-backed tools use the configured OpenAI-compatible endpoint and may make several calls.
Build a scene
- Inspect the image or sample video frames. If installed,
coreprovidesread_videoandsave_view. Otherwise use available image/video tools and provide local frame paths. - Ground the relevant objects in each frame, with one list per frame:
{"label":"chair","bbox":[400,300,600,800],"depth_m":2.3}. - Call
build_scene(frames=..., objects_by_frame=...). All frames must have the same dimensions. Use uniqueframe_indiceswhen retaining original video indices. - For a moving camera, supply one
camera_motionsentry per adjacent pair. Translations are in the preceding camera's frame:forward_mis positive forward,right_mpositive right, andyaw_degpositive for a right turn. Omitting motions assumes a static camera; it does not estimate motion. - Pass the returned
sceneJSON to subsequent tools, or save it and passscene_file.
Coordinates:
- Boxes are
[x1,y1,x2,y2], top-left to bottom-right, with x horizontal and y vertical. The default is 0–1000 normalized at every image resolution. Explicitly setbbox_format="pixels"orbbox_format="normalized"for pixels or 0–1 coordinates. Do not mix units in one scene. depth_mis a positive camera-to-object distance estimate in meters. Preserve uncertainty in the answer; visual depth estimates are not calibrated measurements.- World coordinates use +Y up, +X right, and -Z forward from the first camera. BEV uses
(x,z).pos_bev=[x,z]andyaw_degdescribe each camera in that common world frame. - The scene assumes a 60° horizontal field of view and planar camera motion. Pitch is recorded but does not change the planar geometry. Large pitch or uncertain motion makes the BEV less reliable.
Choose the analysis
| Question | Tools and checks |
|---|---|
| Relative layout | visualize_bev(scene, viewpoint={"frame":N}) shows the selected camera's forward/right axes. |
| Virtual viewpoint | Use viewpoint={"at":"door","facing":"table"} or facing_away for the opposite direction. Resolve an object's actual facing before using its perspective. |
| Distance | triangulate(scene,target,frame_a,frame_b) needs the same stationary object and a real camera baseline. Check reliable; parallel, opposed, or backward rays cannot establish a reliable position. |
| Known-size calibration | calibrate_scale(scene,target,known_size_m,dim) returns a scale factor to apply to distance estimates. |
| Object motion | object_world_motion compares estimated world positions after accounting for supplied camera poses. Recheck correspondence and camera estimates before declaring movement. |
| Camera motion | camera_motion summarizes the poses already in the scene. It is not independent evidence validating the original motion estimates. |
| Facing | orient_facing(image,target) uses VLM calls. Inspect the full image for body/object orientation. |
| Count and identity | count_objects counts grounded instances; match_entities groups sightings across views. Inspect repeated nearby objects before accepting deduplication. |
| Choose frames | select_keyframes supports uniform, motion, coverage, and covisibility strategies. |
| Temporal changes | scene_map provides cognitive maps, appearance order, frame differences, and VLM event localization. |
| Visibility | view_reason reasons from an object's viewpoint. Sparse boxes cannot establish a complete occlusion model. |