Documentation menu

Video-Spatio

model-free 3D spatial reasoning: a skill plus stateless geometry/VLM MCP tools (build_scene, triangulate, visualize_bev, camera_motion, orient_facing, ...) driven by the host model. No inner agent and no perception server.

Token estimates

Skill instructions
About 1,468 tokens
Tool definitions
About 8,144 tokens

Estimated text size—not usage or cost.

SKILL.md
Permalink

Video-Spatio — spatial reasoning

Use this Skill when a question needs spatial reasoning: distance, size, facing, left/right/front/behind, camera versus object motion, cross-frame counting, or viewpoint-dependent visibility. Answer appearance questions directly from the images. Read the qwen-mm-plugins-video-spatio tool schemas for arguments.

The host model supplies object boxes, distance estimates, and camera-motion estimates. Geometry tools calculate from those inputs; they do not independently measure the scene. No GPU perception server is required. VLM-backed tools use the configured OpenAI-compatible endpoint and may make several calls.

Build a scene

  1. Inspect the image or sample video frames. If installed, core provides read_video and save_view. Otherwise use available image/video tools and provide local frame paths.
  2. Ground the relevant objects in each frame, with one list per frame: {"label":"chair","bbox":[400,300,600,800],"depth_m":2.3}.
  3. Call build_scene(frames=..., objects_by_frame=...). All frames must have the same dimensions. Use unique frame_indices when retaining original video indices.
  4. For a moving camera, supply one camera_motions entry per adjacent pair. Translations are in the preceding camera's frame: forward_m is positive forward, right_m positive right, and yaw_deg positive for a right turn. Omitting motions assumes a static camera; it does not estimate motion.
  5. Pass the returned scene JSON to subsequent tools, or save it and pass scene_file.

Coordinates:

  • Boxes are [x1,y1,x2,y2], top-left to bottom-right, with x horizontal and y vertical. The default is 0–1000 normalized at every image resolution. Explicitly set bbox_format="pixels" or bbox_format="normalized" for pixels or 0–1 coordinates. Do not mix units in one scene.
  • depth_m is a positive camera-to-object distance estimate in meters. Preserve uncertainty in the answer; visual depth estimates are not calibrated measurements.
  • World coordinates use +Y up, +X right, and -Z forward from the first camera. BEV uses (x,z). pos_bev=[x,z] and yaw_deg describe each camera in that common world frame.
  • The scene assumes a 60° horizontal field of view and planar camera motion. Pitch is recorded but does not change the planar geometry. Large pitch or uncertain motion makes the BEV less reliable.

Choose the analysis

QuestionTools and checks
Relative layoutvisualize_bev(scene, viewpoint={"frame":N}) shows the selected camera's forward/right axes.
Virtual viewpointUse viewpoint={"at":"door","facing":"table"} or facing_away for the opposite direction. Resolve an object's actual facing before using its perspective.
Distancetriangulate(scene,target,frame_a,frame_b) needs the same stationary object and a real camera baseline. Check reliable; parallel, opposed, or backward rays cannot establish a reliable position.
Known-size calibrationcalibrate_scale(scene,target,known_size_m,dim) returns a scale factor to apply to distance estimates.
Object motionobject_world_motion compares estimated world positions after accounting for supplied camera poses. Recheck correspondence and camera estimates before declaring movement.
Camera motioncamera_motion summarizes the poses already in the scene. It is not independent evidence validating the original motion estimates.
Facingorient_facing(image,target) uses VLM calls. Inspect the full image for body/object orientation.
Count and identitycount_objects counts grounded instances; match_entities groups sightings across views. Inspect repeated nearby objects before accepting deduplication.
Choose framesselect_keyframes supports uniform, motion, coverage, and covisibility strategies.
Temporal changesscene_map provides cognitive maps, appearance order, frame differences, and VLM event localization.
Visibilityview_reason reasons from an object's viewpoint. Sparse boxes cannot establish a complete occlusion model.

Expand folders to explore bundled references, scripts, and assets. Files open at this page’s source snapshot.