Multi-modal models accept video, audio, and images. Engineering multi-modal prompts requires referencing specific timestamps (e.g. 'At 0:15 in the video...') and bounding text contexts (like transcriptions) alongside the media file to align reasoning across modalities.
You are a traffic audit assistant. Analyze the attached video alongside the GPS metadata log below.
GPS Metadata:
Task: Identify the cause of the stop. Focus your visual analysis on the upper-right quadrant of the frame between timestamps 0:10 and 0:15.
Claude 3.5 Sonnet supports high-resolution image inputs. Use clear spatial coordinate terms (e.g. 'lower-left quadrant') in prompts.