Overview
Jumper uses four types of machine learning models to help you understand and search through your footage:Visual search models
AI systems that can “see” and understand images and video frames, then connect that understanding to text queries or other images.
Speech models
AI systems that can transcribe spoken words into searchable text.
Face detection models
AI systems that can identify and group similar faces together, allowing you to search for specific people.
Summary Analysis models
AI systems that turn analyzed media into layered, navigable descriptions for editorial research and AI-assisted workflows.
Visual Search Models
Visual search models are AI systems that can “see” and understand images and video frames, then connect that understanding to text queries or other images. This is what makes Jumper’s visual search work. How visual search models work: These models learn to find correlations between text and visual data from their training material. When you search for something like “a person walking through a door,” the model has learned to understand:- What “a person” looks like visually
- What “walking” means in terms of motion and pose
- What “a door” is and how it appears in different contexts
- How these elements relate to each other in a scene
- Text search: Enter a natural language query to find matching visual content
- Image search: Use an image or frame as your search input instead of text. The model uses similar algorithms to find visually similar content across your footage
- You can search using natural, conversational language
- You can search using images or frames from your footage
- The model understands context and relationships between objects
- It works across different languages (multilingual models)
- Higher resolution models can detect finer details like text on signs or small objects
- Size: Larger models generally offer better accuracy but require more memory and processing power
- Resolution: Higher resolution models (384×384, 512×512) analyze frames in more detail, useful for detecting text, small objects, or fine visual details
- Speed: Smaller models analyze faster, while larger models take more time but provide better results
Speech Models
Jumper uses different local speech-to-text models for different languages. Select the language spoken in your footage before analysis, and Jumper routes it to the most appropriate model for your platform.How languages are routed
Languages on the same line use the same speech model:- English
- Swedish (shoutout kb-whisper)
- Japanese
- Moroccan Arabic (Darija)
- Arabic, Cantonese, Chinese (including Mandarin), Filipino/Tagalog, Hindi, Korean, Malay, Persian, Russian, Thai, and Vietnamese
- All other supported languages
The specialized Swedish, Japanese, and Moroccan Arabic models achieve state-of-the-art word error rates (WER) for their target languages. Lower WER means fewer transcription errors. The shared multilingual base model is also highly competitive across its supported languages.
How speech analysis works
The selected speech model listens to your media’s audio track and converts speech into text during analysis. Once the media is analyzed, you can search for any word or phrase that was spoken. Key features:- Supports more than 100 languages and language variants
- Selects a model based on the language you choose
- Handles different accents, background noise, and audio quality
- Transcribes dialogue with timestamps
- Works entirely offline with no cloud processing required
Face Detection Models
Face detection requires a Jumper Pro license
- Detect faces in the frame
- Group similar faces together, even when lighting, angles, or expressions change
- Automatically identify all appearances of a person across your entire footage library
- Search for specific people using the
@syntax (e.g.,@John sitting on a bench) - Organize people into Collections to keep different productions separate
- Find every scene where someone appears, even if they’re in the background
Summary Analysis Models
Summary Analysis creates a concise whole-clip overview followed by progressively more detailed sections connected to source time ranges. This gives editors and AI agents context before they perform targeted searches. The model is downloaded and runs locally. Choose from Small, Medium, Large, or Extra large in the Summary Analysis configuration. Visual analysis is required; optional speech and face analysis can make the result more useful.Choosing the Right Visual Model
Depending on your specific workflow, usecase and hardware setup, you might want to choose a specific model other than the default. Some suggestions are listed below, or feel free to explore from all the available models.Fastest
Quick analysis with good accuracy, 256x256 resolution. Great for most workflows.
Fast
Larger model with higher accuracy, 256x256 resolution.
Accurate
Top-tier accuracy with 384×384 resolution.
Benefits from more detailed and verbose searches, great multilingual support.
Most Accurate
Highest accuracy with 512x512 resolution among the V1/V2 presets.
Benefits from more detailed and verbose searches, great multilingual support.
Ultra Accurate
State-of-the-art visual search quality using a different model architecture than the presets above. Available on Apple silicon Macs and Windows. Requires a separate download, more compute, and produces larger analysis files. Choose this when search quality matters more than analysis speed.

