Skills
Vision applications often need structured spatial results, such as bounding boxes or 2D points. Moondream exposes these common workflows as dedicated skills, so applications do not have to extract coordinates from free-form text.
Available Skills
Query
The most general-purpose skill. Ask questions about images and get intelligent answers. Great for:
- Visual Q&A systems
- Image analysis
- Content verification
Caption
Generate natural language descriptions of images. Perfect for:
- Creating alt text for accessibility
- Cataloging visual content
- Understanding image context
- Retail item descriptions
Point
Identify and locate specific elements within images by coordinates. Useful for:
- UI automation
- Interactive image annotations
- Precise element selection
Detect
Detect and identify objects, people, and elements in images. Ideal for:
- Object recognition
- Scene understanding
- Content moderation
Segment
Generate precise SVG path segmentation masks for objects. Perfect for:
- Image cutouts and masks
- Precise object selection
- Background removal
- Interactive editing