A hands-on introduction to prompting and controlling vision models for image segmentation, object detection, image generation, in-painting, and personalized image generation.
Provider
DeepLearning.AI
Level
Beginner
Duration
1 hour, 22 minutes
Instructor
Abby Morgan, Jacques Verré and Caleb Kaiser
Format
Online, Self-paced, Video lessons, Interactive code examples
About
Prompt Engineering for Vision Models is a beginner-level short course from DeepLearning.AI in collaboration with Comet. It extends prompt engineering beyond text-based models and demonstrates how vision models can be controlled using natural language, pixel coordinates, bounding boxes, segmentation masks, and generation parameters. Learners work with technologies including Meta's Segment Anything Model, OWL-ViT, Stable Diffusion, and DreamBooth while exploring image segmentation, object detection, image generation, in-painting, fine-tuning, and experiment tracking.