Skip to main content

AI courses

Prompt Engineering for Vision Models

A hands-on introduction to prompting and controlling vision models for image segmentation, object detection, image generation, in-painting, and personalized image generation.

Provider
DeepLearning.AI
Level
Beginner
Duration
1 hour, 22 minutes
Instructor
Abby Morgan, Jacques Verré and Caleb Kaiser
Format
Online, Self-paced, Video lessons, Interactive code examples

About

Prompt Engineering for Vision Models is a beginner-level short course from DeepLearning.AI in collaboration with Comet. It extends prompt engineering beyond text-based models and demonstrates how vision models can be controlled using natural language, pixel coordinates, bounding boxes, segmentation masks, and generation parameters. Learners work with technologies including Meta's Segment Anything Model, OWL-ViT, Stable Diffusion, and DreamBooth while exploring image segmentation, object detection, image generation, in-painting, fine-tuning, and experiment tracking.

View course on provider website

Opens an external website in a new tab.

Learning Outcomes

  • ✓Understand how prompting differs across text and vision models
  • ✓Prompt vision models using text, coordinates, and bounding boxes
  • ✓Adjust image generation parameters such as guidance scale, strength, and inference steps
  • ✓Use image segmentation models to isolate parts of an image
  • ✓Use natural-language prompts for zero-shot object detection
  • ✓Combine segmentation, object detection, and image generation for in-painting
  • ✓Understand how DreamBooth can personalize diffusion model outputs
  • ✓Use experiment tracking to compare visual prompting and hyperparameter configurations

Syllabus

1. Introduction

Introduces visual prompt engineering and the goals of the course.

2. Overview

Explains the vision models, prompting methods, and workflows used throughout the course.

3. Image Segmentation

Explores prompting image segmentation models with positive and negative coordinates and bounding boxes.

4. Object Detection

Uses natural-language prompts with object detection models to locate and isolate objects within images.

5. Image Generation

Explores diffusion-based image generation and parameters such as guidance scale, strength, and inference steps.

6. Fine-tuning

Introduces personalization of diffusion models using DreamBooth and demonstrates how fine-tuning can provide greater control over generated images.

7. Conclusion

Reviews the main visual prompt engineering techniques and workflows covered in the course.

Skills Covered

🧠 Prompt Engineering🧠 Computer Vision🧠 Multimodal Prompting🧠 Image Generation🧠 Image Segmentation🧠 Object Detection🧠 Diffusion Models🧠 Fine-Tuning🧠 In-painting🧠 AI Personalization🧠 Experiment Tracking