If you want your portfolio project to stand out, you need to do more than just wrap a commercial API. It’s important to know how to load, run, and optimize models on your own machine. In this tutorial, I’ll show you how to build a Vision AI app with Python and a strong open-source vision-language model.
We’ll use Florence-2 by Microsoft. It’s a powerful model that works well even if you don’t have high-end server hardware.
How Vision-Language Models (VLMs) Actually Work
Older computer vision systems needed different models for tasks like object detection, text reading (OCR), and scene description. Modern Vision-Language Models can do all of these jobs with one unified system.
No matter what computer vision task you want to do, Florence-2 treats it as a sequence-to-sequence problem. You give it an image and a text prompt, and it returns text as the output. Behind the scenes, it uses a DaViT vision encoder to turn your image into visual data, and a text encoder to process your prompt. Then, a transformer model combines these to create the final answer.
Florence-2 is accurate and flexible because of its training data. It was trained on the huge FLD-5B dataset, which has over 5.4 billion labels on 126 million images. So, right away, it can find objects, read documents, and describe complex scenes.
Building Your Vision AI App with Python
Now, let’s get started with the code. We’ll use the Hugging Face transformers library to load the model and handle our inputs.
Step 1: Install the Dependencies
First, set up your Python environment with these free and open-source libraries:
pip install transformers torch torchvision pillow requests accelerate
Step 2: Initialize the Model and Processor
import torch
from transformers import AutoProcessor, Florence2ForConditionalGeneration
from PIL import Image
import requests
# Load the model and processor
model_id = "florence-community/Florence-2-base"
model = Florence2ForConditionalGeneration.from_pretrained(
model_id,
device_map="auto"
)
processor = AutoProcessor.from_pretrained(model_id)When running in production, it’s important to manage your hardware resources. We use device_map="auto" so the library can automatically spread the model weights across your GPU and CPU memory.
If you’re interested in building more production-ready AI apps, check out my book, Hands-On GenAI, LLMs & AI Agents. It covers LLMs, RAG, AI Agents, and multimodal AI with hands-on projects.
Step 3: Load and Prepare the Image
# Load a sample image from the web
url = "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/tasks/car.jpg"
image = Image.open(requests.get(url, stream=True).raw).convert("RGB")One mistake I’ve made, and often see in deployment scripts, is not handling image alpha channels. If you give a vision model an RGBA image (like a transparent PNG) when it expects three color channels, you’ll get a tensor shape mismatch error and your app will crash. Always convert images to RGB format.
Here’s the image we are using:

Step 4: Define the Task and Process Inputs
Florence-2 uses specific prompts to decide what task to do. In this example, we’ll use <OD> for Object Detection. You can also use <OCR> to pull text from a document or <CAPTION> for a general description:
task_prompt = "<OD>"
# Prepare the inputs for the model
inputs = processor(
text=task_prompt,
images=image,
return_tensors="pt"
).to(model.device)Step 5: Generate and Decode the Output
# Generate the output
generated_ids = model.generate(
input_ids=inputs["input_ids"],
pixel_values=inputs["pixel_values"],
max_new_tokens=1024,
num_beams=3
)
# Decode the raw text
generated_text = processor.batch_decode(
generated_ids,
skip_special_tokens=False
)[0]
# Parse the text into structured coordinates
parsed_answer = processor.post_process_generation(
generated_text,
task=task_prompt,
image_size=(image.width, image.height)
)
print(parsed_answer){'<OD>': {'bboxes': [[34, 160, 597, 371], [272, 241, 303, 247], [454, 276, 553, 370], [96, 280, 198, 371]], 'labels': ['car', 'door handle', 'wheel', 'wheel']}}
We set num_beams=3 to turn on beam search. This helps the model create more accurate text by trying out several possible answers before picking the best one.
After the model creates the token IDs, we turn them into plain text. Then, we use the processor’s post-processing function to change that text into a Python dictionary with the bounding box coordinates.
When you run this code, parsed_answer will give you the labels of the objects found in the image (like “car”, “tire”, “door”) and their exact coordinates. You can use this data in a frontend dashboard, save it to a database, or draw bounding boxes on the image with a library like matplotlib.
Recommended Courses to Strengthen Your Vision AI Skills
If you want to learn more about computer vision and Vision AI, here are two beginner-friendly courses you might like:
- Introduction to Computer Vision and Image Processing: Learn the fundamentals of image processing, feature extraction, object detection, and modern computer vision techniques used in real-world AI applications.
- Deep Learning for Computer Vision: Build practical deep learning skills for image classification, object detection, and vision-based applications using state-of-the-art neural networks.
These courses go well with this project. They’ll help you learn how to use Vision AI models and understand the ideas behind them.
Final Thoughts
When you build your own Vision AI app with Python and an open-source model like Florence-2, you get real experience with machine learning. You’ll learn how to work with tensors, manage memory, and turn raw outputs into useful data for your app.
Start with something simple. Copy the code above, change the image URL to a file on your computer, and try out different prompts. When you see the model pull out data on your own machine, you’ll see how easy it is to use production-level AI.
I hope you enjoyed this article about building a Vision AI app with Python.
For more AI and machine learning tips, follow me on Instagram. My book, Hands-On GenAI, LLMs & AI Agents, can also help you grow your AI career.





