I Deployed an On-Device Vision AI App, Here’s How

A major change happening in the industry right now is the move from building bigger models to running strong models right on your own device. When I mentor new developers, I often notice they use paid cloud APIs for almost every computer vision task. However, using the cloud can lead to delays, privacy concerns, and unexpected costs. That’s why I optimized and deployed my Vision AI app, which runs on-device. In this article, I’ll walk you through how I did it using Microsoft’s Florence-2-base model to create a fully functional, locally hosted app.

Here’s How I Deployed an On-Device Vision AI App

If you want to move away from relying on the cloud and build AI that works directly on your device, this guide is for you.

Setting Up for Local Execution

Before starting to write the app, we need the right tools. For on-device inference, it’s best to use lightweight and open-source options. I picked Streamlit for the frontend since it lets me quickly build interactive UIs in Python. For the machine learning backend, I used PyTorch and the Hugging Face transformers library.

Here is the exact setup command I used:

pip install torch torchvision transformers pillow accelerate streamlit

Once installed, we set up our script and configure the basic Streamlit UI:

import torch
import streamlit as st
from PIL import Image
from transformers import AutoProcessor, Florence2ForConditionalGeneration

st.set_page_config(
    page_title="Local Vision AI",
    page_icon="👁️"
)

st.title("Local Vision AI")
st.write("Florence-2 running locally on your device.")

This setup creates a simple web interface that runs only on your own computer.

Want to build more practical AI applications like this? Hands-On GenAI, LLMs & AI Agents covers LLMs, RAG, and AI Agents through hands-on projects.

Model Loading and Memory Optimization

A common mistake I see is loading the machine learning model inside the main loop. In Streamlit, the script restarts whenever someone interacts with the UI. If you load a large model every time, your computer can freeze, run out of memory, and crash.

To fix this, I used Streamlit’s @st.cache_resource decorator. This makes sure the model loads only once and stays in memory for later use.

Also, you can’t always know what hardware your code will run on. Some people have NVIDIA GPUs, others use Apple Silicon (M1, M2, or M3), and some only have standard CPUs. I wrote the load_model function to automatically find and use the best hardware available:

@st.cache_resource
def load_model():

    if torch.cuda.is_available():
        device = torch.device("cuda")

    elif hasattr(torch.backends, "mps") and torch.backends.mps.is_available():
        device = torch.device("mps")

    else:
        device = torch.device("cpu")

    model_id = "florence-community/Florence-2-base"

    processor = AutoProcessor.from_pretrained(model_id)

    model = Florence2ForConditionalGeneration.from_pretrained(
        model_id,
        torch_dtype=torch.float16
        if device.type == "cuda"
        else torch.float32
    )

    model = model.to(device)
    model.eval()

    return processor, model, device

processor, model, device = load_model()

Why is this code important? Look at the torch_dtype part. If a CUDA GPU is found, I load the model in float16, which uses half the memory and speeds up inference without losing accuracy. If the code runs on a CPU or Apple’s MPS, I use float32 for stability. Setting model.eval() puts the model in evaluation mode and turns off layers like dropout that are only needed during training.

Processing the Image and Running Inference

Next, we need a way for users to upload an image and start the AI process. I used Streamlit’s file_uploader and changed the input to standard RGB format so it works with the model:

uploaded_file = st.file_uploader(
    "Upload an image",
    type=["jpg", "jpeg", "png", "webp"]
)

if uploaded_file:

    image = Image.open(uploaded_file).convert("RGB")

    st.image(
        image,
        caption="Input image",
        use_container_width=True
    )

Now it’s time for the AI to make predictions. Florence-2 is a prompt-based vision model, so it does different tasks depending on the text prompt you give it. For Object Detection, I use the <OD> token:

if st.button("Analyze Image"):

        task_prompt = "<OD>"

        with st.spinner("Running Florence-2 locally..."):

            inputs = processor(
                text=task_prompt,
                images=image,
                return_tensors="pt"
            )

            inputs = {
                key: value.to(device)
                for key, value in inputs.items()
            }

            with torch.inference_mode():

                generated_ids = model.generate(
                    input_ids=inputs["input_ids"],
                    pixel_values=inputs["pixel_values"],
                    max_new_tokens=256,
                    num_beams=1
                )

            generated_text = processor.batch_decode(
                generated_ids,
                skip_special_tokens=False
            )[0]

            result = processor.post_process_generation(
                generated_text,
                task=task_prompt,
                image_size=(image.width, image.height)
            )

        st.subheader("Detection Result")
        st.write(result)

Take note of torch.inference_mode(). You may have seen torch.no_grad() in older guides. While no_grad() stops PyTorch from saving gradients, inference_mode() does even more by skipping view tracking and version counters. It’s only for the forward pass, which makes your on-device app much faster and more efficient with memory.

After creating the raw token IDs, we turn them back into text. Then, we use the processor’s post_process_generation method to match the model’s bounding box coordinates to the size of the original uploaded image.

To execute your file, make sure to run:

streamlit run app.py

Recommended Courses for On-Device Vision AI

If you want to keep building on this code and learn more about optimizing models for local use, I recommend two programs:

  1. PyTorch for Deep Learning Professional Certificate: We used PyTorch a lot in this project to handle memory, data types, and inference modes, so this program is a great next step. It shows you how to optimize models, work with tensors, and deploy deep learning pipelines just like a machine learning engineer does in real projects.
  2. Deep Learning Specialization: If you want to really understand how vision-language models work, this is the top choice. It explains the basics of Convolutional Neural Networks (CNNs) and Transformers, so you can learn to train, debug, and compress your own computer vision models for edge devices.

Both of these programs help you move from running simple scripts on your computer to building AI applications that are ready for real-world use.

Final Thoughts

Building this app reminded me of an important lesson in AI engineering: bigger models aren’t always better. What really matters is efficiency.

When you begin working in AI, it’s easy to think that only huge language models in the cloud do the important work. But real engineering is about taking a small, open-source model like Florence-2, making it fit your local hardware, and serving it through a simple user interface.

I hope you enjoyed this article about how I set up an on-device Vision AI app.

If you want more tips on AI and machine learning, you can follow me on Instagram. My book, Hands-On GenAI, LLMs & AI Agents, is also a helpful resource for growing your AI career.

Aman Kharwal
Aman Kharwal

AI/ML Engineer | Published Author. My aim is to decode data science for the real world in the most simple words.

Articles: 2218

One comment

Leave a Reply

Discover more from AmanXai by Aman Kharwal

Subscribe now to keep reading and get access to the full archive.

Continue reading