DEV Community

Cover image for Building an Intelligent Image Analyzer with Google Cloud Vision API and Cloud Run
Mukhtar Salim
Mukhtar Salim

Posted on

Building an Intelligent Image Analyzer with Google Cloud Vision API and Cloud Run

Modern digital platforms from e-commerce marketplaces to content hubs handle massive volumes of user-uploaded imagery daily. Manually reviewing images, writing catalog tags, and catching policy-violating content quickly becomes an operational bottleneck. While training custom computer vision models from scratch is an option, it requires massive labeled datasets, specialized hardware, and continuous model maintenance.

What is an Image Analyzer?

An Image Analyzer is a software system or application that processes visual data (digital images) to extract meaningful, structured information. Instead of treating an image as merely a collection of raw pixels or a static file, an analyzer uses computer vision techniques and machine learning models to "understand" what is inside the picture.

Depending on the implementation, an image analyzer can perform several tasks:

  • Object & Entity Detection: Identifying items, landscapes, people, or products.

  • Semantic Labeling: Generating descriptive keywords and tagging the image (e.g., identifying "beach", "ocean", or "sunset").

  • Content Moderation: Detecting unsafe, sensitive, or policy-violating imagery (e.g., adult content, violence, or hate symbols).

  • Metadata Extraction: Extracting dominant color palettes, aspect ratios, image clarity, and visual properties.

What is Google Cloud Vision API?

Google Cloud Vision API is a fully managed computer vision service provided by Google Cloud. It allows developers to integrate pre-trained machine learning models into their applications through simple REST or gRPC API calls.

Training computer vision models from scratch requires massive datasets, expensive GPU compute clusters, and continuous model fine-tuning. Cloud Vision API eliminates this barrier by giving you immediate access to Google’s state-of-the-art vision models.

Key capabilities include:

- Label Detection: Generating descriptive tags with associated confidence scores.

- SafeSearch Detection: Evaluating content likelihood across five moderation categories (adult, spoof, medical, violence, and racy).

  • Text Detection (OCR): Extracting printed or handwritten text from images.

- Landmark & Logo Recognition: Identifying well-known public structures and commercial brand logos.

- Image Properties: Evaluating dominant colors, aspect ratios, and visual quality.

What is Google Cloud Run?

Google Cloud Run is a fully managed serverless compute platform that runs containerized applications directly on top of Google’s scalable infrastructure.

With Cloud Run, you package your code, dependencies, and system runtime into an OCI-compliant container (such as a standard Docker container). Cloud Run then manages everything else:

- True Serverless Scaling (Scale-to-Zero): The service automatically spins up instances when requests arrive and scales down to zero when idle, meaning you do not pay for idle server compute.

- Granular Autoscaling: When traffic spikes, Cloud Run provisions new instances in seconds and supports multiple concurrent requests per container instance.

  • Zero Infrastructure Maintenance: There are no virtual machines to patch, no operating systems to update, and no Kubernetes clusters to configure.

- Native Security & IAM Integration: Authenticates securely with other Google Cloud services using Service Accounts and Application Default Credentials (ADC), removing the need to embed secret API keys in your code.

Why We Need Them in This Demo

In the Smart Tagger application, these three components work together to solve a specific production architecture challenge:

[ Image Upload ] ──> [ Cloud Run (FastAPI) ] ──> [ Cloud Vision API ] ──> [ Structured Tags & Flags ]
                      (Microservice Host)          (AI Inference Engine)
Enter fullscreen mode Exit fullscreen mode

Why an Image Analyzer?

Modern applications (like e-commerce platforms, social feeds, and digital asset managers) receive thousands of user-uploaded images. Businesses cannot afford to have human operators manually verify every photo, write catalog search tags, and check for offensive material. An automated Image Analyzer is the core business solution being built.

Why Cloud Vision API?

It provides the AI intelligence without requiring custom model training. It processes complex image labeling, dominant color extraction, and content moderation in a single network round-trip with high accuracy and low latency.

Why Cloud Run?

Cloud Vision API is an external service....it cannot serve user uploads or enforce business rules on its own. Cloud Run acts as the secure, scalable host for your API microservice (built with FastAPI). It:

Provides an authenticated boundary between the client and Google Cloud using IAM (no exposed credentials).

Validates incoming file types and size limits before invoking paid AI calls.

Scales instantly to absorb sudden spikes in user uploads while scaling to zero during periods of inactivity, keeping operational costs low for developer communities and startups.

Getting Started

  1. Accepts a user-uploaded photo.

  2. Uses Vision API to generate hashtags automatically. (This implies using an object/label detection or similar feature).

  3. Filters out "unsafe" images before they are published. (This uses the safe search detection feature of the Vision API).

To create this application, we need to build three main parts:

  1. Frontend (HTML/CSS/JavaScript): To allow users to upload a photo and display the results.

  2. Backend (e.g., Python/Node.js): To handle the file upload and communicate with the Google Cloud Vision API.

  3. Google Cloud Project & Vision API: Set up the necessary cloud resources and credentials.

Application Blueprint: The "Smart Tagger"

A. Prerequisites (What you need to set up)

  1. A Google Cloud Project: Create a new project.
  2. Enable the Vision API: Go to "APIs & Services" and enable the "Cloud Vision API."
  3. Set up Authentication: For a server application, the best method is to create a Service Account and download the JSON key file.

1. You will need to set an environment variable to point to this file

(e.g., export GOOGLEAPPLICATIONCREDENTIALS="/path/to/your/keyfile.json").
Enter fullscreen mode Exit fullscreen mode

Install Python Libraries:

1.  pip install Flask google-cloud-vision
Enter fullscreen mode Exit fullscreen mode

2. Backend Code (Python using Flask)

Create a file named app.py. This code handles the upload, calls the Vision API, and checks for safety.

Python

import os

from flask import Flask, request, jsonify, render\_template

from google.cloud import vision
Enter fullscreen mode Exit fullscreen mode

Initialize Flask app

app = Flask(\_\_name\_\_)
Enter fullscreen mode Exit fullscreen mode

Vision API Functions

def analyze_image_with_vision(image_path):
    """Detects labels and checks for explicit content in the image."""

    client = vision.ImageAnnotatorClient()

    with open(image_path, 'rb') as image_file:
        content = image_file.read()

    image = vision.Image(content=content)

    # 1. Label Detection for Hashtags
    label_response = client.label_detection(image=image)
    labels = label_response.label_annotations

    hashtags = [
        f"#{label.description.replace(' ', '').replace('/', '')}"
        for label in labels[:5]
    ]  # Take top 5 labels

    # 2. Safe Search Detection for Filtering
    safe_response = client.safe_search_detection(image=image)
    safe = safe_response.safe_search_annotation

    # Check for likely or very likely unsafe content
    is_unsafe = False

    # Vision API Safe Search categories and their thresholds
    # for "Unsafe"
    # LIKELY (3) or VERY_LIKELY (4) are generally considered unsafe.
    unsafe_threshold = 3

    if (
        safe.adult >= unsafe_threshold
        or safe.violence >= unsafe_threshold
        or safe.racy >= unsafe_threshold
        or safe.medical >= unsafe_threshold
    ):
        is_unsafe = True

    return hashtags, is_unsafe
Enter fullscreen mode Exit fullscreen mode

Flask Routes

@app.route('/')

def index():

    """Renders the upload page."""

    return render\_template('index.html')

@app.route('/upload', methods=\['POST'\])

def upload\_file():

    """Handles the file upload and analysis."""

    if 'file' not in request.files:

        return jsonify({'error': 'No file part'}), 400

    file = request.files\['file'\]

    if file.filename == '':

        return jsonify({'error': 'No selected file'}), 400

    if file:

        # Save the file temporarily

        temp\_file\_path = 'temp\_uploaded\_image.jpg'

        file.save(temp\_file\_path)

        try:

            # Analyze the image

            hashtags, is\_unsafe = analyze\_image\_with\_vision(temp\_file\_path)

            # Prepare result

            result = {

                'status': 'success',

                'hashtags': hashtags,

                'is\_unsafe': is\_unsafe,

                'message': 'Image analyzed successfully.'

            }

            if is\_unsafe:

                result\['message'\] = '⚠️ \*\*UNSAFE IMAGE DETECTED.\*\* Filtering image from publication.'

                result\['hashtags'\] = \[\] # Clear hashtags if unsafe

            else:

                result\['message'\] = '✅ Image is safe. Hashtags generated.'

            return jsonify(result)

        except Exception as e:

            return jsonify({'status': 'error', 'message': f'An error occurred during analysis: {e}'}), 500

        finally:

            # Clean up the temporary file

            if os.path.exists(temp\_file\_path):

                os.remove(temp\_file\_path)

if \_\_name\_\_ == '\_\_main\_\_':

    # You must have the GOOGLE\_APPLICATION\_CREDENTIALS environment variable set

    if not os.environ.get('GOOGLE\_APPLICATION\_CREDENTIALS'):

        print("!!! WARNING: GOOGLE\_APPLICATION\_CREDENTIALS environment variable is not set. !!!")

        print("Please set it to the path of your Vision API Service Account key file.")

    app.run(debug=True)
Enter fullscreen mode Exit fullscreen mode

3 Frontend Code (HTML/JavaScript)

Create a folder named templates and inside it, a file named index.html. This provides the user interface.

Demo: The "Smart Tagger"
A simple web application that automates image tagging and safety filtering using the Vision API.

Analyze Image

Analysis Results

Upload an image to start the analysis.

function uploadImage() {
    const fileInput = document.getElementById('imageUpload');
    const file = fileInput.files[0];
    const statusMessage = document.getElementById('statusMessage');
    const safetyBadge = document.getElementById('safetyBadge');
    const hashtagsDisplay = document.getElementById('hashtagsDisplay');

    // Reset previous results
    statusMessage.textContent = 'Analyzing...';
    safetyBadge.className = '';
    safetyBadge.textContent = '';
    hashtagsDisplay.innerHTML = '';

    if (!file) {
        statusMessage.textContent = 'Please select a file first.';
        return;
    }

    const formData = new FormData();
    formData.append('file', file);

    fetch('/upload', {
        method: 'POST',
        body: formData
    })
        .then(response => response.json())
        .then(data => {
            statusMessage.textContent = data.message;

            if (data.status === 'success') {
                if (data.is_unsafe) {
                    safetyBadge.className = 'unsafe';
                    safetyBadge.textContent = 'Status: UNSAFE (Filtered)';
                } else {
                    safetyBadge.className = 'safe';
                    safetyBadge.textContent = 'Status: SAFE (Published)';

                    data.hashtags.forEach(tag => {
                        const span = document.createElement('span');
                        span.textContent = tag;
                        hashtagsDisplay.appendChild(span);
                    });
                }
            } else {
                safetyBadge.className = 'unsafe';
                safetyBadge.textContent = 'ERROR';
                console.error(data.message);
            }
        })
        .catch(error => {
            statusMessage.textContent = `An error occurred: ${error.message}`;
            safetyBadge.className = 'unsafe';
            safetyBadge.textContent = 'NETWORK ERROR';
            console.error('Fetch error:', error);
        });
}
Enter fullscreen mode Exit fullscreen mode

4 How to Run the Application

Set Environment Variable: Open your terminal and set the path to your service account key file.

export GOOGLEAPPLICATIONCREDENTIALS="/path/to/your/vision-api-key.json"
Enter fullscreen mode Exit fullscreen mode
  1. Run the Flask App: In the directory where you saved app.py, run:python app.py

  2. Open in Browser: Open your web browser and navigate to the address shown in the terminal (usually http://127.0.0.1:5000/).

Top comments (0)