DEV Community

Nazia Ahmed
Nazia Ahmed

Posted on AI-assisted

perspective

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

My friend is really good at taking photographs of me. I am terrible at taking photographs of her. i mean she says i dont get her sense of style and perspective.

That made me wonder: what does she see through her camera that I don't?
So I built perspective an AI-powered photography reflection tool that compares two collections of photographs:

Photos my friend took of me
Photos I took of my friend

Instead of simply judging which photographs are "better", the project tries to understand the different perspectives behind them.

The goal is to identify patterns in composition, framing, camera angle, distance, subject placement, environment, lighting, and visual mood, and turn those patterns into something understandable in human language.

Eventually, the system will produce insights such as:

Her visual perspective: She seems to give the surrounding environment an important role in the photograph.

My visual perspective: I tend to make her the clear center of attention and focus more tightly on her.

And then go one step further:

What can I learn from her eye?

The point isn't to teach someone to copy another photographer. It is to help them notice what they naturally overlook.

I also want the project to become a small memory-making tool, where the photographs and the insights around them can become a record of how two people saw the same moments differently.

Demo

The project currently runs locally as a Streamlit application.

The application allows the user to upload photographs from both photographers and send them to a locally running vision model for analysis.

A demo prototype will be added here:

Code

https://github.com/nazya-ux/perspective

How I Built It

The application is built with Python and Streamlit.

The core AI component is Qwen3-VL 4B, an open-weight vision-language model running locally through Ollama.

The current pipeline is:

Photograph
↓
Streamlit upload
↓
Ollama
↓
Qwen3-VL
↓
Visual interpretation
↓
Photographer's Eye

The model is prompted to analyze photographs in terms of composition, framing, camera angle, subject placement, background, lighting, candid/posed character, visual mood, and what the photographer appears to be paying attention to.

The next layer of the project will aggregate these observations across both collections and generate a comparison of:

Her perspective
My perspective
Visual philosophy
Emotional character suggested by the photographs
What each photographer tends to notice
What one photographer might learn from the other
A meaningful memory interpretation

The important part is that the AI isn't supposed to take the photograph for you.

It doesn't tell the photographer exactly where to stand or how to pose someone. It acts more like a visual mirror that helps the photographer understand their own habits and perspective.

Why Does Open Innovation Matter?

This project would be very different if the entire experience depended on a closed vision API.

Using an open-weight model running locally makes the project possible in a much more personal way.

These photographs can be deeply personal. They may contain someone's face, home, friends, locations, and private memories. With local inference, the photographs don't need to be sent to a remote commercial AI service just to understand their visual patterns.

Open models also make the project replaceable and experimentable.

I can change the vision model, experiment with prompts, modify the analysis pipeline, or eventually fine-tune the system without rebuilding the entire application around a proprietary API.

Most importantly, open innovation lets me build an AI system that is not trying to replace the photographer.

It helps someone understand how they already see.

My Agent Session

Top comments (0)