DEV Community

Bhavika
Bhavika

Posted on

Monument Assistant for my travelsavvy friend

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This project was built for my travel-savvy friend who loves travelling and exploring India’s rich heritage who finds traditional museum plaques dry and standard image-search tools uninformative. So I built an interactive, personal, multi-turn AI tour guide right in their pocket.

What I Built

The Indian Monument Identifier & Interactive AI Guide is a web-based, memory-aware application that allows users to drag-and-drop a photo of any Indian historical landmark to instantly receive a structured, rich cultural guide.

Instant Visual Recognition: Identifies monuments from user-uploaded images without needing a pre-categorized or hardcoded database.

Structured Cultural Output: Generates formatted breakdowns covering Monument Name & Location, Built Era / Ruler, Architectural Style and Key Historical Facts.

Conversational Thread Memory: Maintains multi-turn session context, allowing the user to ask natural follow-up questions (e.g., "What is the best time of year to visit?" or "What other sites are nearby?") without re-uploading the photo

Code

How I Built It

The application bridges a Streamlit front-end with a Python async backend orchestrated by the Backboard API:

Frontend Interface (Streamlit):

Built a drag-and-drop file uploader accepting .jpg, .png, and .webp images.

Handles user inputs, displays live image previews, and renders Markdown response outputs seamlessly.

Backend Orchestration (Backboard SDK):

Assistant Initialization: Spawns a dedicated AI assistant configured with a persistent system_prompt acting strictly as an expert Indian historian.

Thread Management: Initializes a session thread (client.create_thread()) to store conversation history and visual context on Backboard’s servers.

Multimodal Routing: Passes the temporary local image path directly through Backboard (files=[file_path]) to multimodal vision models (gpt-4o or open-weight vision alternatives).

Open-Source AI & Framework Core:

Async Runtime & SDK: Powered by standard open-source Python packages (asyncio, streamlit, tempfile) and the backboard-sdk.

Model Agnosticism: Constructed around open-source agent integration frameworks, allowing the app to route queries across open-weight vision models (e.g., Gemma Vision variants) or commercial endpoints via Backboard’s unified gateway.

Why Does Open Innovation Matter?

Zero-Shot Flexibility vs. Closed Models: Closed vision APIs force you to use rigid, pre-categorized classifiers that only output flat text labels (e.g., Taj_Mahal). Open-source, multimodal AI enables zero-shot visual understanding, eliminating the need to collect, label, and train expensive custom datasets on thousands of monument photos.

No Vendor Lock-In via Backboard: Closed APIs lock you into proprietary SDKs. Using Backboard's open orchestration framework abstracts model provider logic—swapping underlying vision models (or comparing open-weight models) requires changing just a single parameter string (model_name) without rewriting thread memory or upload pipelines.

Democratizing Cultural Access: Open innovation lets developers build low-cost, high-impact tools that turn static historical plaques into personalized, interactive AI tour guides accessible to everyone without expensive subscription costs.

Prize Categories

Best Use of Backboard ($100 USD + Exclusive Winner Badge):
Built using Backboard’s unified API and SDK to manage multimodal visual inputs (files=[...]), maintain session thread context across user queries, and enforce system prompt guardrails for historical accuracy.

Top comments (1)

Collapse
 
sinarezaei profile image
Sina Rezaei •

The part I find more interesting than the monument recognition itself is the separation between the vision model and the conversation state.

Keeping the image context inside a thread means the model doesn't have to solve the same identification problem on every follow-up question. That makes the interaction feel more like an actual assistant instead of a sequence of independent vision calls.

I also like the model-agnostic setup. Swapping the vision model without rebuilding the upload and memory layer is a much more useful abstraction than simply having multiple models available.

One thing I'd be curious to test next is factual reliability. Historical monuments are a pretty good stress test for hallucinations because the output can sound extremely convincing even when a ruler, date, architectural detail, or attribution is wrong. A small retrieval/verification layer could make this much stronger for real travel use.