DEV Community

Divyansh Saharan
Divyansh Saharan

Posted on AI-assisted

GuardMate: 'Leave it with the guard' shouldn't need another phone call

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

What I Built

My friend and I physically go to the office. I live at home, but he is from far away, so he stays in a PG—a paying guest accommodation.

He also orders quite a few things online.

That combination creates a surprisingly repetitive problem: a courier arrives while he is at work, finds his room locked, and calls to ask where to leave the parcel.

Usually, his answer is the same: hand it to security at the guard room near the entrance.

But he cannot always take that call. Sometimes, the parcel ends up being returned. On weekends, he might actually be in his room, so repeating the same instruction regardless of his availability would not solve the problem either.

I built GuardMate, a personal AI delivery assistant around his routine.

It uses saved PG directions, guard-room instructions, weekly availability and a daily override to guide a courier through a delivery conversation. Each courier gets a separate session, so one parcel’s confirmations and approvals cannot accidentally carry over to another.

The challenge was not making an AI say “guard room.” It was teaching the system when not to say it:

  • Is the parcel prepaid?
  • Is security actually available to receive it?
  • Does this delivery require payment, an OTP or a signature?
  • Is the courier proposing somewhere the resident has not approved?

GuardMate can provide directions, remember confirmed facts and request a resident decision. It does not treat a courier’s suggestion as permission.

Demo

Watch the GuardMate product demo — MP4

Listen to the call recording — MP3

Use GitHub’s download control if playback is unavailable. The recording is shared with all speakers’ permission.

There is no public deployment. The application runs locally, and the cellular experiment uses Windows Phone Link with an Android phone.

In one consenting fictional courier call on a vivo T2x 5G running Android 15, we demonstrated the complete path:

Caller audio → local Whisper → hosted Qwen → permission checks → local Piper → caller
Enter fullscreen mode Exit fullscreen mode

The caller confirmed hearing the full greeting and three checked replies.

The most useful moment was a failure: the final speech transcript was ambiguous. Qwen proposed recording a delivery outcome, but GuardMate’s application checks rejected the unsupported completion and asked for clarification instead.

No parcel actually changed hands, and no completed delivery was recorded.

That test used manual answering and transcript review. Automatic conversation after manual answering is now implemented with offline regression coverage, but it has not yet been demonstrated on a live call. Automatic answering and hangup are not implemented.

Read the actual live-call test report.

Code

GuardMate on GitHub

The repository includes setup instructions, the delivery dataset, training and evaluation tools, test reports, and the demo recordings.

How I Built It

The dashboard uses React, TypeScript and Vite. The backend uses FastAPI and Pydantic, with SQLite for resident settings, conversation history and approvals.

At the AI core is the open-weight Qwen3.5-4B, accessed through Tinker.

This is not a single-message classifier. The planner receives conversation history, saved instructions, current availability, confirmed parcel facts and resident-decision state. It returns a structured action plan.

GuardMate then checks that plan before executing it.

For example, an alternative handoff location requires an explicit, expiring resident approval. A courier cannot authorize it by claiming to be the owner. Changes to resident instructions invalidate stale decisions. Delivery outcomes remain labelled courier-reported, not independently verified receipt.

For voice, whisper.cpp performs local transcription and Piper generates speech locally. Capture and playback happen in separate phases to avoid deliberately recording the assistant’s own reply.

Fine-tuning Qwen with Tinker

I used Tinker for a real delivery-specific LoRA training pilot, not just inference.

The pilot used:

  • 15 owner-reviewed fictional training scenarios.
  • 40 structured planner targets.
  • LoRA rank 16 and batch size 4.
  • Three epochs, producing 30 optimizer updates.

The targets teach the next planning action—not just a list of phrases to repeat.

I then ran a fresh, matched comparison between base and tuned Qwen on six development scenarios:

Metric Base Tuned
Strict planner agreement 7/13 — 53.85% 8/13 — 61.54%
Checked scenario success 5/6 5/6
Median text-model/policy latency 2.133 s 3.113 s

Fine-tuning produced one additional correctly matched planner turn: a 7.69-percentage-point improvement on this small rubric.

The limits matter. Checked scenario success did not improve, and median latency worsened. These are fictional, exposed development cases—not evidence of real-world accuracy or generalization.

The recorded live call used base Qwen; the tuned model was evaluated separately. Completing training did not automatically promote the checkpoint.

Read the training recipe, matched results and remaining errors.

Why Does Open Innovation Matter?

For GuardMate, openness means being able to work on the pieces that fail.

During the call test, Whisper sometimes recognized “parcel” as “passion” or “option.” Because the speech pipeline is local and inspectable, I can investigate capture timing and compare installed speech models without replacing the entire application or uploading raw audio to a cloud speech API.

Open-weight Qwen also gave me a model I could adapt to this narrow task through Tinker, then compare against its base version under the same application checks.

That combination gives me three useful kinds of control:

  • Data control: routine speech processing stays local.
  • Model control: I can fine-tune the planner and evaluate a specific checkpoint.
  • Behavior control: permission-sensitive actions remain inspectable application logic rather than an instruction buried inside a prompt.

Closed services can offer fine-tuning too. My reason for choosing this approach was the replaceable open components and the ability to inspect and change the system—not a claim that every closed model would perform worse.

There is also an important privacy boundary: GuardMate is not fully offline. Recognized text, previous messages and saved resident context go to Tinker for planning. Local speech avoids hosted audio processing, not all hosted data sharing.

The prototype still needs better transcription and a real handover test with my friend. I have not claimed that it has reduced returned parcels yet.

The goal is small and personal: help him receive his deliveries without repeatedly interrupting his workday.

Prize Categories

Best Use of Tinker

GuardMate uses Tinker to fine-tune open-weight Qwen for delivery-conversation planning and evaluate it against a fresh, matched baseline.

The measured improvement is in strict planner agreement: 53.85% → 61.54%. The small sample, unchanged checked delivery behavior and slower median latency are reported alongside that gain.

Top comments (0)