DEV Community

Cian
Cian Subscriber

Posted on AI-assisted

Floor: the number you meant to hold

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.

I did not start this weekend building Floor.

I started with another project, Pause. I had a plan, a model to train, and the hope that the next run would make the idea come together. I put time into it. Then I put more time into it.

Eventually, I had to admit something frustrating: I was trying to force a solution.

There is a particular disappointment in looking at work you have already done and deciding it should not be the thing you ship. You want the effort to count. You want one more experiment to justify the hours before it.

But the deadline was getting closer, and I wanted to build something someone could actually find useful. I archived Pause.

After that, I met a friend over the weekend. Our conversation turned to the online marketplace they run, and the work of dealing with buyer inquiries.

They were getting plenty of messages. The difficult part was negotiating well enough for those conversations to become clear agreements. Someone still had to answer the questions, understand the offers, and decide how far to move on price.

I had been so caught up in getting my own idea to work that it was a relief to have a specific person's problem in front of me. Their situation gave me somewhere to start again.

I kept thinking about the person on the other side of that inbox. How many times can you explain the same thing, reconsider your price, and try to keep a conversation alive before it begins to wear on you?

I left that conversation with a clearer idea of what I wanted to build for my friend: a little support between the buyer's message and the seller's reply.

That became Floor.

What I Built

“What's your best price?”

It is such a small message. Answering it can take more thought than writing the listing.

The seller already has an asking price. Somewhere behind it is another number: the amount they can afford to accept. But knowing that number and holding it through a conversation are different things.

One buyer wants delivery included. Another wants to pay half now and the rest next week. Someone else comes back to an old conversation, and the seller has to remember what they already offered.

That was the part of my friend's work I wanted to help with. Every inquiry creates more work before it creates a sale. And although a buyer sees one conversation, the person answering may be keeping track of several.

I kept coming back to a simple question: could I help my friend answer the next message without making them rethink the whole deal?

So I built Floor, a seller inbox agent that keeps the listing, the seller's terms, and separate buyer conversations together. It reads a message, checks the context, and prepares a next move: answer, accept, counter, clarify, or close.

The seller sees the interpreted terms, corrects anything misunderstood, and edits the reply before copying it to their marketplace or chat. After sending it there, they record it as sent in Floor.

I liked the name because it points to something the seller already knows. Floor helps them carry that decision into the next conversation. For my friend, I wanted it to mean a little less second-guessing and a little more room to get on with the rest of the work.

How I Built It

After the first project, I wanted a more concrete reason to use fine-tuning. I needed to be able to point to a message and say: this is the distinction the model should learn.

An offer of KES 11,000 sounds promising. But KES 11,000 including delivery leaves the seller with a different amount. KES 6,000 today and KES 5,000 later is a different payment arrangement again.

A fluent reply would mean very little if the assistant misunderstood the deal.

I used Qwen3-8B, fine-tuned through Tinker, and gave it a narrow job: extract the total offered, the payment type, the costs the seller would pay, and the buyer's intent.

Training combined human buyer messages from CraigslistBargains with generated scenarios covering delivery ambiguity, payment conditions, and different ways of making offers. The human messages have explicit task-specific seller-policy assumptions; the generated examples are marked as generated. Preparation produced 396 training scenarios, rendered with two prompt variants, and 100 validation scenarios.

The application then uses those interpreted terms to calculate prices. It also checks saved listing facts and previous seller quotes. The suggested wording comes from code, and the interpreted terms remain visible for review. FastAPI serves the application; SQLite keeps the desk persistent.

The next run was better. The last one wasn't the best.

I compared the base model and fine-tuned checkpoints on the same 20-case development probe:

Model and prompt Exact interpretation Correct action and price with the shared pricing function
Base Qwen3-8B, detailed prompt 5/20 12/20
Base Qwen3-8B, compact prompt 9/20 16/20
Fine-tuned Qwen3-8B, detailed prompt, three passes 15/20 19/20

The pricing function was identical for every model. Better interpretations gave that same function better inputs.

Seeing the improvement was encouraging. It gave me a reason to keep building the product around this task.

I tried six training passes, hoping more training would help. The three-pass checkpoint was the best on this probe. Six passes finished at 18/20 decisions, so I selected the earlier checkpoint. The six-pass run cost approximately $0.80 in estimated training-token charges, with roughly $0.01 for its checkpoint comparisons, excluding storage.

There was still a mistake I could not ignore: an installment offer interpreted as full payment. These are small development results used to select the checkpoint, not an independent accuracy benchmark. They do not tell me how many sales Floor will help someone make. The training notes document the fuller comparison.

That remaining error helped shape the interface. The seller needs to see what Floor heard before trusting what it suggests.

Demo

The moment in the demo that made me think back to my friend's inbox was a buyer agreeing to an earlier quote.

Watch the 28-second interface walkthrough. It shows saved sample conversations and analyses produced by the selected Tinker model during the demo checks.

Floor's seller inbox, showing a sample conversation and a suggested reply

In the sample, the headphones are listed at KES 12,000, with a private floor of KES 10,000. Kevin offers KES 10,000 in full and will collect. Floor prepares a KES 11,000 counteroffer.

The reply is edited and recorded as sent. Then Kevin agrees to KES 11,000 in full.

Floor recognizes the earlier quote and recommends accepting. The conversation can move forward without another counteroffer.

That is a small piece of memory, but it matters. A seller should not have to reconstruct the whole conversation every time a buyer comes back.

Other sample inquiries ask about condition and availability, include delivery costs, or propose installments. All are labeled examples, not actual customer conversations.

The desk saves listing details, message histories, edited drafts, and inquiry states. I checked that they survived browser refreshes and server restarts. It also works on mobile. I wanted the reply someone had spent time editing to still be there when they returned.

The current version handles one seller and one listing. Messages are pasted in manually, and replies are sent through the seller's existing chat. My friend's handover and feedback are still ahead; I have not measured actual sales or conversion.

Code

Explore the Floor repository.

It includes the application, prompts, generated scenarios, data preparation and training code, attribution, and deployment instructions. Credentials, seller conversations, and private checkpoint identifiers stay outside the repository.

Why Does Open Innovation Matter?

This weekend, being able to change the model mattered to me.

When a message was misunderstood, I could inspect the failure, change the examples, fine-tune the open-weight model through Tinker, and compare the resulting checkpoints with the base model. I had something concrete to work on when an interpretation was wrong.

Tinker made remote training and sampling practical without setting up a local GPU training system. The application keeps readable rules for prices, listing facts, and approved quotes. Those decisions remain inspectable alongside the model's interpretation, and the model can be changed as the project develops.

Inference runs on Tinker, so buyer messages and listing terms are sent there. The local database keeps conversation history, and credentials stay on the server.

Prize Categories

Best Use of Tinker: task-specific fine-tuning of Qwen3-8B, with a recorded improvement over the baseline using the same pricing policy and development cases.

I began the weekend hoping a model would make my first idea work. A conversation with a friend gave me a clearer reason to keep building.

Now I want to put Floor in their hands and find out whether it helps during an ordinary afternoon of selling. Do the replies sound like them? Are the terms easy to check? Does holding a price become a little less tiring?

I hope it gives them a moment to breathe before answering. A place to remember what they already decided. A reply they feel comfortable standing behind.

Archiving Pause felt disappointing. Building something with my friend in mind gave the rest of the weekend a direction.

For now, Floor is a place to keep the number you meant to hold—and a little backup for the next message.

Top comments (0)