Okay, so for ages, our weekly product meeting was a bit of a crapshoot when it came to user feedback. We'd have a Slack channel dedicated to early testers, a few support channels, and then just general chatter where users would often drop gold nuggets of ideas or frustrations. The problem? Getting that information consistently into a digestible format for our Monday morning meeting was a nightmare.
The Feedback Firehose Problem
Every week, before the meeting, I'd block out an hour, sometimes two, to try and sift through hundreds of messages. I'd scroll, highlight, copy-paste into a doc, and try to tag things myself: "Okay, that's a negative sentiment about onboarding. This one's a feature request for dark mode. Oh, and here's a bug report I almost missed." It was painfully manual, incredibly inconsistent, and frankly, I was probably missing half the important stuff. We'd go into meetings with anecdotal evidence or just a 'feeling' about what users wanted, rather than hard data. It felt like we were sailing blind, making decisions based on the loudest voice or whatever I'd managed to skim that morning.
What I Tried (and What Didn't Quite Stick)
My first attempt at automating this was incredibly naive. I thought, "Keywords! That's it!" So I tried some simple regex searches for terms like bug, feature request, broken, love, hate. You can imagine how that went. "This feature is broken right now" was easy, but "I'm absolutely loving the new dashboard, but it would be even better if it had X" got missed entirely for its sentiment, and the feature request was buried in praise. The nuance was completely lost. Plus, people don't always use precise terminology. "It's a bit clunky" is negative, but doesn't contain a keyword.
Then I looked at dedicated feedback tools. They're great, sure, but they meant asking users to go to another place to leave feedback. Our users were already comfortable dropping messages in Slack; forcing them elsewhere felt like an unnecessary barrier. We needed something that worked where the feedback already was.
I even fiddled with basic LLM summarization initially. I'd dump a thread into ChatGPT, ask for a summary. Better than nothing, but it didn't give me the structured data I craved: a clear sentiment score, or a distinct list of feature requests. It was still qualitative, still required me to interpret the summary.
What Actually Fixed It: Structured Extraction with an LLM
The real breakthrough came when I stopped thinking of LLMs as just summarizers and started viewing them as incredibly powerful, flexible data extractors. The key wasn't to ask it to summarize the feedback, but to extract specific pieces of information in a structured format.
Here’s the high-level flow that finally clicked:
Slack Export: First, I set up a simple Python script using the Slack API to pull messages from our designated feedback channels. I focused on specific channels like #beta-feedback and #customer-support-chat for the last week. I grabbed the text field and the user ID. For a rough estimate, this took me about an hour to get right, including setting up the Slack app and OAuth tokens.
-
The Prompt Engineering Magic: This was the most crucial part. My prompt had to be explicit about what I wanted and how I wanted it formatted. I spent a good half-day just refining this. Initially, I kept getting an error like "LLM did not return valid JSON" because I wasn't strict enough with my instructions. Adding clear examples (few-shot learning) and telling it exactly to output JSON was key. I mostly used gpt-4-turbo for this, with a temperature setting of 0.2 to ensure consistency.
python
prompt_template = """
You are an AI assistant designed to extract sentiment and feature requests from user feedback.
Analyze the following Slack message. Determine the overall sentiment (Positive, Negative, Neutral).
Identify any distinct feature requests or suggestions made by the user. If no feature request is present,
return an empty list for 'feature_requests'.Output your response STRICTLY as a JSON object with the following keys:
- 'sentiment': (string, one of 'Positive', 'Negative', 'Neutral')
- 'feature_requests': (list of strings, specific feature requests/suggestions)
Example 1:
User message: "I love the new UI update! It's so much cleaner. Though, it would be amazing if we could export reports as CSV directly instead of PDF."
JSON: {"sentiment": "Positive", "feature_requests": ["Export reports as CSV"]}Example 2:
User message: "The new search function is completely broken, can't find anything. This makes my workflow impossible."
JSON: {"sentiment": "Negative", "feature_requests": []}User message: """""{message}"""""
JSON:
""" -
Python Processing: I wrote a Python script to iterate through the exported Slack messages. For each message, it calls the OpenAI API (or whichever LLM provider you use) with the message interpolated into my strict prompt. I use requests to handle the API calls, catch potential json.JSONDecodeErrors if the LLM occasionally misbehaves, and then parse the JSON output.
python
import os
import json
import requestsOPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
OPENAI_API_URL = "https://api.openai.com/v1/chat/completions"def get_llm_extraction(message_text, prompt_template):
headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {OPENAI_API_KEY}"
}
payload = {
"model": "gpt-4-turbo-preview", # Or gpt-3.5-turbo if you're on a budget
"messages": [
{"role": "user", "content": prompt_template.format(message=message_text)}
],
"temperature": 0.2,
"response_format": {"type": "json_object"} # Crucial for getting JSON consistently
}try: response = requests.post(OPENAI_API_URL, headers=headers, json=payload) response.raise_for_status() return response.json()['choices'][0]['message']['content'] except requests.exceptions.RequestException as e: print(f"API request failed: {e}") return None except KeyError: print(f"Unexpected API response structure: {response.text}") return None... (rest of the script to load slack messages and loop)
Storage and Reporting: I dump the extracted JSON objects into a simple CSV file, which then gets loaded into a Google Sheet. It's not fancy, but it gives us a clear, filterable list of [Sentiment, Feature Request, Original Message Link] that we can sort and review in our meeting. This saves us at least 3-4 hours of manual work every single week.
It's not perfect, LLMs can still hallucinate or misinterpret occasionally, but the consistency is miles better than my manual efforts. Now we come to the meeting with actual counts of positive/negative feedback, a prioritized list of user-requested features, and a much clearer picture of what our users are actually saying. It's been a total game-changer for focusing our product roadmap and just understanding our users better. This setup, from idea to first reliable run, probably took me about two and a half days spread across a week, but the payoff is massive.
Top comments (1)
Nice approach to the triage side. One thing that helped us upstream of this: a lot of feedback that lands in Slack is too thin to classify well ("export is confusing"), and no script can recover context that was never collected. We now ask one clarifying question at capture time, which makes the later clustering much cleaner. That's the product I work on (chatform.in, forms that follow up on vague answers), so take it with that bias. How are you handling the items your script can't confidently place?