Customer feedback is one of the critical sources of information that can be leveraged to improve a product; however, it is often unstructured and noisy. The same issue can be described in different ways, and it may be challenging to single out the most pressing problems that require immediate action. One of the use cases I worked on recently aimed at analysing end-user feedback and extracting meaningful patterns from it. The data set we worked on comprised 65 entries related to reports, dashboard, tasks, and billing issues.
The initial analysis was fairly straightforward, as the data set was easy to read and understand. Nonetheless, eyeballing the spreadsheet did not help highlight any recurring issues on which the product team could agree. It was also challenging to estimate the significance of different problems reported. Grouping entries by a particular product feature helped to a certain extent, as it highlighted the most frequently mentioned areas. Nonetheless, the importance of individual entries varied considerably, and some topics appeared together frequently, hinting at potential clusters.
Working with these data sets further highlighted the importance of choosing appropriate tools for the task at hand. Using python’s pandas library, I explored the data set’s properties and got a rough idea of the data structure. Data frames were also used to group the entries by feedback type and evaluate the distribution of different categories. Finally, a sentiment score was calculated for each entry, which provided some insights into the general attitude toward specific problems reported.
When individual entries were grouped into categories, it became evident that the same issues were often raised in different terms. One of the recurring topics was related to the need for a solution for managing recurring tasks. While many entries contained the word ‘recurring’, some did not mention this term at all, which probably contributed to the fact that this problem was not prioritized. Other terms used in the entries could also be found in data sets related to other products, suggesting that the presence of similar terms does not necessarily imply that the same issue is present.
The process of training the model to identify these patterns was also surprisingly enlightening. First, I realized that the data needed to be prepared and validated before attempting to use it with any machine learning algorithms. Second, I found that it is important to keep track of the model’s training data to ensure that it actually captures the intended information. Both insights highlight the importance of thorough data preparation and validation in machine learning projects.
Keeping these findings in mind, I would approach the task differently if given another opportunity to work on it. In particular, it would be interesting to track temporal trends and see if any issues had been consistently reported over an extended period. The model findings could also be visualized better, perhaps with the use of a dashboard that would allow users to sort entries based on particular topics or dates. Overall, the project was a great opportunity to demonstrate how a few simple python operations can be used to extract meaningful insights from unstructured data. However, it also highlighted the critical importance of proper data formatting. The findings would be more accurate if the model could group similar entries automatically, which is why semantic clustering would be a better choice than the frequency distribution of individual terms.
Top comments (1)