Observation period: August 2026- September 2026
Duration: 2 weeks
Due to the enormous amount of user-generated content shared on Instagram, effective content moderation is a crucial element for online safety and security. This study adopts a user-level observational approach to examine Instagram's content moderation and reporting mechanisms.
Objective: The aim of this study is to observe how Instagram responds to potentially inappropriate types of content and whether moderation results vary depending on the content's format. The study also aims to examine the effectiveness of Instagram's user reporting mechanism and the feedback provided to users following reports.
Research Questions
1)Does Instagram's observable moderation response differ between images, videos, and direct messages?
2)How often do user reports result in visible enforcement actions?
3)What types of enforcement actions are observed following reports?
4)How consistently does Instagram provide
feedback following a report?
Methodology: Over a period of 2 weeks, I reported over 50 pieces of content shared via Instagram Stories, over 20 Reels videos that I considered potentially inappropriate, and over 20 direct messages that I considered potentially inappropriate. I then observed Instagram's enforcement actions and the feedback it provided.
Fraud/fake accounts: Accounts created for fraudulent purposes and structured to resemble real user profiles were appeared more difficult to identify in observational tests based solely on profile and content. This could increase the importance of user reporting, especially if the account has just been created.
Hate speech: Users are given the option to filter specific words or phrases. However, in the tests conducted, some comments expressing the same meaning with different spellings were able to pass through the filtering mechanism.
Sexual/Inappropriate Content: Rapid intervention was observed in some content that could be clearly deemed inappropriate. Conversely, changes such as reducing the size of the content or censoring certain parts were found to make detection difficult in some cases.
User Reporting System: The effectiveness of user reporting appeared more limited in cases where it could not be automatically determined whether the content constituted a violation.
Feedback Consistency: While the user initially received feedback about the action taken after submitting a report, in some tests no additional notification was received regarding subsequent stages.
Findings: A difference was observed between content formats. Image-based content was found to be subject to moderation more frequently than video content. Only a small percentage of reported Reels videos were removed but some videos were subject to moderation after review. Direct messages, however, were different. In many cases, the notification did not result in an account ban. Instead, temporary messaging restrictions lasting a few days were observed, and in some cases, no visible action was taken.
**Limitations: 1)These observations are based on controlled user-level tests and do not provide information about Instagram's internal moderation models, datasets, or application infrastructure. Therefore, the findings should not be interpreted as a comprehensive evaluation of Instagram's content moderation system.
2)The original content was not available for retrospective analysis. Therefore, the study relies on the observations and moderation outcomes recorded during the observation period.
3)To protect user privacy and avoid unnecessarily reproducing potentially harmful material, original reported content and identifiable information were excluded from the published analysis.**
I hope this study contributes to a better understanding of the user-facing aspects of content moderation and provides useful observations for both users and platform stakeholders.
Top comments (0)