Hate Speech Moderation

The article by Wilson and Land, "Hate Speech on Social Media: Content Moderation in Context (2021)," emphasizes the importance of considering context when developing evidence-based hate‑speech policies. Here, context can be the implicit social and culatural norms that shape users' behaviors of behaviors. Even though some users perform profenity to attach other users, platforms might need to see what caused them to behave that way, such as who created dangerous speech. On the other hand, even if singer users' contribution is small and minimal, they can be magnified by the platform's algorithms.

Nowadays, Generative AI is considered as a way to moderate scalable hate speech, yet the effectivess of this approach is still being developed. Fasching and Lelkes found that varied consistency was observed across differen AI models in detecting hate speech and these variations are pronounced for specific demographic groups. The researchers highlighted that different AI models weigh hate speech, intent, or context differently, which can lead to inconsistent moderation outcomes. Although Generative AI can be a powerful tool to reduce workload for human moderators, this inconsistency and lack of explainability of their moderation can lead to make platform space less inclusive and unreliable for users. In 2026 UN special report, still 45 percent of women journalists and media workers self-censor on social media to avoid abuse.

Returning to the authors' claim, I agree that laws can be improved by a better understanding of social and cultural norms. Platform users do not operate within static normative ethics; they continually revise standards based on the communities they trust and beliefs shaped by lived experience. As the authors note, these expectations can be demanding for platforms. Amid the current flood of AI‑generated "slop" on social media, it is unrealistic to expect platforms to shift slowly from automated systems to fully human‑driven moderation. With advanced LLM approaches, automating contextualized moderation may be feasible. At the same time, this could introduce new forms of discrimination — for example, neglecting low‑resource languages and indigenous cultures in model training — and it raises questions about how agents interpret context. Context should be analyzed from multiple perspectives rather than judged at once, so humans can oversee and remain accountable for moderation decisions.