LLMs are the Third Wave of Automated Content Moderation
On September 10, Penn MEDIATED hosted a second convening as part of our ongoing project examining how Large Language Models (LLMs) are shaping civic discourse, that is, how people use LLMs to access and interpret information about social and political topics.
On September 10, Penn MEDIATED hosted a second convening as part of our ongoing project examining how Large Language Models (LLMs) are shaping civic discourse, that is, how people use LLMs to access and interpret information about social and political topics. Our first convening focused on the politics of LLM chatbots and helped inform a Carnegie Endowment white paper co‑authored by Penn Professor Danaé Metaxa and Executive Director Alex Engler. For our second event, we turned to the specific use of LLMs in online platform content moderation.
Automation in content moderation is already pervasive: in the last three months of 2025, 98 percent of YouTube video removals and 99.8 percent of comment removals are now flagged by automated systems. For most major platforms, automation handles a sizable portion of content moderation decisions. After hashing algorithms and supervised machine learning, LLMs mark a third wave in automated moderation. To unpack what this means in practice, we convened academic researchers, technologists building moderation tools, and civil society organizations focused on how these tools shape online speech rights.
Miranda Bogen, Director of the AI Governance Lab at the Center for Democracy and Technology, opened by framing how LLMs are influencing the broader information ecosystem. She pointed out that unlike previous technological changes to information seeking—like personalized social media feeds and search engines—LLMs often appear to be singularly authoritative sources. The potential dangers, she argued, are already visible in critical domains like inaccurate health information and misleading election advice. Miranda also underscored that these design choices should be surfaced and debated in public rather than being “subsumed in technical interventions under the guise of safety.”
We then turn to a series of presentations from practitioners who provide insights into how automated content moderation actually works, and where LLMs fit in. Alice Hunsberger, Head of Trust and Safety at Musubi, shared her perspective on how automation has shaped human moderation, based on her experience working at platforms like Grindr and OkCupid. She noted that automated systems help reduce the burden on human moderators, reducing exposure to traumatic content. But these systems can be rigid and not particularly adaptable to novel content. Alice shared an example from her time at Grindr, where the sexual-content classifier assumed a binary male versus female taxonomy that did not fit the platform's largely LGBTQ user base.
LLMs, in Alice’s view, are more flexible and adaptable than prior automation. Rather than inheriting a fixed label set, trust and safety teams can now feed their own policy text into a model and ask for decisions that are explicitly grounded in that policy. Alice admits that LLMs are far from perfect—they are subject to bias, gaps in language coverage–but the key advantage in her assessment is that they can give more agency to trust and safety teams and help prioritize cases needing human review.
To show what policy-steerable moderation looks like, Dave Willner, co-founder of Zentropi, demonstrated their “Bring Your Own Policy” CoPE model. In his demo, Dave showed how a BYOP hate speech model can interpret a natural‑language policy and label content in line with those criteria. Zentropi also built a tool that provides AI-assisted recommendations on refining policy language to ensure the model better understands the goals of the policy.
Juliet Shen, Head of Product at Roost, turned to the wider ecosystem of open-source safety infrastructure. Roost is a non‑profit building free, open‑source online safety tools deployed across platforms like Matrix, Bluesky and Notion. Roost develops technical infrastructure that help platforms surface patterns of abuse, support human reviewers, and structure enforcement actions. Juliet emphasized that developing these tools in public under permissive open‑source licenses allows researchers, smaller platforms, and civil society groups globally to inspect, adapt, and improve moderation systems for their own languages and contexts.
Two academic presentations then examined how these systems behave once deployed, and how they impact users. Penn researchers Neil Fasching and Yphtach Lelkes presented research on how different moderation models disagree when labeling the same hate‑related text. Analyzing roughly 1.3 million sentences, they found that models can arrive at different moderation decisions for the same piece of content. Neil made the case that while the BYOP approach may make it easier to customize policies, it is still unclear whether the underlying model will interpret policies consistently. Next, Penn’s Danaé Metaxa and Haverford’s Sorelle Friedler’s research presentation examined the threat of identity speech suppression by LLM systems. They examined how identity-related speech is more likely to be incorrectly flagged as toxic or hateful, and how this burden falls disproportionately on marginalized groups.
Dia Kayyali, a digital rights advocate and former Meta Oversight Board member, opened the group discussion. Dia discussed civil society concerns around the gap between platforms’ public content policies and the far more detailed, often opaque instructions given to LLMs, making it hard for affected communities to know or contest what norms are being enforced. They argued that over‑removal carries serious human‑rights costs, including censorship, loss of human rights documents, and stifling dissent.
In closing, we looked at how public policy developments may compel de facto automation requirements for platforms. Laws like Vietnam’s latest cybersecurity law and amendments to India’s IT rules impose strict takedown deadlines and expansive liability, making it functionally impossible for large platforms to rely on human workflows alone. Shrinking content takedown windows coupled with penalties for non-compliance may inadvertently lead platforms to over moderate to avoid legal risk, and the burden of those errors often falls unevenly on marginalized communities and sensitive political speech.
Collectively, our speakers delivered a clear message that LLMs will reshape online content moderation. The convening also highlighted the importance of interdisciplinary collaboration. It is imperative to ensure that the developers building content moderation tools are in conversation with the academic researchers and civil society advocates who are documenting its impact. This convening and our ongoing LLMs and Civic Discourse project has and will continue to bring these stakeholders together to advance a pro-democratic vision for these technologies.

