An audit of Hindi and Hinglish content moderation classifiers on three major platforms
- Dr. Rohit Malhotra · Microsoft Research IndiaORCID 0000-0001-5567-8813
- Dr. Priya Venkataraman · IIIT HyderabadORCID 0000-0002-3319-7745
- Aisha Khan · IIIT Hyderabad
- Journal
- Journal of AI Ethics & Society
- Volume
- 8, issue 1
- Pages
- 22–50
- Licence
- CC BY 4.0
Abstract
Content moderation classifiers are evaluated overwhelmingly on English. We construct a 24,000-item benchmark of Hindi, Romanised Hindi and Hinglish social media text, annotated by nine trained annotators against a published codebook with substantial inter-annotator agreement, and audit the public moderation endpoints of three major platforms against it. All three perform far worse on Romanised Hindi than on Devanagari Hindi, and worst of all on code-mixed Hinglish, which is the register most actually used. False negative rates on caste-based slurs reach 71 per cent. We release the benchmark and the annotation codebook.
Keywords
Introduction
Automated content moderation now operates at a scale that makes human review of its decisions impossible in principle, not merely in budget. The question of how well these systems work has therefore become a question about their aggregate statistical behaviour, and that behaviour is documented almost exclusively for English text.
India is the largest user base for each of the three platforms we audit. The dominant registers of Indian social media text are Romanised Hindi and Hinglish — Hindi written in Latin script, freely code-mixed with English. Neither is well represented in the public evaluation sets these systems report against.
Benchmark construction
We sampled 24,000 public posts across the three platforms between March and September 2025, stratified to give equal representation to Devanagari Hindi, Romanised Hindi and code-mixed Hinglish. Nine annotators, all first-language Hindi speakers with prior moderation experience, labelled each item against a published five-category codebook.
Krippendorff's alpha across the full set was 0.79, and 0.74 on the caste-slur subcategory, which is the lowest-agreement category and the one where the codebook required the most revision. All disagreements were adjudicated by a tenth senior annotator whose decisions are recorded separately in the release so that others can reconstruct or contest them.
Audit results
Against Devanagari Hindi, the three endpoints achieve F1 of 0.81, 0.77 and 0.74 on the binary harmful/not-harmful task. Against Romanised Hindi the same endpoints achieve 0.62, 0.58 and 0.55. Against Hinglish they achieve 0.51, 0.49 and 0.44 — close to, and in one configuration below, the majority-class baseline.
The caste-slur subcategory is worse than the aggregate figures suggest. False negative rates are 64, 68 and 71 per cent respectively, and inspection of the misses shows a consistent pattern: slurs that are orthographically close to innocuous words survive Romanisation in a form the classifiers do not recognise, while their meaning to a reader is unambiguous.
Discussion
The gap between the Devanagari and Romanised results is the practically important one, because Romanised text is the majority register and Devanagari the minority. A system evaluated on the script its developers can most easily source annotation for will report numbers that do not describe its deployment.
We release the benchmark, codebook and adjudication record under CC BY 4.0. We deliberately do not release the raw post text with user identifiers, and we discuss the tension between reproducibility and the re-identification risk that a public corpus of moderated speech creates for the people who wrote it.
References
- 1.Sap, M., et al. (2019). The risk of racial bias in hate speech detection. Proceedings of ACL, 1668–1678.
- 2.Bender, E. M., & Friedman, B. (2018). Data statements for NLP. TACL, 6, 587–604.
- 3.Raji, I. D., & Buolamwini, J. (2019). Actionable auditing. Proceedings of AIES, 429–435.
- 4.Krippendorff, K. (2018). Content Analysis: An Introduction to Its Methodology (4th ed.). SAGE.
Cite this article
Malhotra, R., Venkataraman, P., Khan, A. (2026). An audit of Hindi and Hinglish content moderation classifiers on three major platforms. Journal of AI Ethics & Society, 8(1), 22–50. https://doi.org/10.53456/jaies.2026.8.1.22
Related in Computer Science
The people behind the label: working conditions in India's data annotation industry
Prof. Daniel Okonkwo, Aisha Khan, Dr. Hannah Weiss
Journal of AI Ethics & SocietyVolume 7, issue 42025pp. 401–42852 citations
Machine learning depends on annotation labour that its literature rarely describes. Drawing on 68 interviews with annotators and eleven with managers across nine firms in Bengaluru, Kochi and Bhubaneswar, we document a labour process organised around quota, surveillance and a quality regime that transfers the cost of ambiguous data to the worker. We argue that annotation guidelines function as an unacknowledged site of normative decision-making: annotators routinely resolve genuine moral ambiguity under time pressure, and their resolutions are laundered into training data as ground truth.