Skip to content

Computer Science

Algorithms, systems, artificial intelligence, security and the theory underneath them.

1 journal2 articles1 book4 records in total

Sub-fields

Journals

Articles

All 2
Research ArticleOpen access

An audit of Hindi and Hinglish content moderation classifiers on three major platforms

Dr. Rohit Malhotra, Dr. Priya Venkataraman, Aisha Khan

Journal of AI Ethics & SocietyVolume 8, issue 12026pp. 22–5064 citations

Content moderation classifiers are evaluated overwhelmingly on English. We construct a 24,000-item benchmark of Hindi, Romanised Hindi and Hinglish social media text, annotated by nine trained annotators against a published codebook with substantial inter-annotator agreement, and audit the public moderation endpoints of three major platforms against it. All three perform far worse on Romanised Hindi than on Devanagari Hindi, and worst of all on code-mixed Hinglish, which is the register most actually used. False negative rates on caste-based slurs reach 71 per cent. We release the benchmark and the annotation codebook.

Research ArticleOpen access

The people behind the label: working conditions in India's data annotation industry

Prof. Daniel Okonkwo, Aisha Khan, Dr. Hannah Weiss

Journal of AI Ethics & SocietyVolume 7, issue 42025pp. 401–42852 citations

Machine learning depends on annotation labour that its literature rarely describes. Drawing on 68 interviews with annotators and eleven with managers across nine firms in Bengaluru, Kochi and Bhubaneswar, we document a labour process organised around quota, surveillance and a quality regime that transfers the cost of ambiguous data to the worker. We argue that annotation guidelines function as an unacknowledged site of normative decision-making: annotators routinely resolve genuine moral ambiguity under time pressure, and their resolutions are laundered into training data as ground truth.

Books