A harm-detection research tool

RAIUZ reads Uzbek-language YouTube comments and explains how harm shows up in them.

Built for researchers studying online harm in Uzbek-language spaces — it teaches, it does not moderate.

A comment, in miniature
"Qishloqidan kelib shaharda yashaydiganlar…"
Regional · Rural / internal migrants

A derogatory remark aimed at people who moved from rural areas to the city — a specific fault line in Uzbek discourse.

Why this matters

Harm in Uzbek-language spaces does not map cleanly onto Western harm-detection models. The fault lines are culturally specific: the role of the daughter-in-law, tensions between rural and urban life, region and migrant status, and public policing of religiosity.

A generic model trained on English toxicity will miss most of this — either flagging nothing, or flagging the wrong things. RAIUZ starts from a taxonomy built for this context, so researchers can see harm as it is actually expressed.

How it works

  1. Scrape

    Comments are collected from public Uzbek-language YouTube videos.

  2. Keyword match

    Each comment is checked against a culturally-specific keyword taxonomy for a fast first pass.

  3. LLM fallback (Claude)

    Uncertain cases are escalated to a language model, which returns a category and its reasoning.

  4. Human review

    Researchers review flagged comments in the dashboard, correcting the taxonomy over time.

The taxonomy

Sixteen subcategories across two top-level categories. Browse them like a field guide — each entry has a plain-language definition.

Hate & Abuse

4 groups · 16 subcategories

General misogyny
General derogatory or demeaning language directed at women.
Marital status policing
Judging a woman's worth or choices based on her marital status.
Modesty / appearance policing
Policing a woman's appearance, dress, or modesty.
Bride / daughter-in-law role
Reducing a woman's value to her role as a bride or daughter-in-law.
Failure of masculinity
Mocking a man for not meeting traditional masculine expectations.
Rural / internal migrants
Derogatory remarks targeting people who moved from rural areas to cities.
Region / ethnic stereotyping
Stereotyping based on region, ethnicity, or nationality.
Regional language differences
Mocking regional dialects or accents.
Labor migrants
Derogatory remarks targeting labor migrants.
Class
Derogatory remarks based on perceived social or economic class.
Profession
Derogatory remarks targeting a person's occupation.
Disability
Derogatory remarks targeting disability.
Ageism
Derogatory remarks based on age.
Education
Derogatory remarks based on educational background.
Too religious
Mocking someone for being visibly or strictly religious.
Not religious enough
Mocking someone for not being religious enough.

PII & Data Breach

4 identifier types

PINFL
Uzbekistan's 14-digit personal identification number shared in a comment.
Passport number
A passport series-and-number pattern shared in a comment.
Phone number
A personal phone number shared in a comment.
Card number
A bank card number (e.g. Uzcard, Humo) shared in a comment.

Team

Aziza Mirsaidova
Research LeadStanford Ethics Fellow, Oracle AI Scientist.
Fahed Daibes
EngineeringInfrastructure, frontend, deployment.

Try the taxonomy on a comment

Paste a YouTube link or a line of text and see how RAIUZ reads it.

Open the Classify tool