A harm-detection research tool
RAIUZ reads Uzbek-language YouTube comments and explains how harm shows up in them.
Built for researchers studying online harm in Uzbek-language spaces — it teaches, it does not moderate.
"Qishloqidan kelib shaharda yashaydiganlar…"A derogatory remark aimed at people who moved from rural areas to the city — a specific fault line in Uzbek discourse.
Why this matters
Harm in Uzbek-language spaces does not map cleanly onto Western harm-detection models. The fault lines are culturally specific: the role of the daughter-in-law, tensions between rural and urban life, region and migrant status, and public policing of religiosity.
A generic model trained on English toxicity will miss most of this — either flagging nothing, or flagging the wrong things. RAIUZ starts from a taxonomy built for this context, so researchers can see harm as it is actually expressed.
How it works
Scrape
Comments are collected from public Uzbek-language YouTube videos.
Keyword match
Each comment is checked against a culturally-specific keyword taxonomy for a fast first pass.
LLM fallback (Claude)
Uncertain cases are escalated to a language model, which returns a category and its reasoning.
Human review
Researchers review flagged comments in the dashboard, correcting the taxonomy over time.
The taxonomy
Sixteen subcategories across two top-level categories. Browse them like a field guide — each entry has a plain-language definition.
Hate & Abuse
4 groups · 16 subcategories
- General misogyny
- General derogatory or demeaning language directed at women.
- Marital status policing
- Judging a woman's worth or choices based on her marital status.
- Modesty / appearance policing
- Policing a woman's appearance, dress, or modesty.
- Bride / daughter-in-law role
- Reducing a woman's value to her role as a bride or daughter-in-law.
- Failure of masculinity
- Mocking a man for not meeting traditional masculine expectations.
- Rural / internal migrants
- Derogatory remarks targeting people who moved from rural areas to cities.
- Region / ethnic stereotyping
- Stereotyping based on region, ethnicity, or nationality.
- Regional language differences
- Mocking regional dialects or accents.
- Labor migrants
- Derogatory remarks targeting labor migrants.
- Class
- Derogatory remarks based on perceived social or economic class.
- Profession
- Derogatory remarks targeting a person's occupation.
- Disability
- Derogatory remarks targeting disability.
- Ageism
- Derogatory remarks based on age.
- Education
- Derogatory remarks based on educational background.
- Too religious
- Mocking someone for being visibly or strictly religious.
- Not religious enough
- Mocking someone for not being religious enough.
PII & Data Breach
4 identifier types
- PINFL
- Uzbekistan's 14-digit personal identification number shared in a comment.
- Passport number
- A passport series-and-number pattern shared in a comment.
- Phone number
- A personal phone number shared in a comment.
- Card number
- A bank card number (e.g. Uzcard, Humo) shared in a comment.
Team
- Aziza Mirsaidova
- Research LeadStanford Ethics Fellow, Oracle AI Scientist.
- Fahed Daibes
- EngineeringInfrastructure, frontend, deployment.
Try the taxonomy on a comment
Paste a YouTube link or a line of text and see how RAIUZ reads it.
Open the Classify tool