Analyzing workplace discrimination
KAIST has something called an “AI Specialized Major”, which you can fulfill by taking AI/ML/robotics classes from different departments. I was once on that track, and for that I took a class called Statistical Methods with Computer (MAS456). It was based on the An Introduction to Statistical Learning book (ISLR/ISLP for R and Python versions, respectively). It covered the statistical basics for ML, and the final project was to analyze workplace discrimination. We were basically asked to do some EDA, and it was a fun little project where I had to think extra to give the statistical metrics some real-world interpretation.
The dataset is called “7th Wave of Korean Labor and Income Panel Study” from 2004, and it covers background info and discrimination experienced by individuals. The data columns are:
| No | Variable | Name/Description | Possible Answers |
|---|---|---|---|
| 1 | disc_hire | Response to the question: “Have you ever experienced discrimination in getting hired?” | 0: ‘No’, 1: ‘Yes’, NA: ‘Not Applicable’ |
| 2 | Gender | Gender | 0: male, 1: female |
| 3 | Age | Age | 0: 16–24, 1: 25–34, 2: 35–44, 3: 45–54, 4: 55–64, 5: 65+ |
| 4 | Edu_cat | Education level | 0: middle school graduate or less, 1: high school graduate, 2: college graduate or more |
| 5 | Marriage | Marital status | 0: never married, 1: currently married, 2: previously married |
| 6 | Emp_fin | Employment status | 0: permanent, 1: non-permanent |
| 7 | Income_quartile | Total household income ÷ √(household size) | 0: Q1, 1: Q2, 2: Q3, 3: Q4 |
| 8 | Birth_region | Birth region | 1: Jeolla-do, 0: other regions |
| 9 | Self-rated health | Response to: “How would you rate your health?” | 0: ‘very good’, 1: ‘good’, 2: ‘poor’, 3: ‘very poor’ |
| 10 | Disability | Response to: “Do you have any impairment or disability?” | 0: ‘No’, 1: ‘Yes’ |
| 11 | Residence | Residential areas | 1: Seoul, 2: Pusan, 3: Daegu, 4: Daejeon, 5: Incheon, 6: Gwangju, 7: Ulsan, 8: Kyunggi, 9: Kangwon, 10: Choongbuk, 11: Choongnam, 12: Jeonbuk, 13: Jeonnam, 14: Kyungbuk, 15: Kyungnam |
| 12 | disc_wage | Experience of discrimination in receiving income | 0: ‘No’, 1: ‘Yes’, 2: ‘Not Applicable’ |
| 13 | disc_jobedu | Experience of discrimination in training | 0: ‘No’, 1: ‘Yes’, 2: ‘Not Applicable’ |
| 14 | disc_promotion | Experience of discrimination in getting promoted | 0: ‘No’, 1: ‘Yes’, 2: ‘Not Applicable’ |
| 15 | disc_resign | Experience of discrimination in being fired | 0: ‘No’, 1: ‘Yes’, 2: ‘Not Applicable’ |
| 16 | disc_edu | Experience of discrimination in obtaining higher education | 0: ‘No’, 1: ‘Yes’, 2: ‘Not Applicable’ |
| 17 | disc_home | Experience of discrimination at home | 0: ‘No’, 1: ‘Yes’, 2: ‘Not Applicable’ |
| 18 | disc_social | Experience of discrimination at general social activities | 0: ‘No’, 1: ‘Yes’, 2: ‘Not Applicable’ |
Logistic Regression: What Drives Discrimination?
The first step was to train a logistic regression model to predict hiring discrimination by using the remaining 17 factors. The columns have relatively same range, so we can kinda interpret the coefficients as feature importance (but there are some “but”s like multicollinearity). The feature coeffitients are
[(‘disc_wage’, 3.148), (‘disc_social’, 1.041), (‘emp_fin’, 0.3865), (‘disc_jobedu’, 0.2166), (‘health’, 0.2076), (‘disc_promotion’, 0.1670), (‘disc_resign’, 0.099223), (‘age’, 0.05661), (‘residence’, 0.03373), (‘disability’, 0.03290), (‘disc_edu’, -0.009576), (‘gender’, -0.01735), (‘edu_cat’, -0.1494), (‘disc_home’, -0.1836), (‘income_quartile’, -0.2065), (‘marriage’, -0.2084), (‘birth_region’, -0.2385)]
And the result? Exactly what you’d expect 😅:
- If someone is discriminated in other areas (training, promotions, income), they are usually discriminated at their workplace as well.
- Socioeconomic factors such as employment status, marital status, health, education, and income predict the discrimination very well.
- Part-time workers, in particular, had significantly higher risk.
This makes sense: discrimination rarely happens in isolation. It tends to cluster and compound across different dimensions of life.
PCA and Clustering: Who Gets Grouped Together?
Regression tells us which features matter individually, but I also wanted to know: how do people cluster when you look at the whole feature space?
To do this, I ran a Principal Component Analysis (PCA) and reduced the data to 6 dimensions (explaining ~70% of the variance). Then I applied k-means clustering with k=4, chosen as the smallest number that produced visually distinct clusters.
Here’s what stood out:
- Cluster 1: Mostly middle-to-low income workers without college degrees. They had the highest rates of discrimination during hiring and often left questions blank when asked about promotion/training discrimination.
- Cluster 2: Young, permanent workers with high income and at least some college education. Very low discrimination reported.
- Cluster 3: Middle-aged, married, middle-to-low income workers with health issues. Surprisingly low discrimination reported, but high rates of self-reported poor health.
- Cluster 4: Workers with disabilities. This group stood out sharply as their experiences with the workplace were dramatically different - high rates of discrimination, low income, and overrepresentation of elderly workers.
One of the most sobering findings: disability was essentially a cluster-defining characteristic. These individuals were socioeconomically disadvantaged across the board.
Underreporting: Do Women Stay Silent?
One striking finding came from analyzing individuals who refused to answer whether they had experienced hiring discrimination. Using both logistic regression and random forest classifiers, I tried to predict their responses.
The results showed:
- Women who didn’t answer were overwhelmingly predicted to have faced discrimination.
- Men also underreported, but at lower rates.
This aligns with what’s been discussed in sociology: women often face social or workplace stigma that discourages them from speaking openly about discrimination.
Health and Discrimination: A Statistical Link
Finally, I tested whether there’s an association between health status and discrimination experiences. Using ANOVA and pairwise Tukey’s HSD tests across four groups (experienced discrimination, no discrimination, declined to answer but predicted yes/no), the results were clear:
- Workers who reported discrimination had significantly worse health outcomes than those who didn’t.
- The groups with missing answers followed predicted distributions, suggesting the signal is real.
This is important because it shows that discrimination doesn’t just affect income or career, it correlates with physical well-being too.
Why This Matters
This project showed me how statistical and computational tools can be used to uncover systematic inequalities hidden inside large datasets.
Some takeaways:
- Discrimination is multi-faceted: the same people often face multiple types at once.
- Disability, education, age, and income all intersect to shape unequal outcomes.
- Silence (underreporting) is itself a signal - especially for marginalized groups.
- The consequences are not just professional, but also health-related.
While this dataset was from 2004 Korea, the methods can be applied anywhere - and sadly, I suspect many of the same patterns would still show up today.
Enjoy Reading This Article?
Here are some more articles you might like to read next: