Monday 03 March 2025
In recent years, large language models (LLMs) have become increasingly prevalent in various industries, including hiring and recruitment. These models are designed to analyze resumes and job postings, and provide recommendations for potential matches. However, researchers have raised concerns about the fairness of these models, particularly when it comes to demographic groups such as women and minorities.
To investigate this issue, a team of researchers conducted a study on the performance of LLMs in generating resume summaries for different demographic groups. The study used a dataset of over 10,000 resumes from Kaggle, a popular platform for data science competitions, and evaluated the models’ ability to generate summaries that were fair and unbiased.
The researchers found that the LLMs exhibited significant differences in their performance across demographic groups. For example, the models tended to favor male candidates with more extensive work experience over female candidates with similar qualifications. Similarly, the models were less likely to select resumes from minority candidates, even when those candidates had stronger qualifications.
The study also explored the impact of name perturbations on the models’ performance. The researchers found that when they replaced male names with female names or white names with black names, the models became more biased in their selection of candidates. This suggests that the models are not just making mistakes due to a lack of data, but are actually learning biases from the patterns they see in the training data.
To address these issues, the researchers proposed several metrics for evaluating the fairness of LLMs in generating resume summaries. One metric, called non-uniformity, measures how evenly the models distribute candidates across demographic groups. Another metric, called exclusion, measures whether the models are more likely to exclude certain demographic groups from their recommendations.
The study’s findings have important implications for the use of LLMs in hiring and recruitment. If left unchecked, these biases could lead to discriminatory outcomes and perpetuate existing inequalities in the workforce. However, by using metrics like non-uniformity and exclusion, researchers can identify areas where the models are biased and work to mitigate those biases.
The study’s authors hope that their research will inform the development of more fair and unbiased LLMs. They suggest that future studies should explore ways to improve the diversity of training data, as well as develop new metrics for evaluating fairness in machine learning models.
Ultimately, the goal is to create LLMs that are not only accurate but also fair and unbiased.
Cite this article: “Fairness Concerns in Large Language Models: A Study on Resume Summarization Biases”, The Science Archive, 2025.
Language Models, Fairness, Bias, Recruitment, Hiring, Resumes, Job Postings, Demographic Groups, Women, Minorities, Machine Learning







