Biases in large language models (LLMs) such as ChatGPT, particularly in health and HIV contexts, remain underexamined. This study adopted a psychological lens to examine sociodemographic biases in ChatGPT by comparing its predictions of HIV knowledge with actual responses from adolescents and young adults using nationally representative survey data from the Philippines. Prompt-based simulations using GPT-4o were conducted using participants’ sociodemographic profiles (n = 1,393) and analyzed using binary and multinomial logistic regression models. ChatGPT was more likely to inaccurately predict HIV knowledge for certain groups, including LGBTQ individuals (adjusted odds ratio [aOR] = 1.65, p = .025), older participants (aOR = 1.08, p = .005), urban residents (aOR = 1.78, p < .001), and those with higher formal education (aOR = 3.58, p < .001). These inaccuracies reflected overestimations. Findings suggest the presence of implicit biases, and underscore the need to evaluate LLMs for equitable application in health education.
Keywords: artificial intelligence; ChatGPT; gender and sexuality; health education; HIV and AIDS; natural language processing; youth
Bilon, X. J. (2026). Can large language models support health literacy? Examining sociodemographic biases in ChatGPT’s representation of HIV knowledge. Journal of HIV & Social Services, 24(1), 8–28. https://doi.org/10.1080/15381501.2025.2547229