A new study has raised concerns about how widely used artificial intelligence tools respond to people searching for help with mental health problems. Researchers found that popular large language models show clear signs of gender bias when assessing borderline personality disorder, despite appearing accurate and empathetic on the surface. The findings, published in Personality and Mental Health, matter because many people turn to mental health apps and AI chatbots for information on anxiety, personality disorders, and self-diagnosis, often hoping to avoid the stigma they fear in face-to-face services.
The research compared responses from ChatGPT-3.5, ChatGPT-4, and Google Gemini with the reactions of 218 mental health care practitioners. Each system or human participant reviewed the same clinical vignette describing either a woman or a man with symptoms consistent with borderline personality disorder. The team examined how accurately the systems identified the disorder and whether they displayed signs of public stigma, such as mistrust, pessimism, or assumptions about attention-seeking behaviours.
The study found that the AI models were more accurate than clinicians when identifying the disorder, with all systems recognising borderline personality disorder in both the woman and the man described in the vignette. Human clinicians were less consistent. This will interest people searching for “borderline personality disorder symptoms”, “mental health apps”, or “AI diagnosis”, since accurate identification is often what users expect these tools to provide.
Despite this high accuracy, the researchers found that all three AI systems held gender-skewed assumptions once they moved beyond diagnosis. The models were more likely to describe a man as having additional co-occurring conditions, even though no extra disorders were presented in the vignette. The woman was more often described as having chronic and unchangeable symptoms, suggesting a more pessimistic view of her ability to recover. Google Gemini showed the strongest bias, rating the woman as more dramatic, less trustworthy, and more likely to seek attention than the man. These patterns point to the kind of public stigma commonly associated with borderline personality disorder.
Researchers also found that the AI systems struggled with understanding the nuance of symptoms. They sometimes misidentified behaviours such as brief sexual encounters or binge drinking as separate diagnostic features, showing how mental health chatbots can misinterpret context when users ask questions about impulsivity, stress, or relationships. This is particularly relevant for young people using AI tools to explore questions about mood swings, anxiety, or self-harm, believing these systems to be neutral and objective.
While the models expressed more empathy than human clinicians overall, they also rated people with the disorder as less trustworthy and less capable of forming stable relationships. This suggests that users searching for support may receive subtly discouraging or inaccurate guidance from systems intended to help them. The study highlights the continued problem of misinformation in mental health apps and reinforces the need for transparency in how these tools are trained.
Experts in the study call for developers to work more closely with clinicians and people with lived experience to reduce bias in future versions of these systems. With more people turning to AI tools for advice on stress, anxiety, and personality disorders, the researchers argue that the technology must avoid reinforcing old stereotypes.