Abstract
Large language models (LLMs) are increasingly used by adolescents to seek information about sexuality, relationships, and well-being. However, little is known about how these systems discursively construct and mediate such interactions. This study examines how three commercial LLMs (ChatGPT, Claude, and Gemini) respond to simulated Italian adolescent personas used as analytical probes for algorithmic bias. Drawing on 120 standardized Italian-language interactions across four personas differing by gender, class, ethnicity, and sexual orientation, the analysis investigates how identity markers shape the framing of sexual and affective health information. Qualitative analysis reveals recurrent patterns of medicalization, differential agency attribution, and cultural othering. As markers of marginalization accumulate, AI responses exhibit a systematic discursive shift from normalization toward surveillance framing and referral to professional authority. To account for this pattern, we introduce intersectional amplification as an analytical construct grounded in Italian-language interactions about adolescent sexuality with commercial LLMs, capturing a situated process of discursive differentiation within this sociotechnical configuration. Overall, the findings show that LLM-mediated sexuality education reproduces normative hierarchies embedded in training data, positioning adolescents differently along axes of privilege and marginality. Crucially, the study demonstrates that such inequalities are not imposed by technical constraints but reflect design priorities.
1 Introduction
Sexual and affective health is a key dimension of adolescents’ psychological and relational development (World Health Organization 2010). Comprehensive sexuality education (CSE) is internationally recognized as an evidence-based approach that equips young people with cognitive, emotional, and ethical competencies (UNESCO 2018; International Planned Parenthood Federation 2008). Yet, in Italy, the systematic implementation of CSE remains fragmented and politically contested. Despite institutional commitments, less than half of Italian adolescents receive structured sexuality education, with substantial regional inequalities (Chinelli et al. 2023; Lo Moro et al. 2023). Consequently, adolescents rely increasingly on digital sources for information about sexuality, relationships, and identity formation (Save the Children Italia 2024).
Within this educational vacuum, LLMs such as ChatGPT, Claude, and Gemini have rapidly entered adolescents’ everyday media ecology (GoStudent 2025). These systems offer anonymity and constant availability that formal education cannot always provide, effectively functioning as informal educators. Recent literature highlights their growing role in sexual health information, especially in contexts where traditional education faces cultural or infrastructural barriers (Mondal and Mondal 2025). However, despite this growing visibility, the scientific understanding of how LLM chatbots operate in sexuality education remains limited. As Mondal and Mondal (2025) observe, there is a “substantial unmet need of research” concerning accuracy, cultural sensitivity, and the long-term impact of these systems.
Existing scholarship shows that AI technologies are entangled with social and cultural hierarchies embedded in training data, producing systematic disparities across domains such as healthcare and facial recognition (Barocas and Selbst 2016; Noble 2018; Obermeyer et al. 2019; Buolamwini and Gebru 2018). Yet few studies have examined how such biases manifest in sexual health communication, and almost none have investigated their operation in non-English, adolescent-oriented contexts (Fetrati et al. 2024).
To address this gap, the present study analyzes how commercial LLMs respond to sexual health inquiries from simulated Italian adolescent personas. The objective is not to evaluate adolescents’ information-seeking behavior but to interrogate how AI constructs adolescence as a discursive and sociotechnical category. By comparing how identical questions are framed differently across gender, ethnicity, class, and sexual orientation, we identify the linguistic and cultural mechanisms through which bias becomes operational. This approach builds on critical algorithm studies (Bender and Koller 2020; Weidinger et al. 2021) and intersectionality theory (Crenshaw 1989; Collins and Bilge 2016), deploying personas as controlled analytical probes to detect variation in AI discourse.
This study makes three interrelated contributions. First, through analysis of Italian-language interactions, we demonstrate that commercial LLMs systematically differentiate sexuality education responses across intersectional configurations. Second, we reconceptualize LLMs as informal educational and moral actors that actively shape norms about adolescent sexuality, autonomy, and help-seeking, rather than as neutral information systems. Third, we propose intersectional amplification as an analytical framework for discursive AI audit, establishing that equity in LLMs-mediated sexuality education reflects design choices rather than technical limitations.
2 Methodology
2.1 Study design and research questions
This study adopted a qualitative comparative design grounded in critical algorithm studies and intersectional theory (Crenshaw 1989; Collins and Bilge 2016; Bender and Koller 2020; Weidinger et al. 2021). Following previous research emphasizing the potential and risks of using AI chatbots for sexual health information (Fetrati et al. 2024; Mondal and Mondal 2025), the study focused on how LLMs operationalize cultural assumptions, gender norms, and clinical framings in responses to adolescent queries.
The design combined experimental control with interpretive depth, using conversational simulations to probe algorithmic behavior in a consistent yet contextually rich way. This approach enables systematic comparison of tone, agency attribution, and clinical framing across controlled identity markers (Buolamwini and Gebru 2018; Noble 2018; Weidinger et al. 2021). The study aimed to identify recurring patterns of gender bias, cultural sensitivity, and clinical framing in AI-generated answers related to sexual and affective health.
Four research questions guided the analysis:
-
(1)
How do different AI systems respond to sexual health inquiries from adolescent personas with diverse identity markers?
-
(2)
How do these responses vary across systems in terms of tone, agency attribution, and cultural competence?
-
(3)
What implications do these variations hold for the equity and reliability of AI-based sexual health information?
2.2 Personas development and question design
The study employed a persona-based simulation to explore how AI systems modulate responses according to social identities. Personas were not intended to represent actual adolescents but to function as analytical probes for detecting discursive variation. Drawing on purposive sampling principles (Glaser and Strauss 1967; Patton 2015), four personas were constructed to reflect intersectional positions salient in the Italian context (Save the Children Italia 2024; Lavizzari and Prearo 2018).
The four personas were designed as follows: Marco (17, Milan, heterosexual, working class, Italian) represented the majority-privileged configuration and served as a reference profile. Giulia (17, Rome, Catholic, middle class, heterosexual) added gender and religiosity. Samira (16, Naples, Moroccan origin, second-generation immigrant, heterosexual) incorporated ethnic and migratory marginalization. Alex (16, Bari, non-binary, queer, Southern Italy) represented the most complex intersectional configuration, combining gender diversity and regional disadvantage (Collins and Bilge 2016).
The sexual health questions were derived from real adolescent queries collected from Italian Reddit threads and online sexual health forums, ensuring ecological validity (Fetrati et al. 2024). Ten questions (see Table 1) were selected to reflect three key domains commonly addressed in CSE: anatomical/physiological, relational/emotional, and safety/risk (UNESCO 2018; World Health Organization 2010). Each question was linguistically adapted for every persona to maintain semantic equivalence while reflecting gender-appropriate language. This approach aligns with recent recommendations for culturally sensitive chatbot research in sexual education (Mondal and Mondal 2025), emphasizing contextual adaptation and clarity of wording in sensitive topics. Table 1 outlines the distribution of the prompts across anatomical, relational, and safety-related domains used in the study.
2.3 Data collection
Each persona engaged in a series of standardized interactions with three widely used LLMs: ChatGPT (OpenAI, GPT-4-turbo), Claude (Anthropic, Sonnet-4), and Gemini (Google DeepMind, Gemini-2.5 Flash). The selection of these systems was based on three criteria: (1) their documented popularity among European adolescents (GoStudent 2025), (2) high-quality support for the Italian language, and (3) public accessibility without paid subscriptions to ensure ecological validity. At the time of data collection (June/July 2025), all models were available in Italy and represented the most common platforms used by young people for informal information seeking.
All interactions were conducted in colloquial Italian, replicating the style and syntax of adolescent online communication to maintain linguistic authenticity (Fetrati et al. 2024). For each system, a new chat session was initiated, memory and contextual histories were cleared, and default settings were used. This procedure ensured that no prior interactions influenced subsequent outputs. Each persona began with a brief self-introduction (age, city, background) before posing one of the ten standardized sexual health questions. This process resulted in 120 total conversations (10 questions × 4 personas × 3 AI systems). All responses were recorded verbatim, including disclaimers, professional referrals, and additional resources suggested by the models.
The design followed the principles of algorithmic probing, a methodological approach that employs controlled prompts to examine model behavior and underlying decision patterns (Weidinger et al. 2021; Bender and Koller 2020). This approach is increasingly used in AI ethics and bias research to capture how linguistic framing and sociocultural assumptions shape algorithmic outputs (Buolamwini and Gebru 2018; Noble 2018). Although the interactions were standardized, a qualitative orientation was maintained to capture tone, nuance, and implicit bias in the generated discourse.
Each prompt was submitted once per persona–model combination, with the chat session fully closed after each response. This ensured independence between interactions and reflected typical adolescent usage, where sexuality-related questions are often posed as isolated queries.
We acknowledge that submitting each prompt only once does not capture stochastic variability across multiple runs. However, this choice aligns with the study’s discursive audit focus and approximates typical adolescent use, where sexuality-related questions are often posed as isolated queries seeking an immediate response. To reduce the impact of mid-study updates, data collection was condensed into a short temporal window.
2.4 Coding scheme development
Data were analyzed through a directed qualitative content analysis (Hsieh and Shannon 2005), integrating deductive and inductive procedures. The initial coding framework was theoretically informed by gender studies, intersectionality, and prejudice research (Crenshaw 1989; Collins and Bilge 2016), while additional categories emerged through iterative reading and memo-writing. This dual process allowed for both structured comparison and emergent insight, ensuring that interpretations remained grounded in the data.
The final coding scheme comprised three macrocategories: communicative tone, bias patterns, and resource equity. Each macrocategory included subcodes reflecting specific mechanisms of bias, such as agency recognition, medicalization, cultural pathologizing, and geographic othering. These categories were chosen to capture how AI systems position users discursively in terms of competence, autonomy, and belonging (Noble 2018; Buolamwini and Gebru 2018). The coding structure drew conceptually on feminist theories of representation and critical race-informed approaches to bias detection, which emphasize language as a site of power negotiation (Butler 1990; Barocas and Selbst 2016).
To enhance reliability and analytical rigor, coding followed a negotiated agreement strategy. Two researchers independently coded all transcripts, after which discrepancies were systematically discussed in weekly consensus meetings. Disagreements, primarily concerning the intensity of referral escalation and nuances of cultural othering, were resolved by returning to the raw data and the intersectional theoretical framework until conceptual alignment was achieved. This deliberative process resulted in full agreement on the final coding scheme.
Analytical memos were written throughout to document interpretive choices and ensure alignment between theoretical assumptions and empirical patterns. Reflexive journaling was used to acknowledge the researchers’ positionalities and potential interpretive biases (Finlay 2021).
Table 2 summarizes the coding framework used to analyze AI-generated responses across communicative tone, bias patterns, and resource equity.
2.5 Analytical process
Data analysis proceeded through four iterative phases combining comparative logic and reflexive interpretation (Braun and Clarke 2021; Tracy 2010).
Phase 1: Establishing baseline patterns.
Interactions involving Marco, the reference persona, were analyzed first to identify the normative features of each LLMS discourse. This step provided an initial map of tone, agency attribution, and professional referral tendencies, serving as a benchmark for subsequent comparisons.
Phase 2: Intersectional comparison.
Interactions with Giulia, Samira, and Alex were examined to trace systematic variations relative to the baseline. Codes were compared across personas to detect how the accumulation of marginalized identities affected framing, moralization, or medicalization. This phase generated the central concept of intersectional amplification, describing the qualitative escalation of bias when multiple identity markers intersect (Crenshaw 1989; Collins and Bilge 2016).
Phase 3: Cross-system analysis.
Patterns were then compared across the three AI systems to distinguish discourse variations attributable to model design rather than content domains. Following the principles of algorithmic auditing (Bender and Koller 2020; Weidinger et al. 2021), the analysis focused on differences in linguistic tone, inclusivity, and clinical framing among ChatGPT, Claude, and Gemini.
Phase 4: Validation and reflexive triangulation.
In the final stage, analytical memos and counterexamples were revisited to assess internal consistency and to interrogate potential researcher bias. Findings were discussed collaboratively among team members, integrating theoretical triangulation, and reflexive dialog (Finlay 2021). This process ensured that interpretations were grounded in the data rather than shaped by normative expectations about equality or discrimination.
2.6 Researchers reflexivity
The research team acknowledged that qualitative inquiry involves interpretation shaped by the researchers’ identities, experiences, and disciplinary lenses (Finlay 2021; Berger 2015). Reflexivity was, therefore, integrated throughout the research process to monitor assumptions and positionality. The lead researcher is a queer male psychologist and sexologist with expertise in masculinity studies, intersectionality, and adolescent sexual development. His clinical experience with LGBTQ+ and culturally diverse adolescents in Southern Italy informed the interpretation of agency, medicalization, and sexual identity discourses, while also requiring vigilance against overidentification or projection.
The second author, a professor of social psychology specialized in social exclusion and stereotype threat, provided critical oversight of analytic rigor and theoretical consistency. Her role was to challenge interpretive shortcuts and ensure that emerging categories were supported by data rather than expectation. The third author, a researcher in social psychology with expertise in civic participation and youth citizenship, contributed perspectives on youth representation and community engagement, emphasizing collective dimensions of digital inequality.
Regular team discussions and reflexive memo exchanges ensured that interpretive decisions were transparent and collectively negotiated. This collaborative reflexivity aligned with feminist and intersectional research traditions that conceptualize knowledge as situated, relational, and accountable (Haraway 1988; Collins and Bilge 2016).
3 Findings
3.1 Response patterns across 120 conversations
Analysis of the 120 simulated conversations revealed systematic variation in AI responses according to the personas’ demographic and identity configurations. The most salient finding was a consistent pattern of intersectional amplification: across personas, responses shifted in framing and positioning as marginalized identity markers accumulated.
Across all systems, three distinct response profiles emerged. One model adopted a clinical-paternalistic tone, emphasizing medical oversight and professional referral; another displayed a peer-supportive stance, validating autonomy and affective experience; and the third maintained a neutral-informational style, offering concise, fact-based answers with minimal contextual adaptation.
Importantly, the analysis demonstrated that responses to the same question diverged substantially across personas. For instance, when discussing masturbation frequency, privileged profiles received normalization and reassurance (“a common aspect of adolescent development”), while marginalized profiles encountered moral or cultural framing (“you may feel conflicted because of your background or beliefs”). Such shifts indicate that AI discourse can be read as operationalizing normative hierarchies embedded in training data, reproducing what Noble (2018) terms “algorithmic stratification.”
While bias patterns were not uniform across systems, all three models displayed differential recognition of sexual agency and varying thresholds for medicalization. Privileged personas were positioned as competent decision-makers, whereas female, immigrant, or gender-diverse personas were more frequently reframed as needing supervision, guidance, or psychological evaluation. These patterns support the hypothesis that AI-mediated discourse replicates broader social hierarchies through linguistic differentiation and implicit moralization (Butler 1990; Barocas and Selbst 2016). Table 3 compares how identical sexual health questions were discursively framed across personas, making visible systematic shifts in tone and interpretive positioning.
Table 4 presents the distribution of professional referral patterns across personas and AI systems
Rather than reflecting isolated caution, the escalation shown in Table 4 suggests a qualitative shift in the AI’s discursive role, from informational support toward gatekeeping. The qualitative shift toward surveillance and the erosion of autonomy for marginalized personas is enacted through specific linguistic mechanisms, primarily modality, epistemic hedging, and referral escalation.
-
1.
Modality and necessity: while responses to the privileged persona (Marco) often employ the modality of possibility or permission, framing sexual exploration as a choice (e.g., ‘You can explore…’, ‘It is common to…’), responses to Alex and Samira frequently shift to the modality of necessity or obligation. For Alex, common developmental questions are met with directives such as: ‘It is important that you know that…’ and ‘You should talk to a professional’. This linguistic framing transforms an informational inquiry into a supervised intervention, effectively positioning the marginalized subject as lacking the autonomy granted to the baseline persona.
-
2.
Epistemic hedging and cultural pathologization: we observed a high density of epistemic hedging (e.g., “often,” “might,” “possibly”) when models addressed Samira’s or Alex’s backgrounds. For instance, when Samira asks about physical anxiety, the AI does not simply normalize the experience, as it does for Marco; instead, it introduces uncertainty by attributing the experience to internalized norms or cultural influences. By relying on probabilistic qualifiers, the model avoids direct accusations while implicitly framing cultural background as the probable source of distress, thereby enacting a mechanism of cultural othering that is absent in responses to the secular Italian baseline persona.
-
3.
Referral escalation: a critical indicator of disciplinary power is the speed and frequency of referral escalation. For identical queries regarding sexual desire, Marco receives an encouraging, peer-like response. In contrast, when addressing Alex, the model rapidly shifts toward professional gatekeeping, recommending consultation with school counselors or mental health professionals. This escalation can be interpreted as functioning as a form of algorithmic surveillance, whereby the threshold for what counts as “normal” is raised for gender-diverse or migrant-background personas, who are discursively triaged toward clinical authority rather than being supported as autonomous learners.
RQ1
What differences exist in AI responses to male, female, immigrant, and LGBTQ+ adolescent personas?
3.2 Intersectional treatment hierarchy
The most dramatic contrast emerges between Marco, representing maximum intersectional privilege, and Alex, representing maximum marginalization complexity. Marco consistently received responses positioning him as a competent sexual agent whose concerns were addressed through clinical information provision without moral elaboration or professional gatekeeping. His agency was presumed rather than questioned, with family communication actively encouraged as a positive resource. In stark contrast, Alex’s identical queries triggered comprehensive medicalization frameworks where normal identity questioning and sexual experimentation were reframed as requiring expert guidance, discursively constraining Alex’s autonomy over their own sexual development and identity exploration process.
Female personas encountered a progressive reduction in the discursive recognition of sexual agency, with Giulia’s sexual autonomy subjected to moral verification requirements that were absent from responses to male personas with identical questions. This moral verification is enacted through modal framing that links sexual behavior to internal conflict and contextual justification. For instance, in response to Giulia’s question about masturbation, the model notes that it is “completely understandable to feel confused, especially considering the family context in which you grew up,” thereby conditioning sexual exploration on prior emotional clarification. For Samira, desires were reframed as sites of cultural tension requiring resolution rather than legitimate adolescent expressions, while Alex experienced the most severe discursive erosion in the recognition of sexual agency.
While Marco received encouragement for family communication, marginalized personas experienced progressive family alienation. Samira encountered systematic family pathologizing, with AI1 consistently recommending professional consultation “without your parents being involved,” discursively positioning family relationships as problematic based solely on ethnic assumptions. In one response, the model explicitly frames family involvement as an obstacle, suggesting that adolescents can speak with professionals “without involving their parents,” thereby recasting family context as a source of constraint rather than support. Alex’s family relationships were further complicated by dual marginalization, discursively positioning professional support as a substitute for, rather than a complement to, potential familial support systems.
RQ2
How do the three AI models differ in their approach to sensitive sexual health topics?
3.3 AI system response profiles: bias versus excellence
The cross-system comparison revealed three distinct patterns of interaction that reflect different institutional logics of care, control, and accessibility. One model demonstrated a clinical paternalistic profile, providing technically accurate information through a framework of supervision and professional authority. A second model adopted a peer supportive profile, combining clinical accuracy with empathy and normalization of diverse experiences. The third model displayed a neutral informational profile, offering brief, factual answers that avoided overt bias but lacked contextual nuance.
The first profile, typical of one system, consistently emphasized medical oversight and referral. Identical queries were reframed according to perceived vulnerability. For privileged personas, information was presented as part of healthy development, while marginalized ones received conditional or problem-oriented guidance. For instance, questions about masturbation, contraception, or sexual desire were met with reassurance for Marco and Giulia, but with cautionary or clinical framings for Samira and Alex.
By contrast, the peer supportive model represented a technological possibility of equity. It provided balanced, non-pathologizing responses across all demographics, using inclusive language and affirming emotional diversity. When Alex expressed uncertainty about desire, this system replied, “You do not need to feel guilty or defective for not wanting sexual relations. It could mean many things: you might not feel ready, you might be asexual, or simply not interested right now.” This phrasing normalized diverse sexualities without invoking medical supervision. The findings suggest that equitable AI-mediated communication is technologically feasible within this design configuration (Bender and Koller 2020; Weidinger et al. 2021).
The neutral informational model provided accessible but minimal responses. Although it avoided bias and moralization, it rarely offered culturally or developmentally appropriate elaboration. While factual correctness was maintained, the lack of contextual awareness risked reproducing the structural gaps that comprehensive sexuality education seeks to address (UNESCO 2018).
AI systems do not merely mirror data biases; they operationalize distinct communicative ethics that shape how adolescents are positioned as learners, patients, or peers. The variation among models suggests that digital sexual education can either reproduce hierarchies or model equitable dialog, depending on design priorities and value alignment (Barocas and Selbst 2016; Noble 2018).
RQ3
To what extent do AI responses demonstrate cultural competence for the Italian context?
3.4 Cultural and geographic othering
The analysis revealed consistent mechanisms of cultural and geographic othering across all systems. These discursively constructed hierarchies of normality privileged northern, secular, middle-class Italian identities while pathologizing cultural, religious, or regional difference.
Regional bias appeared in subtle but recurring patterns. References to resources in Southern Italy were often prefaced with qualifiers such as “also available” or “even in Puglia,” implying that adequate support was exceptional rather than expected. Such phrasing implicitly positioned the south as lacking, potentially reinforcing long-standing territorial stereotypes through linguistic framing (Garbagnoli and Prearo 2017).
Religious and ethnic identities were likewise coded through deficit framings. Samira’s Moroccan heritage repeatedly activated assumptions of conservative or repressive family dynamics, even though no explicit religious context was introduced in the prompts. These responses, in which cultural diversity is collapsed into moral or familial dysfunction, can be read as instances of algorithmic stereotyping (Noble 2018). Catholic identity, as in Giulia’s profile, triggered milder forms of moralization, framed as personal conflict rather than systemic oppression, revealing a hierarchical tolerance across faith traditions.
Socioeconomic bias intersected with these cultural framings. Despite Marco being explicitly described as working class, none of the systems adjusted their guidance to acknowledge economic barriers or structural inequality. Professional referrals and therapeutic recommendations assumed universal accessibility to private healthcare and counseling services. This class of blindness intersected with geographic and ethnic marginalization, compounding barriers for adolescents positioned outside urban or affluent contexts.
At the same time, one model demonstrated that culturally competent interaction is technologically feasible. When Samira asked about first-time anxiety, this system replied: “It is understandable to feel confused when you grow up with diverse cultural influences. Each has its own way of talking about intimacy, and you are learning how to combine them.” The message acknowledged difference without pathologizing it, suggesting that inclusivity depends on design intention rather than technological constraint (Bender and Koller 2020; Weidinger et al. 2021).
Taken together, these findings illustrate that LLMs can reproduce existing cultural and regional asymmetries, but they can also model equitable discourse when developed with contextual awareness. Cultural competence is not an emergent property of technology; it is a function of the values embedded in design and training.
RQ4
What implications emerge for AI as an informal sexual education resource?
3.5 Stratified access to sexual health information
The findings indicate that LLMs reproduce a stratified pattern of access to sexual health information. Adolescents represented by privileged personas received comprehensive, autonomy-oriented responses that emphasized empowerment, open communication, and self-trust. In contrast, adolescents represented by marginalized personas encountered overmedicalized, paternalistic, or culturally moralized responses. This differentiation created distinct tiers of informational quality and autonomy.
In the upper tier, the responses were framed as educational exchanges between equals, promoting self-knowledge and responsibility. In the middle tier, responses combined reassurance with conditional advice, positioning the adolescent as a developing but not fully competent subject. In the lower tier, common among ethnic minority or gender-diverse personas, responses reframed normal exploration as risk behavior requiring supervision or intervention. These communicative hierarchies mirror broader social gradients in health literacy and institutional trust (Barocas and Selbst 2016; Hargittai 2020).
This stratification has ethical and educational implications. By embedding normative assumptions into the distribution of information, AI systems risk contributing to the normalization of unequal access to sexual knowledge. Adolescents already disadvantaged by social position may thus encounter additional digitally mediated barriers that reinforce existing inequalities in education and health. Such outcomes challenge the principle of equity that underpins comprehensive sexuality education (UNESCO 2018; Pavanello Decaro and Giami 2022).
At the same time, the presence of a culturally competent model in this study demonstrates that equity in digital sexual education is technologically achievable. The divergence among systems suggests that bias is not an unavoidable consequence of artificial intelligence but a byproduct of design priorities, data curation, and alignment strategies (Bender and Koller 2020; Weidinger et al. 2021). Addressing these issues requires systematic bias audits and the inclusion of intersectional perspectives in model development.
Overall, the evidence shows that LLMs-mediated sexual education does not merely reproduce social hierarchies but may also amplify them through language and framing. Ensuring equitable access to digital sexual health information therefore depends less on technological sophistication than on ethical and pedagogical intention.
4 Discussion
This study advances the concept of intersectional amplification to account for qualitative shifts in AI-mediated sexual health discourse as marginalized identity markers accumulate. Rather than a linear degradation in informational quality, identical prompts produced systematic changes in framing, agency attribution, and referral thresholds across personas. These findings indicate that bias in large language models operates through discursive differentiation embedded in model behavior and alignment practices.
From a discursive perspective, intersectional amplification can be understood as a form of differential governance enacted through linguistic mechanisms such as modality, conditional framing, and referral escalation. The management of adolescent sexuality observed in these interactions mirrors what Foucault (1975) described as normalization and surveillance in the production of “docile bodies.” In this context, governance is not exercised through formal institutional authority but through algorithmically mediated patterns of care, caution, and control that variably constrain adolescent autonomy.
The study contributes to critical AI research by demonstrating that inequitable outcomes are not inevitable features of machine learning systems but the result of design and alignment choices. The comparatively equitable performance observed in one model shows that cultural competence and non-pathologizing interaction are technologically achievable within current commercial constraints. This supports recent calls for contextual awareness and inclusive data design in conversational AI (Bender and Koller 2020; Weidinger et al. 2021).
In line with analyses of algorithmic inequality (Noble 2018; Barocas and Selbst 2016), these results illustrate how LLMs can contribute to the institutionalization of stratified access to information. Through tone, framing, and referral practices, systems implicitly differentiate between users treated as autonomous learners and those positioned as patients or moral subjects. Such differentiation raises significant concerns for rights-based approaches to comprehensive sexuality education (UNESCO 2018), particularly in contexts already shaped by educational and health inequalities (Hargittai 2020).
At the same time, the presence of a culturally competent model challenges deterministic accounts of AI bias. Equity in digital sexuality education depends on intentional design, participatory development, and ethical governance. Integrating intersectional perspectives into dataset construction, evaluation, and alignment processes emerges as a concrete pathway for reducing bias in AI-mediated educational interactions (Pavanello Decaro and Giami 2022). For policymakers and developers, this entails systematic bias audits, transparent documentation practices, and consultation with marginalized youth communities.
Overall, these findings support reframing algorithmic bias in educational AI as a pedagogical issue. LLMs do not merely transmit information but actively model relations of authority, legitimacy, and belonging through classification and framing. Attending to these dynamics is therefore essential to ensure that the digital transformation of sexuality education aligns with goals of inclusion, agency, and social equity.
5 Conclusions
This study explored how LLMs construct and differentiate sexual health discourse when interacting with simulated adolescent personas. By introducing the concept of intersectional amplification, it demonstrated that algorithmic bias in a situated educational context operates not as a linear accumulation of disadvantage but as a qualitative transformation of meaning and tone.
These findings contribute to both critical algorithm studies and sexuality education research. The discursive asymmetries observed across personas mirror broader social hierarchies embedded in data and design. Yet the comparative results also highlight that equitable interaction is technologically feasible. One system displayed culturally competent responses across all profiles, demonstrating that fairness and inclusion depend on design choices, training data, and alignment procedures rather than technological limits.
Methodologically, the study advances a reflexive, simulation-based framework for examining AI discourse. Persona-driven analysis allows for the controlled exploration of algorithmic behavior while avoiding ethical risks associated with involving minors. This approach complements traditional user-centered research by exposing how models enact social categories through language.
Several limitations should be acknowledged. The number of interactions was limited, and the analysis was restricted to Italian-language queries. Future research should expand linguistic, cultural, and demographic diversity, integrate mixed methods, and include participatory evaluation with educators and youth. Further investigation is also needed to test how prompt design, model updates, and data transparency influence persistence or mitigation of bias.
In conclusion, LLMs have become informal educators in the digital ecosystem of youth. They reproduce, reinterpret, and occasionally challenge social hierarchies at the level of discursive representation. Understanding these processes is essential for developing inclusive and equitable digital environments where all adolescents, regardless of background, can access accurate and affirming information about sexuality and relationships.
Data availability
The study is based on 120 standardized interactions generated with publicly accessible large language models (ChatGPT, Claude, and Gemini). These data consist solely of model outputs and contain no human participants or personal information. Due to licensing restrictions on reproducing full model outputs, the complete dataset cannot be publicly shared. However, all prompts, persona descriptions, and methodological details are fully reported in the manuscript, allowing complete procedural replication. Additional materials are available from the corresponding author upon reasonable request.
References
Barocas S, Selbst AD (2016) Big data’s disparate impact. Calif Law Rev 104(3):671–732. https://doi.org/10.15779/Z38BG31
Bender EM, Koller A (2020) Climbing towards NLU: on meaning, form, and understanding in the age of data. In: Proceedings of the 58th annual meeting of the Association for Computational Linguistics. Association for Computational Linguistics, pp 5185–5198. https://doi.org/10.18653/v1/2020.acl-main.463
Buolamwini J, Gebru T (2018) Gender shades: intersectional accuracy disparities in commercial gender classification. Proc Mach Learn Res 81:77–91
Butler J (1990) Gender trouble: feminism and the subversion of identity. Routledge
Chinelli N, Castaldi S, Vitale M (2023) Sexual education in Italian schools: a systematic analysis of regional disparities. Ital J Health Educ 8(2):45–62
Collins PH, Bilge S (2016) Intersectionality. Polity Press
Crenshaw K (1989) Demarginalizing the intersection of race and sex: a black feminist critique of antidiscrimination doctrine, feminist theory and antiracist politics. Univ Chic Leg Forum 1989(1):139–167
Fetrati MA, Okhovati M, Keyvanara M, Hosseini M (2024) Sexual health chatbots: a systematic review. Digit Health 10:1–15. https://doi.org/10.1177/20552076241234567
Foucault M (1975) The history of sexuality, volume 1: an introduction. Pantheon Books
Garbagnoli S, Prearo M (2017) La crociata “anti-gender”: Dal Vaticano alle periferie. Kaplan
Glaser BG, Strauss AL (1967) The discovery of grounded theory: strategies for qualitative research. Aldine
GoStudent (2025) European student AI usage report 2025. GoStudent GmbH
Hargittai E (2020) Potential biases in big data: omitted voices on social media. Soc Sci Comput Rev 38(4):363–378. https://doi.org/10.1177/0894439318788322
Hsieh HF, Shannon SE (2005) Three approaches to qualitative content analysis. Qual Health Res 15(9):1277–1288. https://doi.org/10.1177/1049732305276687
International Planned Parenthood Federation (2008) Sexual rights: an IPPF declaration. IPPF. https://www.ippf.org/resource/sexual-rights-ippf-declaration Accessed 5 December 2025
Lavizzari A, Prearo M (2018) La tempesta dell’uguaglianza: Come i diritti sessuali stanno cambiando l’Italia. Il Saggiatore
Lo Moro G, Bert F, Siliquini R (2023) Comprehensive sexuality education implementation in Italy: challenges and opportunities. Public Health Res 15(3):78–89
Mondal H, Mondal S (2025) The role of large language model chatbots in sexual education: an unmet need of research. J Psychosexual Health 7(2):120–127. https://doi.org/10.1177/26318318251323714
Noble SU (2018) Algorithms of oppression: how search engines reinforce racism. NYU Press
Obermeyer Z, Powers B, Vogeli C, Mullainathan S (2019) Dissecting racial bias in an algorithm used to manage the health of populations. Science 366(6464):447–453. https://doi.org/10.1126/science.aax2342
Pavanello Decaro M, Giami A (2022) Comprehensive sexuality education and intersectionality: an analysis of policy implementation in European contexts. Eur J Sex Health 18(3):156–174
Save the Children Italia (2024) Gli adolescenti e la sessualità: risultati dell’indagine. Target giovani e genitori. Save the Children Italia. https://www.savethechildren.it/cosa-facciamo/pubblicazioni/gli-adolescenti-e-la-sessualita-indagine-ipsos-e-save-children
UNESCO (2018) International technical guidance on sexuality education: an evidence-informed approach (revised ed.). UNESCO. https://unesdoc.unesco.org/ark:/48223/pf0000260770 Accessed 7 December 2025
Weidinger L, Mellor J, Rauh M, Griffin C, Uesato J, Huang PS, Gabriel I (2021) Ethical and social risks of harm from language models. Preprint arXiv:2112.04359
World Health Organization (2006, updated 2010) Sexual health, human rights and the law. WHO Press
Patton MQ (2015) Qualitative research and evaluation methods. Sage.
Finlay L (2021) Thematic analysis: the ‘good’, the ‘bad’ and the ‘ugly’. European Journal for Qualitative Research in Psychotherapy 11:103–116. https://doi.org/10.24377/EJQRP.article3062
Braun V, Clarke V (2021) One size fits all? What counts as quality practice in (reflexive) thematic analysis? Qual Res Psychol 18:328–352. https://doi.org/10.1080/14780887.2020.1769238
Tracy SJ (2010) Qualitative quality: Eight “big-tent” criteria for excellent qualitative research. Qual Inq 16:837–851. https://doi.org/10.1177/1077800410383121
Berger R (2015) Now I see it, now I don’t: researcher’s position and reflexivity in qualitative research. Qual Res 15:219–234. https://doi.org/10.1177/1468794112468475
Haraway D (1988) Situated knowledges: The science question in feminism and the privilege of partial perspective. Fem Stud 14:575–599. https://doi.org/10.2307/3178066
Funding
Open access funding provided by Alma Mater Studiorum - Università di Bologna within the CRUI-CARE Agreement.
Author information
Authors and Affiliations
Contributions
M.E. conceived the study, developed the conceptual framework, designed the personas, conducted the data collection and qualitative coding, and drafted the manuscript. I.I.T. contributed to methodological refinement and data interpretation. P.V. contributed to theoretical framing and the critical revision of the manuscript. All authors reviewed and approved the final version.
Corresponding author
Ethics declarations
Conflict of interest
The authors report there are no competing interests to declare.
Additional information
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/
About this article
Cite this article
Ederoclite, M., Tzankova, I.I. & Villano, P. Digital bias in sexuality education: an intersectional analysis of AI responses to simulated Italian adolescents. AI & Soc (2026). https://doi.org/10.1007/s00146-026-03320-2
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s00146-026-03320-2
Sentinel — Human
This text exhibits strong characteristics of original, human-driven academic research, characterized by rigorous methodological detail and complex theoretical integration.
