Can AI reduce gender bias? The case of AI-generated audio conversations in Google's NotebookLM
Malekzadeh, Milad (2026-08)
Malekzadeh, Milad
08 / 2026
Computers in Human Behavior Reports
101246
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi:tuni-202608249236
https://urn.fi/URN:NBN:fi:tuni-202608249236
Kuvaus
Peer reviewed
Tiivistelmä
Voice-based generative AI increasingly mediates how people encounter scientific knowledge, but its conversational designs may reproduce gendered patterns of authority even when content appears neutral. This study audits Google NotebookLM's Audio Overview feature to examine whether AI-generated podcast-style conversations distribute speaking time, epistemic authority, and communicative roles equitably across woman-coded and man-coded voices. 100 default audio summaries were generated from anonymized scientific papers, then transcribed, diarized, and coded the conversations for speaking time, technical level, speech acts, utterance-level epistemic positioning, and conversation-level roles. To contextualize the findings, the same workflow was applied to 100 episodes from highly ranked human-produced science podcasts. NotebookLM exhibited near parity in airtime: woman-coded voices accounted for 48.9% of speaking time and man-coded voices for 51.1%, with no statistically reliable difference after correction for multiple comparisons. By contrast, in the human podcast reference corpus, woman-coded speakers accounted for 38.4% of gender-coded speaking time and man-coded speakers for 61.6%, and this imbalance remained statistically detectable after correction. For the principal comparisons concerning self-positioning as expert, conversation-level epistemic authority, communicative roles, technical level, question proportion, and speech-act distributions, the corrected analyses did not provide statistically reliable evidence of gender-coded differences. These findings indicate that the clearest contrast between the audited NotebookLM outputs and the human podcast reference corpus concerned airtime: NotebookLM showed substantially greater gender-coded speaking-time parity, while evidence of broader interactional asymmetry was limited.
Kokoelmat
- TUNICRIS-julkaisut [25742]
