Abstract
What makes artificially generated podcasts feel natural, and what are their current limitations? In this exploratory study, we compare podcasts generated by Google’s NotebookLM Audio Overview feature with naturally produced podcasts to explore two key ingredients of human dialogue, turn-taking timings and contextually appropriate role behavior. Humans excel at predicting where a previous speaker ends their turn, which allows them to chime in with minimal overlap or gap between speaker turns. Humans also follow unwritten rules of conversation to avoid redundancy, provide relevant information, and adhere to socially appropriate behavior within the dynamics of an unfolding conversation. These key ingredients to human conversation pose significant challenges to artificial dialogue generation despite recent advancements in speech recognition, appropriateness of content delivery, and voice production. NotebookLM overcomes some of the traditional pitfalls of spontaneous dialogue generation through avoiding endpoint prediction by including turn-taking and backchanneling in its script writing. It does so by strictly adhering to avoidance of overlaps or gaps. This is unlike what we observe in human podcast dialogues. Where humans show an acute awareness of the changing dynamics between hosts and guests in conversation, and even overlap to show engagement, AI shows little flexibility. We argue that naturalness in dialogue requires a contextually appropriate and addressee-oriented approach to respond to how a conversation unfolds and not just getting the timings right. A novel human intervention feature shows that pre-determining role behavior and turn taking only goes so far in generating natural dialogue: NotebookLM clearly hits its limits when responding to unpredictable intervention. Variable turn-taking timings and adapting conversational role behavior in spontaneous conversation remain unmatched human skills that are indispensable for natural dialogues.
1 Introduction
In an age of blurring boundaries in human–machine interaction, the central question remains: What makes our human language experience unique? When we converse with a human being, we follow the rules of a game well-rehearsed and well-known. So much so that we typically know what to expect from our interlocutor and when to chime in. In this paper, we compare the conversational role behavior and turn-taking timings of natural dialogues in human podcasts with those of dialogues generated by Google’s NotebookLM. Initially released in 2023, NotebookLM (previously Project Tailwind) is Google’s attempt to enhance AI-driven knowledge interactions by grounding its responses entirely in sources provided by the user (Google 2023). While initially designed as a text-based research assistant, NotebookLM expanded its capabilities in September 2024 with the introduction of Audio Overview, or “Deep Dive” podcasts (Google 2024a), a feature that allows users to generate spoken explanations and summaries in the format of a podcast between two people. Our aim is to describe and quantify how the latter compares to human dialogue. The two central measures for this human–machine comparison are two hallmarks of human conversation: the near perfection of timings in turn taking, which typically has only one speaker speak at a time (Sacks et al. 1974) and requires a sophisticated mechanism for projecting points of transitions (Levinson and Torreira 2015), and the reliability of our behavior in conversation, which has been the subject of investigation of linguists ever since Wittgenstein formulated the first sketch of the rules of the language game, which has brought forth modern-day Dynamic Pragmatics (Murray and Starr 2021). The present investigation builds on this tradition of analyzing conversational dialogue with an integration of turn-taking timings and observations aligned with tools used in conversation analysis. We propose that machines need to go beyond perfecting the sound and meaning of human language to generate natural conversations; they also need to adhere to the dynamic rules of human conversation.
The remainder of this section provides an overview of previous research on the dynamics in human conversations, research on human turn taking behavior, advancements in machine-generated speech, and some recent attempts at creating naturalistic conversation experiences through artificial intelligence (AI). In Section 2, we then describe how we compared human and AI dialogue performance in Google NotebookLM’s approach through a range of linguistic measures of natural dialogue behavior, such as turn durations, overlaps, and conversational role behavior, i.e. taking on the behavior humans ascribe to hosts or guests on a podcast. Section 3 presents our results based on two human podcast excerpts compared with two NotebookLM podcasts generated from their transcripts. Section 4 briefly shows the impact of a new feature that allows live human involvement for the naturalness of the podcasts. Section 5 discusses the differences observed between human and AI discussants. In Section 6, we conclude.
1.1 The unwritten rules of the language game
Human conversation is structured around conventions to which participants adhere without ever having been told what they are. In a landmark publication, Grice (1975) described them with his four maxims centered on the cooperative principle: be informative, truthful, relevant, and clear. Lewis (1979) added the idea of scorekeeping to determine what is not only truthful, but also appropriate or acceptable to say in a conversation. So-called felicity conditions link particular constructions to the conventions and contexts in which they are used (Searle 1975; Farkas 2024). Later models then refined these proposals to give us elaborate descriptions and analogies of when we expect to hear, for example, a question vs. a statement, and how we respond to them. A vast body of research across different analytical frameworks has also investigated how we manage the timings of the back and forth between speaker and hearer (e.g., Auer 2005; Riest et al. 2015; Levinson and Torreira 2015; Kendrick et al. 2023). Together, these accounts have cemented the idea that conversations are rule-governed negotiations of how to grow a set of shared beliefs or to address a question under discussion (Farkas and Bruce 2010; Malamud and Stephenson 2015; Roberts 2012; Heim 2019). These conventions make conversational behavior predictable if its interlocutors are reliable in the sense that they conform to the unwritten rules. We will unpick each of these two aspects, predictability and reliability, in turn.
Human-to-human exchanges are predictable to the extent that they are governed by an intricate system of turn-taking conventions. These conventions allow participants to transition smoothly between speaking and listening while simultaneously allowing for the addressee to engage with the speaker through paralinguistic features, such as backchannels and laughter, which add to the dynamic and expressive quality of human dialogues without claiming the speaker role (Sacks et al. 1974; Schegloff 2000). Humans can exploit various morphosyntactic, semantic, and prosodic cues, which help listeners anticipate when they are expected to engage and take up the turn (Riest et al. 2015; Bögels and Torreira 2015). In co-present conversation, humans additionally benefit from a range of non-verbal cues, such as gaze and gesture, that help them to identify turn transitions (Cassell et al. 2001; Jokinen 2010; Kendrick et al. 2023). The management of this system is at the core of real-world conversations (Arora et al. 2025). Speakers anticipate subsequent contributions, so much so that responses are likely being prepared while the previous speaker still finishes their turn (Garrod and Pickering 2015). This is unexpected because average turns themselves are rather short: Levinson and Torreira (2015) report a median duration of 1227 ms (mean: 1680 ms) in their analysis of telephone conversations. To increase the likelihood of efficient conversational development, interlocutors also project where conversations go next (Auer 2005; Clayman 2012). Humans seem hardwired to excel at predicting conversational developments, with traces of conversational turn-taking already evident from birth (Dominguez et al. 2016; Legerstee et al. 1990; Reddy et al. 1997). Levinson and Torreira (2015) highlight that sensitivity to turn-end cues also develops early in life, with young children showing the ability to anticipate conversational boundaries with the help of these cues. Even before the arrival of multiword speech, children shift their gaze from speaker to addressee in anticipation of a response to a question (Casillas and Frank 2017). From the start, turn-taking in human conversation is a complex game of anticipating transitions based on typical conversational behavior and getting ready to respond.
Humans perform so well at this game that globally speaking, very little goes wrong. The seminal work of Sacks et al. (1974) established the understanding of turn-taking as an organized system governed by rules rather than random altercations. The strong hypothesis was that an average transition between interlocutors would not display a significant break (‘gap’) or any interruptions (‘overlap’). Although brief overlaps may well be common, the norm would be that only one party speaks at a time and turn transitions are characterized by a sensitivity from each party to encode and identify when to change turns, and to do so effectively (Sacks et al. 1974). Likewise, disruptive overlaps—through simultaneous starts of mid-turn vocalizations—should be avoided (Schegloff 2000) unless they signal addressee engagement. Non-intrusive overlaps during a speaker’s turn in the form of backchanneling can be added to further increase a sense of naturalness (Clark 1996), as long as they coincide with transition-relevance places (TRPs), i.e. natural points of potential transitions, to maintain a high degree of dialog fluidity.
The rigidity of this “no-gap, no-overlap” principle has been questioned by subsequent research which has shown that overlaps account for up to 40% of speaker transitions and that gaps and overlaps often exceed the reported 200 ms characterized as the average turn transition time (Heldner and Edlund 2010). Further research suggests that the gap between turns in conversations of speakers who share a native language tends to be around 200–250 ms (Gratier et al. 2015; Stivers et al. 2009). Levinson and Torreira (2015) report that in their sample of the Switchboard corpus (Calhoun et al. 2010) overlaps are common, but short (around 275 ms) as are overall turn durations (mean 1680 ms). A recent meta-analysis of 41 turn-taking studies revealed that mean and mode of turn transition are around 68–222 ms (Hoogland 2025). Regardless of the exact average, then, the overall pattern still holds: the average time between turns significantly falls below the average time we require to prepare our response (around 600 ms if we add up times for conceptual preparation, lemma retrieval, and form encoding (Indefrey and Levelt 2004, p. 108, cited in Levinson and Torreira 2015). It is important to note, however, that much of the debate around turn-transitions is tainted by different measurement techniques (Hoogland 2025) and that large deviations from these averages can be pragmatically meaningful and speech act specific (Roberts et al. 2011). Hoogland et al. (2023), for instance, report that closed questions tend to have shorter inter-speaker gaps than open questions, which seems pertinent in the context of a podcast format where open questions regularly serve as a means to provide a platform for a participant (Heiselberg and Have 2023). This would predict longer gaps in host–guest exchanges. Question–answer sequences, in general, show protracted gaps, especially in free conversation (Hoogland et al. 2023). Duncan’s (1972) signaling model offers an extension of Sacks et al. by highlighting the role of prosodic and non-verbal cues in managing turn-transitions (see also: Bögels and Torreira 2015), arguing that turn-taking is not only governed, but also mediated by these ‘turn-yielding signals’ (falling intonation, slowing speech rate) and ‘attempt suppressing signals’ (increased loudness, floor-holding gestures), and demonstrating how speakers use multimodal signals to anticipate and manage conversational flow (Hjalmarsson 2011). Expectations toward average turn-taking transitions should, therefore, reflect the type of conversation and who contributes to it. Some settings may require longer transition time, and if inter-speaker gaps are long (over 700 ms), they may indeed be pragmatically meaningful (Roberts et al. 2011).
Besides the predictability of turn-taking timings, interlocutors can also base their expectations for a conversation on the conversational role behavior of its participants. Who the interlocutor is and what they know, or rather, what the speaker assumes their interlocutor knows, significantly shapes the flow of dialogue (Sidnell 2012). Depending on whether or not the interlocutor counts as a reliable source, we adjust our own expectations as to what information they can provide (Gunlogson 2008; Heim 2025). Conversations can, therefore, be conceptualized as reducing a knowledge asymmetry between the interlocutors whereby this asymmetry denotes the differences in expertise, information access, or understanding between individuals engaged in interaction (Sidnell 2012; Osa Gomez del Campo 2020; Heim and Wiltschko 2020). These differences can manifest as explicit roles—such as doctor-patient, or expert-novice—or emerge organically within conversations, based on participants' backgrounds and experiences (Jacobsen 2014; Kim 2019). Knowledge asymmetry is grounded in the concept of common ground (Stalnaker 1978; Murray and Starr 2021)—the shared body of assumptions and information that speakers rely on for mutual understanding, which is dynamically updated as a speaker asserts, challenges, or accommodates new information (Bavelas et al. 2012). For a speaker to contribute to expanding the common ground, they need to have information that addresses the current topic of conversation (Roberts 2012), and that the addressee does not already know (Grice 1957; Westera 2017).
Communication requires an ongoing process of alignment to ensure understanding between participants, a process which has been described as grounding (Clark and Brennan 1991; Bavelas et al. 2012). Grounding emphasizes that for effective communication, participants must not only share information, but also ensure mutual understanding through continuous feedback and adjustment. This is facilitated through interactional language, such as sentence-final intonation, backchannels, confirmation requests, and other discourse particles (Heim et al. 2016; Wiltschko 2021). Through this, conversation participants collaboratively maintain and modify their common ground (Clark and Wilkes-Gibbs 1986), either reinforcing the existing knowledge asymmetries or working to bridge them through clarification and elaboration. Knowledge asymmetry can manifest through the roles of those participating in a conversation—where one speaker typically has more expertise than the other. Experts tend to control the topic progression, ask fewer clarification questions, and introduce more information through declarative sentences (Murray 2014), whereas novices rely more on questions and confirmation-seeking strategies (Heritage 2012). Moreover, asymmetric turn-taking rights feature in many institutional(ized) settings—where the expert speaker allocates turns and manages the flow of a discussion (Drew and Sorjonen 1997). Again, roles are not static and can change as a conversation develops. Interlocutors manage role transitions through linguistic cues, such as hedging, evidentials, and self-repair (Clark 1996; Yngve 1970; Sidnell 2012; Heim et al. 2025). We see then that established roles, while dynamic, not only affect who is distributing the knowledge, and the content of a dialogue, but also influence the entire structure and mechanics of dialogical interaction, including our ability to project its development. Awareness of common ground development and management provides a context that allows interlocutors to project how conversations will continue and how knowledge asymmetries may be resolved. This is less a matter of predicting the timing of individual turns than it is one of what type of contribution we can expect in response to our own contributions.
Predictability in public dialogue settings, such as podcasts, is particularly high due to the expectations any listener has toward the medium. It is worth remembering the broadcasting roots of podcasting to explain its register (Ehret et al. 2024) and conversation dynamics (Jarrett 2009). Despite modern attempts of breaking through to the listener through creating a growing sense of an intimate knowledge of the host, these roots shape the expectations we have toward conversational behavior. Although podcasts straddle the line between intimate and institutionalized engagement (Berg 2025), they follow social scripts and dialogic protocols that provide opportunity for projection. Hosts take on the role of the expert in establishing the topic of conversation and then soon transition into a novice role because they seek to provide a platform for the invited guest or expert to contribute to the conversation (Heiselberg and Have 2023; Schaefer and Loewe 2024). Podcast hosts, for example, will introduce and provide a conversational platform for a guest that is known to its listeners for particular perspective or area of expertise. Listeners can also expect hosts to take on a leading role in structuring the conversation through their questions, attempts of clarification, and encouragements to expand (Kangulu 2025). Self-disclosure is another prominent host feature, and taking on an expert perspective is just as much part of their podcasting repertoire (Jarrett 2009). Because expert perspective and subjective disclosure can go hand in hand, the same speaker can shift from expert to novice within a single turn if they ask their interlocutor to relate to their perspective through their own experiences. The role expectations baked into the genre of dialogical podcasts, therefore, add to the predictability of interlocutor behavior despite the frequent role shifts from expert to novice.
1.2 Simulating conversation through artificial intelligence
When designing artificial spoken dialogue systems, the main aim is to “achieve a high degree of human-likeness” (Mitsui et al. 2023; Liu 2024; Kolekar et al. 2024) so that users feel as if they are talking to a real human rather than a computer, as would be expected by the user (Lala et al. 2018). While machines have recently managed to master generating easily understandable speech, there is more to generating or simulating natural and meaningful conversation (Kolekar et al. 2024). Following from the description of natural dialogue dynamics above, we can infer that generating spoken dialogue involves not only the crafting of coherent linguistic content, but also the coordination of timing, prosody, and turn-taking (Liu 2024)—all of which contribute to the perceived naturalness of dialogue (Hung et al. 2009; Skantze and Hjalmarsson 2013). In non-scripted, live generation of language, spoken dialogue systems typically employ automated speech recognition software (ASR), which transcribes the incoming audio into text—making it easier to process. Modern ASR systems, often based on end-to-end neural models need to determine when to segment the stream of speech signal into meaningful units that can be processed. Much of the research surrounding spoken dialogue systems has focused on addressing the critical issue at the interface between ASR and dialogue management, that is, how exactly the system can determine when a user has stopped speaking (Lala et al. 2018). This is commonly associated with endpointing (Chang et al. 2022), and end-of-turn detection—the former focusing on detecting the end of speech segments, while the latter considers conversational cues to more accurately predict turn completion. Endpointing can, therefore, be understood as the equivalent of anticipating a turn transition in human-to-human dialogue. Though relatively common in the latter, the impracticability of interrupting a user mid-turn leads many systems to be conservative (Lala et al. 2018). In human-to-AI interaction, the most common approach is for the conversational agent to be silent for an extended period, therefore, meaning that the model takes significantly longer to take the floor than a human participant would, leading to unnaturally long transitions.
Conversational AI systems are trained on large corpora of data, and are designed to engage in fluid, natural interactions. Since their development in the 1960s (Wilks et al. 2011), they have evolved from early dialogue systems to high-quality end-to-end neural models (McTear 2020; Liu 2024) capable of exhibiting proficiency in producing coherent, contextually relevant, and stylistically adaptable textual outputs (Kokala 2024). Those early systems relied heavily on explicitly defined rules, fixed contexts, and constrained knowledge bases (Hung et al. 2009), which as a result meant that their dialogue was inherently rigid, as they lacked the flexibility needed to manage and respond meaningfully to open-ended, dynamic interactions or extend beyond their predefined domains, therefore, rendering them rather ineffective in real-world scenarios where nuance, spontaneity, and general contextual awareness are paramount. Contemporary systems are more refined and can generate syntactically fluent and semantically appropriate responses no matter the topic or task (Mitsui et al. 2023). One significant advancement that redefined the capabilities of these systems was the introduction of the transformer model (Vaswani et al. 2017), which now underpins the architectures of many state-of-the-art models, such as GPT, BERT, and Gemini. With the introduction of the transformer model came a self-attention mechanism, which allows the model to weigh the relevance of every word in an input sequence in parallel irrespective of its position, essentially enabling a more effective way to model context over extended spans of text by seeing how each word relates to the others. The result of this is that transformer models significantly enhance a system’s ability to understand and generate seemingly coherent responses across multiple turns. AI can manipulate symbols without any internal grasp of what those symbols mean (a phenomenon known as the ‘Chinese Room Problem’ (Searle 1980)). Hence, the term “understand” here denotes a model’s ability to capture and represent statistical patterns and contextual dependencies in language, not human-like semantic comprehension (Bender et al. 2021).
The arrival of the transformer model marked a shift away from the traditional recurrent neural networks and long short-term memory models that processed data sequentially and would struggle to capture long-range dependencies due to issues, such as vanishing gradients (where earlier information in a sequence becomes harder to retain during training) and limited memory capacity. For continuity in conversational role behavior and dialogue coherence, this advancement was essential. Nevertheless, coherence is primarily a context-dependent measure. Human-to-human dialogue is a jointly constructed activity (Le Bigot et al. 2007), whereas machine-generated dialogue lacks any real-world grounding. Therefore, where human speakers can seamlessly track conversational context, AI lacks implicit conversational memory—that is, the human ability to remember and mentally track what has been discussed, what the other conversation participant knows, and how the conversation is developing over time. This causes AI models to struggle with maintaining consistency, meaning that they often repeat information, forget certain details mentioned earlier in the dialogue, or shift topics disjointedly (Brabra et al. 2022). This problem arises from finite context windows. Essentially, the model can only see a certain portion of the conversation at a time, and when information falls outside of that window, its content is effectively non-existent. Hence, despite the tremendous advancements that have been made in the field of natural speech and dialogue generation (Kulkarni et al. 2019; Mahlsela et al. 2024), there remain areas where these dialogue systems struggle to truly replicate the seamless flow of human conversation.
What is more, AI lacks the intuitive real-time processing abilities that humans use to navigate conversational flow, and thus it is one of the primary reasons for the limitations of naturalness in AI dialogue systems (Skantze 2021). One particularly relevant limitation is the balancing of knowledge asymmetries between interlocutors, which affects how human speakers adjust their language, level of detail, and conversational role depending on their expertise relative to their interlocutor (Sidnell 2012; Jacobsen 2014). This is an area that AI seems to find challenging to handle consistently, as it is unable to reason the way a human would in conversation (Darapaneni et al. 2023). Therefore, a system may respond as an expert in one turn, but then ask a novice-like question about the same subject in the next turn, creating inconsistencies that disrupt the perceived naturalness of the interaction that would not necessarily exist in human-to-human exchanges. In the latter, role changes are often predictable through topic shift or genre-specific, social protocols of conversational behavior.
1.3 The present study: natural vs. generated conversations
Human dialogue is largely predictable in its timings, transitions, and interlocutor behavior. This includes role reversals that follow social scripts of particular genres whereby interlocutors alternate between expert and novice behavior at predictable points in conversation. The state of the art on research in human dialogue behavior, therefore, provides robust models for generating natural dialogue. Average transitions can be set around 200–250 ms (Levinson and Torreira 2015) and average turn durations at around 2 s (Meyer 2023). Non-intrusive overlaps, such as backchannels, should largely coincide with TRPs (Schegloff 2000; Levinson and Torreira 2015) rather than occurring at random points during a speaker’s turn. Speakers follow conversational scripts that respond to knowledge asymmetries and are sensitive to genres and (semi-) institutionalized contexts. Such a high degree of predictability of dialogical behavior is a potential solution to one of the remaining issues in artificial dialogue generation. Turn-by-turn podcast generation has struggled with coherent content crafting and situation-specific sound production due to large computational demands (Xiao et al. 2025). Google’s NotebookLM gets around this problem by planning dialogues not spontaneously, thereby reducing computational demands needed for endpointing. NotebookLM generates the dialogue based on a text source provided by the user; all timings, transitions, and role behavior are, therefore, pre-planned. In pre-planned language generation, such as a podcast simulation, determining the timing of turn transitions can, therefore, allow for more complex scenarios than spontaneous human-AI interactions. Accidental overlaps are avoidable in pre-planned dialogue, and the turn transitions can be included in the generation of the dialogue. Overall, therefore, the pre-planning of Google NotebookLM’s dialogue generation should make the key attributes of human conversation, such as avoiding intrusive overlaps and delivering appropriateness of conversational role behavior within genre-specific scripts, comparable between humans and agents.
The present study employs the turn taking model of one-party-at-a-time and the notion of context-sensitive, conventionalized role behavior as linguistic heuristics for natural, dialogical conversation. These heuristics map onto testable hypotheses for timings and speaker contributions, with a clear sense from the extensive literature that average turn taking and role behavior show considerable variation. We provide an exploratory comparison of spontaneous human dialogue with pre-planned, AI-generated dialogues in standard host and guest configurations characteristic of the podcast genre. When turning to NotebookLM’s attempts of modelling these dialogues artificially, we are mindful of several foundational differences between natural and artificial dialogue types, particularly in the podcast context that we chose to examine. Machine-generated dialogue is first and foremost based on the most likely continuation of a construction of topic independently of the immediate conversational context. Predictability of a conversation or any turn within a spontaneous conversation takes on a very different form in any spontaneous dialogues, regardless of whether they are produced by humans or machines. Hence, predictability is less context- and interlocutor-dependent in artificially generated dialogues than in human dialogue contributions. Second, we acknowledge that our exploratory comparison tests one specific genre of human dialogue—long-form podcast conversations where alignment between contributors and audiences is likely—and one specific strategy to generate dialogue, NotebookLM’s approach to pre-plan the entire dialogue before turning text to speech. These choices will unavoidably limit the generalizability of our findings to other genres or tools. Nevertheless, the genre and tool of choice speak clearly to the central question of what makes human language unique—or in our view, more natural than artificial language—and uses a popular tool to compare one possible strategy of AI dialogue generation to a well-grounded linguistic model on what governs human dialogues: timing, appropriateness, and predictability. This approach provides a novel perspective on how we can assess naturalness of machine-generated dialogue independently of the standard behavioral methods.
2 Methods
2.1 Data source
Our study is based on four podcast-styled dialogues based on two different topics. Two of those dialogues are excerpts from natural, unscripted conversation related to human productivity; two are machine-generated versions whose content is entirely based on their transcripts. Recordings of the original podcasts are freely available online.Footnote 1 Each podcast features a recurring host and an invited guest who is known for producing content on the episode’s topic. The overall tone and conversational style of the podcasts are conversational, relaxed, and somewhat intimate. Based on the two podcast transcripts, Google’s NotebookLM Audio Overview feature then generated the dialogue scripts with the help of Gemini AI and added conversational features, such as backchannels. The resulting scripts were voiced by two conversational agents following a podcast format using Google’s own text-to-speech capabilities. No further training data were supplied to maximize comparability between the original podcasts and machine-generated podcasts. The two human dialogue excerpts and the machine-generated podcasts were approximately ten minutes in length (M = 562.33 s, SD = 46.11). An exploratory three-speaker configuration with AI-human interaction and a 10-min podcast excerpt involving two guests and one hostFootnote 2 were later added for testing a new feature added to NotebookLM which allows for spontaneous human intervention (see Sect. 4).
2.2 Data annotation
Using Praat (Boersma and Weenink 2025), we annotated all dialogues (including those reported in Sect. 4) at the level of speaker contribution, potential role shifts and various timings. Annotations on speaker contributions and role shifts were completed independently by both authors; any deviations on speaker contribution were resolved through discussion. We only deviated on 5 out of 315 contribution labels for speakers, which corresponds to an almost perfect agreement of 98.41% (Cohen’s Kappa: κ = 0.94, 95% CI [0.92, 0.99]). Overlap contributions showed similarly high levels of agreement with 5 deviations out of 105 labels (93.3% agreement; Cohen’s Kappa: κ = 0.87, 95% CI [0.78, 0.96]). Finalized annotations and all timings measures were then exported into a CSV table, which formed the basis for later graphing and statistical analyses via SPSS and R, respectively (see Sect. 3).
Figure 1 provides a snapshot of the annotation procedure. The Praat screenshot shows the acoustic waveform (top), a spectrogram with pitch and intensity contours (mid) and text annotations (bottom) with separate tiers for speaker type (1 = guest, 2 = host), speaker role dominating the individual turn (N = novice, E = expert, C = clarification, B = backchannel), and possible role shift within turn. The speaker role reflected their knowledge whereby a novice is at a knowledge disadvantage compared to an expert. Backchannels were brief, non-intrusive overlaps that did not result in a turn change; clarifications were brief questions or comments that invited or provided an expansion of the current topic. Role labels was decided based on the assumed intent of each speaker turn, drawing on syntactic structures, prosodic cues, discourse context (whether the turn is elaborating, following up, or redirecting the topic), and the speaker’s epistemic stance (knowledgeable about the topic or not).Footnote 3 Where coding between novice and clarification contribution was ambiguous, we turned to role shifting, which was annotated on a separate tier to distinguish between instances where there was a noticeable deviation in the speakers’ assumed stance. Role shifts include overt transitions (an expert explicitly expressing uncertainty) and more implicit changes in knowledge positioning, like a novice speaker unexpectedly offering authoritative information. Minor shifts in tone or hedging were not coded as shifts from expert to novice unless they were accompanied by further indicators that the role had changed. In cases where a single turn contained a noticeable change in epistemic stance, a single role was applied based on the primary function of the turn. For instance, a host asking a question may incorporate their own experience in expressing curiosity, thereby exhibiting both novice and expert behavior; the primary function nevertheless remains being a novice. However, where these shifts impacted the course of conversation significantly, they were marked using a dedicated role shift tier, so that the speaker’s evolving stance could be captured—allowing for the overall role distributions to be decided. We also annotated speaker roles for overlapping turns on a separate tier to separate contributions from speakers and those overlapping.
For timings, our annotations also captured turn durations and transitions. Turn transitions received positive values for gaps and negative values for overlaps. Timings can be read from the lower bar or queried via the Praat menu for selections or at cursor position. Transitions were also annotated distinguishing gaps, smooth transitions, and overlaps. Gaps marked transitions where pauses exceeded 250 ms, reflecting the natural turn-taking timing grounded in Sacks et al. (1974) and later Heldner and Edlund (2010). Smooth transitions which involved minimal gaps between 100 and 250 ms. Anything below 100 ms was also annotated as smooth, but flagged in a separate notes tier, as these transitions were generally imperceptible in natural conversation, but were still worth noting for complete transparency during the annotation process. Overlaps marked transitions where a speaker began talking before the previous speaker had finished their turn, including interjections, interruptions, and backchannels.Footnote 4 Overlaps were annotated both when they marked a turn transition and when the original speaker carried on their turn; however, intra-turn pauses occurring within the same speaker’s turn were considered irrelevant for the purposes of this study, and hence not annotated.
2.3 Statistical modelling strategy and methods
Where observations were sufficiently large, our statistical modelling followed a keeping-it-maximal strategy (Barr 2013), reducing the maximal model until it converged or did not have a singular fit. Our analysis in R 4.5.1 (R Core Team 2025) used the lme4 package (Bates et al. 2015) for multimodal logistic regression modelling of numeric outcomes and categorical outcomes. Descriptives were estimated using the emmeans package (Lenth 2025); Confidence Intervals (CIs) and p values for effects were computed using a Wald z-distribution approximation.
3 Analysis
Broadly speaking, human podcasters and conversational agents sounded and behaved similarly: both podcast types exhibited regular back and forth between hosts and guests; questions served as a structure-lending device; and topics were expanded through coherent follow-ups, personal perspectives, and individual experiences. Both human and AI-generated hosts and guests sought clarification and provided conversational feedback through backchannelling. Turn transitions occurred where expected, and overlaps were mostly non-intrusive, particularly in AI-generated podcasts.
Each podcast began with the introduction of a topic by the hosts, which then gave the guest a platform to develop their perspective. This introduction was longer in human podcasts, which may be explained through the longer format of which only the first 10 min were extracted. The main section of each podcast consisted of a lively back-and-forth between the two interlocutors with an almost equal number of contributions per speaker. Because both interlocutors were informed speakers, speaker contributions changed throughout each dialogue, with a slightly more balanced distribution of contributions in the machine-generated podcasts compared to the human podcasts. The machine-generated podcasts also ended with a close by the host at the end of a segment; that close occurred outside of the excerpt of the human podcasts. Table 1 provides an overview of the number of speaker contributions and overall durations (excluding overlaps) for each speaker separated by podcast script and speaker role.
On the surface, AI-generated podcasts did not stand out as notably different from the human podcasts that provided the content for the dialogue generation. If anything, the former felt more balanced in how contributions were distributed across speakers, which grounds in speaking times rather than number of contributions. What is required, therefore, is a more nuanced investigation of what people contribute and how this reflects the trajectory of the conversation. In the following, we will discuss (1) how different types of speaker contributions were distributed across their roles and podcast types; (2) what speakers contributed during overlaps; and (3) how turn-taking timings differed across roles and types.
3.1 Types of speaker contributions
To assess role behavior of hosts and guests across the two podcast types (humans conversing with humans and AI agents conversing with other AI agents), we annotated one contribution for each turn (see Sect. 2.2 for details), distinguishing between expert contributions (E), notice contribution (N), backchannels (B), and clarification questions (C). The current discussion does not include contributions made during overlaps. A typical exchange consisted of role changes between participants, sometimes including a transition from one role to another in a longer sequence. In example (1), the host makes both novice contributions in the form of an open question, and several attempts to clarify. The guest, who was invited to speak on the topic of walking, provides expert contributions throughout.
Figure 2 provides an overview of the total number of speaker contributions for guests and hosts for each podcast type across the two topics selected for the NotebookLM generation. The figure shows a stark difference between the number of turns between humans and agents, albeit with similar proportions of turn contributions. It is worth noting that the overall count was influenced by how a turn was identified as such. Backchannels in AI rarely overlapped with speaker turns. Instead, the current speaker paused, waited for the backchannel to be uttered, and then took up their turn again. Such backchannels, therefore, resulted in two turn counts where an overlapping backchannel would only result in one turn.
Both human guests and AI guests predominantly contributed as experts (92.8% and 87.5% of their contributions, respectively) with only few novice contributions (human guest: 4.7%; AI guests: 6.3%) and even fewer clarification contributions (human guest: 2.4%; AI guests: 6.4%). Human and AI hosts, on the other hand, had more novice contributions than guests (29.4% and 13.6% of their contributions, respectively). Compared to their guest counterparts, clarification contributions were only higher for human hosts (11.8%) not AI hosts (2.3%), but expert contributions continued to make up for the majority of human (84.1%) and AI (58.8%) hosts.
While there was a (non-significant) trend of different role distributions between AI and human speakers when speakers were collapsed (X2(2) = 4.8908, p = 0.08669), we also ran an exploratory statistical analysis with a multimodal logistic regression for the interaction between the two variables. Although the model converged, coefficient estimates were unstable and associated with large standard errors, which renders it unreliable. We there refrain from interpreting this comparison. To address the strong asymmetry of observations for human podcasts compared to the NotebookLM-generated podcasts, which may have made the multimodal model unstable, we ran a second model using multinomial logistic regression to compare expert contributions against other contributions (novice and clarification). For this, we fitted a logistic mixed model to predict the turn contribution with podcast type and speaker role as fixed effects. The model also included dialogue transcripts as a random effect. Contrasts were sum-coded. The model's total explanatory power is moderate (conditional R2 = 0.22), and the part related to the fixed effects alone (marginal R2) is of 0.14. Within this model, we observed an effect of the speaker role whereby guests had more expert contributions than hosts (β = 1.29, 95% CI [0.14, 2.44], p = 0.028), in line with the visual trends in Fig. 2.
3.2 Overlap contributions
For overlaps, we also wanted to know whether speaker contributions differed across podcast types, now with the added contribution of backchannels. Number of overlap contributions across the two transcripts are visualized in Fig. 3. Even at first glance, we can observe that overlaps were more frequent, and backchannels played a greater role in human dialogues than in AI-generated dialogues. The former is worth noting because the smaller number of turns in human dialogues (see Fig. 2) provided fewer opportunities for overlapping, while variation in turn length should not affect the opportunities for backchanneling. Among their overlap contributions, human (87.5%) and AI-generated (90%) guests mostly used backchannels, notably more so than human (57.1%) and AI-generated hosts (25%). Clarifications made up only a small portion of overlaps among AI-generated (10%) and human (6.3%) guests. AI-generated hosts did not clarify, but human hosts did so once (6.7%). Expert contributions accounted for 50% of the AI-generated host overlaps, compared to 14.3% of the contribution of human hosts. Neither human nor AI-generated guests used clarifications at all during overlaps. Only human guests made novice contributions during overlaps (6.3%), AI-generated hosts used 25% of their overlaps on novice contributions, compared to 21.4% of human host overlaps.
Given the low number and large asymmetry of overlap contributions (see Fig. 3), we did not run any statistical analyses here. Visual trends suggest that asymmetries in the overlapping behavior of guest and host were fairly similar between the two podcast types.
3.3 Timings
We were also interested in comparing the turn timings of human and AI-generated dialogues to see the extent to which they adhere to a strong version of the ‘no-gap-no-overlap’ strategy. Here we present the results for the durations for average turns, overlaps, and transitions. Intuitively, we would expect NotebookLM to be ‘better’ at turn taking, i.e., with fewer overlaps and shorter transitions, because the full transcript was available at the point of podcast generation. Endpointing was self-determined for machines. Humans, on the other hand, had to await or anticipate each turn end individually.
3.3.1 Duration of turns
Figure 4 provides an overview of the distributions for total turn durations for speakers and hosts across the two original podcasts and the two NotebookLM simulations. Average human turn durations were higher for both human hosts (M = 22.3 s, SE = 4.14, 95% CI [14.15, 30.5]) and guests (M = 49.4 s, SE = 4.27 s, 95% CI [40.99.11, 57.9]) than for AI-generated hosts (M = 11.3 s, SE = 2.57 s, 95% CI [6.18, 16.4]) and guests (M = 14.8 s, SE = 2.63 s, 95% CI [9.63, 20.1]), respectively. Despite the stark differences between humans and AI dyads, turn durations for guests were consistently higher than those for hosts. It is also worth noting that human distributions were skewed toward longer turns, while AI-generated turn durations were much more symmetrically distributed.
For the statistical analysis, we fitted a linear mixed model to predict the average turn durations with speaker role and podcast type as fixed effects. The model included transcripts as a random effect. Contrasts were treatment-coded with the intercept corresponding to the average overlap duration of an AI-generated guest. The model's total explanatory power was substantial (conditional R2 = 0.58), and the part related to the fixed effects alone (marginal R2) was of 0.46. Within this model, we found a statistically significant difference for podcast type (β = 39.11, 95% CI [23.76, 54.47], t(113) = 45.05, p < 0.001), marking an increase from AI-generated turn durations to human turn durations, no statistically significant difference for speaker role (β = − 1.20, 95% CI [− 6.78, 4.38], t(113) = 4.62, p = − 0.06), and a statistically significant interaction of the two fixed effects (β = − 25.51, 95% CI [− 36.11, − 14.90], t(113) = − 4.77, p < 0.001) suggesting that speaker role and podcast type combined in changing the turn durations compared to AI guests.
3.3.2 Overlaps
We also compared human and AI-generated overlapping behavior. A representative overlap example from both a human-to-human dialogue is provided in (2). Overlaps, such as these, seem to signal agreement or emotional engagement (as in the earlier example in (1)). AI-to-AI dialogues had hardly any overlaps. Backchannels, like the one in (3), resembled individual turns with smooth transitions, despite their non-intrusive function. The few, very brief overlaps observed in AI-generated podcasts were all backchannels, notably not occurring at turn-transition moments, but occurring mid-turn.
Human dialogues had a lower number of turns (33) and transitions (25) than AI-generated dialogues, which aligns with its long average turn duration (Fig. 5). Of these transitions, 18 were smooth and 7 were gap transitions. The AI-generated dialogues featured almost double the number of turns compared to the human dialogues (86). Smooth transitions were more common (43), but the dialogues still exhibited 39 instances of gap transitions. The distribution of gap transitions to smooth transitions here suggests that NotebookLM does not fully mirror the rhythm of human-to-human turn-taking. Figure 5 shows the distribution of overlap durations (in seconds) of guests and hosts compared across podcast types with now familiar differences between human and AI-generated behavior. While both AI-generated hosts (M = 60 ms, SE = 219 ms, 95% CI [32.11, 66.76]) and guests (M = 63 ms, SE = 219, 95% CI [− 380, 506]) had very short overlaps, human hosts overlapped quite notably (M = 2109 ms, SE = 352, 95% CI [1413, 2806]) and human guests even more so (M = 4216 ms, SE = 362 ms, 95% CI [3498, 4934]).
For analyzing the difference between AI and human overlapping behavior across speaker roles, the few observations meant that the random effect of transcript had to be dropped. We, therefore, fitted a linear regression model to predict overlap duration with podcast type and speaker role. The model explains a statistically significant and substantial proportion of variance (R2 = 0.52, F(3, 115) = 41.43, p < 0.001, adj. R2 = 0.51). Factors were treatment-coded with the model’s intercept corresponding to an AI-generated guest. Within this model, humans had a significantly longer duration than agents (β = 4153.44, 95% CI [3309.74, 4997.13], t(115) = 9.75, p < 0.001). There was also a significant interaction between the effects of podcast type and speaker role reducing the duration of overlaps across speakers and podcast types (β = − 2104.37, 95% CI [− 3280.97, − 927.76], t(115) = − 3.54, p < 0.001).
3.3.3 Transitions
Finally, we also compared transition times between human speakers and AI-generated speakers. Figure 6 presents the distribution of transition times, with AI-generated speakers taking slightly longer to transition (M = 272 ms, SE = 18.7, 95% CI [235, 309]) than human speakers (M = 231 ms, SE = 33.2, 95% CI [165, 297]). These transitions included gaps and smooth transitions whereby human podcasts included fewer gaps (7) and smooth transitions (19) than AI-generated podcasts (39 and 43, respectively). This asymmetry between the two podcast types follows from the more frequent turn changes and shorter turn durations for AI-generated podcasts described above.
For the analysis of transition times, we fitted a linear mixed model to predict transition times with podcast type. The model also included transition type as random effect. The difference between human dyads and AI agents was not statistically significant. The model’s explanatory power related to the fixed effects alone (marginal R2) was 0.01. Within this model, the effect of podcast type is statistically non-significant and negative (β = − 0.04, 95% CI [− 0.12, 0.03], t(104) = − 1.08, p = 0.281). Visual trends in Fig. 6 nevertheless suggest that human transitions showed less variance than AI transitions.
4 Comparison with hybrid version
In December 2024, Google enhanced NotebookLM further by giving the user the ability to join a conversation between its two AI-generated podcast speakers (Google 2024b). This intervention feature significantly adds to the complexity of generating natural dialogue because content and transitions are no longer fully plannable. It, therefore, serves as a test case for how linguistic measures of naturalness in generated dialogue depend on pre-planning. With the integration of this “interactive mode”, an unknown contributor, whose role and knowledge state is not clearly defined, can enter the conversation at any point, and their contribution must be integrated into the pre-planned course of the conversation. Given Notebook LM’s approach of conservative transition estimation, it can be expected that turn transitions notably slow down and overlaps are largely avoided in the context of human intervention.
Figures 7 and 8 show that this prediction pans out across the timing measures. Transition timings increase while overlaps remain short. A three-way conversation between humans is provided as a reference for both transitions and overlaps. The three-way conversation between humans was one between a host and two guests discussing a movie; the NotebookLM conversation with human involvement was a conversation between one AI guest and one host based on the transcription of the latter, with a human calling in. Just as with a two-way conversation between humans, transitions are short with little variance, and all three interlocutors show some overlapping (with the host showing the smallest overlapping durations, similarly to the two-way conversations).
Outliers in Fig. 8 mark those instances where the human intervened. Any intervention is (somewhat unwieldily) integrated into the narrative of the pre-planned dialogue. Overlaps are completely absent when a human intervenes as the dialogue generation is interrupted. We also found audio-glitches causing sudden switches in speaker voice (impacting role allocation and stability), an inability of the user to interrupt or overlap due to the “call-in” structure (preventing the spontaneous co-construction of dialogue between human and AI), and a prolonged, silence-based end-of-turn detection (meaning that turn-taking cannot be initiated in a natural way and is delayed).
While this comparison can only serve as a snapshot of human intervention in AI-generated dialogues, it suggests that the new feature prevents NotebookLM from maintaining the fluidity observed for the dialogues generated without human intervention. A specific example is provided in (4). The long breaks before and particularly after the human intervention signal a clear interruption of the scripted dialogue. This interruption receives an instant reaction as the breaking off in the hosts ongoing contribution suggests. After a long delay, the host acknowledges the human’s intervention. After a few more turns, the agents return to the discussion of the pre-planned script (not shown in the excerpt in (4)).
We leave it at this surface-level comparison because the overall picture is clear: intervening in the pre-planned dialogue interrupts what worked so well for NotebookLM’s generation of a classic host–guest podcast dialogue. The interruption results in a dialogue that lacks the smoothness of transition previously recorded and introduces distortions and inaccuracies which we can ascribe to the fact that low-resource pre-planning is no longer an option and requests resource-intensive, spontaneous updating.
5 Discussion
The aim of this paper was to provide an exploratory investigation of how pre-planned, AI-generated dialogues compare to spontaneous human conversations. Google’s NotebookLM Audio Overview featured as a case study to investigate two central dimensions of natural dialogue: turn-taking timings (Arora et al. 2025; Sacks et al. 1974) and knowledge asymmetry (Sidnell 2012). Our comparisons showed that role behavior was similar between human and machine-generated dialogues with a significant effect of expert contributions occurring more frequently for guests, just as we would expect from the social scripts of the podcast genre: an invited guest usually is a source of expertise that makes them attractive as a guest in the first place. Number of contributions and turn-taking durations strongly differed between two podcast types despite their similar length overall. Human podcast excerpts had longer durations and fewer contributions per speaker than those in AI-generated podcasts of a similar length. This may have been the result of the known overall duration of the podcasts chosen where speakers knew that they had more time to develop their thoughts. The more important point, however, is that fewer speaker changes did not result in fewer or shorter overlaps. Overlaps of various kinds were more frequent in human podcasts, likely signaling a high degree of engagement from the interlocutor. Somewhat surprisingly, human guests, not hosts, had the longest overlap durations. This may be a reflection of the lower information density of host contributions, which aligns with the conventional role of the host to provide a platform for the guest to develop their input. Previous research has shown that overlaps increase with lower information density (Dethlefs et al. 2016). Backchanneling was a common feature across speakers and podcast types, but they showed differences in their overlapping behavior. Duration of transitions was well within the range observed in the previous literature (at around 200 ms). Finally, any intervention of a human speaker in AI-generated podcasts (a novel feature which we added for comparison only) revealed that NotebookLM shows performance issues when integrating human interventions; it reverts to conservative endpointing strategies and temporarily interrupts the pre-planned script with a short acknowledgement only to return to the original script notably quickly.
These findings suggest that despite the surface-level realism of the conversations that NotebookLM can generate, its generated podcasts underperform in several areas investigated in this study, albeit more so on qualitative than on quantitative levels. The agents seemed to follow a conservative interpretation of the no-gap-no-overlap strategy and kept turn duration consistently around what others have reported in the previous literature. Guest and host behavior also showed a comparable number of role shifts. Yet, where human exchanges show speaker roles and turn-taking timings as being highly interdependent on previous turns, the AI exchanges show a much smaller degree of natural variation. Knowledge profiles of AI-generated guests and hosts were very similar, and overlapping was significantly rarer than in humans despite additional overlapping opportunities based on a higher number of turns. Put differently, AI agents not only spoke shorter and exhibited fewer overlaps than their human counterparts; they also exhibited more similarity in speaker-specific behavior, which goes against the expectations set by the podcast genre of longer exchanges between guest and host as well as different role behavior. In the case of AI-generated dialogue with human intervention, where dialogue generation has to cope with spontaneous speech, it becomes clear that mechanical limitations inhibit both turn-taking timings and surface-level metrics, such as role stability, voice consistency and coherence. These mechanical limitations negatively impacted NotebookLM’s ability to reflect and recreate the shifting understanding and dynamic interaction patterns needed for natural and coherent conversational flow. The boundary between simulation of dialogue and integration of human contributions is defined by technical and resource limitations that are otherwise absent in NotebookLM, presumably because the spontaneous interventions require resource-intensive context integration and the prepared scripts need revising.
The data from our explorations supports the argument that turn-taking timing is generally smooth in human dialogue (Stivers et al. 2009; Gratier et al. 2015), with all human dialogues examined having very few transitions over 250 ms. This is where NotebookLM podcasts and human podcasts clearly align. Across all six dialogues investigated in the present study, the mean length of gaps and overlaps corresponded with Heldner and Edlund’s (2010) suggestion that gaps and overlaps can sometimes exceed 200 ms. Despite some concerns in the previous literature that transition averages lack nuance, the 200–250 ms window is nevertheless a fair description for the transition in human podcasts and the machine-generated podcasts investigated. Yet, while both human and machine-generated dialogues featured smooth turn transitions, only the human dialogues consistently demonstrated a high ratio of smooth to gap transitions. This pattern suggests that favoring smooth transitions is not a learned behavior or a surface-level feature of interaction, but is a core competency—which is in line with previous research suggesting that turn-taking management is intertwined with the cognitive and social development of humans (Reddy et al. 1997; Dominguez et al. 2016; Levinson and Torreira 2015).
NotebookLM dialogues, despite their smooth transitioning, contained proportionally more gap transitions, highlighting the limits of mechanical imitation. As stated by Ward and DeVault (2017), there is only so much an AI can replicate naturally without being able to genuinely interact with a human being. While smooth transitions dominate the human conversations, they are not enforced, and gap transitions do occur. However, they are rare and have contextual justification (unlike in AI dialogues), reminiscent of Wilson et al.’s (1984) idea that turn size emerges dynamically depending on what is needed to maintain conversational coherence. The current findings are not without precedent. Previous research on spontaneous AI systems show that they do not have this natural inclination towards fluent timings, and do not possess predictive modelling of turn endings that are perceived as natural (Skantze 2021), a pattern that resurfaces in dialogues that mixed AI and human contributions. Overall, the turn-taking timing data in this study suggests that where humans predictively process turns (Garrod and Pickering 2015), AI generates transitions mechanically—leading to an almost perfectly synchronized conversational rhythm with only few overlaps. NotebookLM’s conversational dynamics appear rigid due to its strict adherence to the no-gap-no-overlap strategy and limited variation in turn duration, both within podcast type and across roles (guests vs. hosts). While both podcast types display transition times that match expectations from the literature, it is the human capacity to manage and adapt to any shifts to preserve conversational naturalness, something that may be beyond pre-planned dialogue generation.
Though the role shifts for human and machine-generated dialogues mirror each other, only the instances of epistemic flexibility demonstrated in the human dialogues were functionally responsive to the context, supporting Stalnaker’s (1978) view that maintaining common ground depends on adjusting assumptions and knowledge displays, and Clark and Brennan’s (1991) model of conversational grounding, where mutual understanding is sustained through subtle, ongoing coordination between participants. In the NotebookLM dialogues, knowledge role flexibility was too rigid and mechanically inconsistent, maintaining static expert stances regardless of context or producing abrupt, short-lived, scripted shifts that lack interactional relevance. NotebookLM’s failed to simulate contextually appropriate role shifts, which in turn disrupts the overall dynamic adjustment necessary to maintain common ground, leading again to unnatural conversational rhythms and a breakdown of cooperative dialogue. Much of the generated podcast dialogues sounded like a discussion between two equally informed participants with overlapping areas of expertise, rather than a host who wanted to give a platform to the guest to share their expertise. While this may be in line with the tool’s primary delivery goal (the podcast feature is a later addition to a text-based research assistant), the greater balance in expertise undermines the naturalness of the generated podcasts through breaking with genre convention. Adding human intervention to AI-generated dialogues further increased the sense of a lack of flexibility. The absence of meaningful role flexibility meant that conversational expectations grounded in reducing knowledge asymmetry were repeatedly violated at the point of intervention, which destabilized the flow of information and made exchanges feel disjointed.
Taken together, the results of this study invite reflection on both the limitation of state-of-the-art in AI podcast generation and the nature of human dialogue itself. NotebookLM’s failure to truly replicate and adopt the natural flow and conversational dynamics of human dialogue highlights that the purpose of human communication is more than just the transmission of information in a rigid pattern dedicated to reducing any knowledge asymmetry between interlocutors. Human conversation is an emergent and collaborative process shaped by mutual attention (Sidnell 2012), verbal feedback (Clark 1996), social alignment (Clark and Brennan 1991), and shared and growing context (Stalnaker 1978, 2014). NotebookLM attempts to increase naturalness in dialogue through adding mostly appropriately timed backchannels and avoiding overlaps, but these strategies are applied more strictly than humans would. The human ability to time responses fluidly and adjust epistemic stance dynamically to construct and maintain common ground points towards conversation being a socially situated activity shaped by subtle cues and embodied awareness. The podcast genre serves as a reminder that such activity can be informed by social scripts that lend additional structure to a conversation. Backchanneling and even “intrusive” overlapping can add to a sense of engagement and overall coherence rather than introduce disruptions. With that in mind, the over- and imperfections observed through AI systems, such as Google’s NotebookLM, offer an indirect insight into what exactly humans do when they converse: Human conversation is not optimized for seamless information exchange, but for building common ground.
6 Conclusion
The findings of this study suggest that NotebookLM can simulate individual features of human conversation patterns dialogues through pre-planning, but lacks the ability to coordinate these features responsively in real time. Human dialogue is defined by the fact that turn-taking timing and knowledge role allocation are not mutually exclusive, but rather work in tandem. Contrastingly, NotebookLM’s dialogues seem to separate these key ingredients of natural conversation, which can result in mechanical and rigid exchanges that lack the rhythm and adaptability seen in human dialogues. This study was limited by its small dataset, reliance on a single platform (Google’s NotebookLM) for assessing the abilities of AI in creating human-like data, and the use of podcast recordings and their scripts rather than everyday conversations as a baseline for the human dialogue. The potential for extrapolating our findings to spontaneous dialogue generation is, therefore, limited to the descriptive comparison of AI-generated dialogues that follow NotebookLM’s pre-planning strategy within the podcast genre. The inclusion of the joining feature shows that the inflexibility of role adapting, and the mechanical execution of turn-taking behavior clearly have their limits in generating natural dialogue; every human intervention led to a conversational breakdown where attempts of integrating human interventions remained incoherent and disfluent. Future research will need to investigate more spontaneous interactions, and different types of dialogue systems. The main finding remains that AI-generated dialogue can mimic features of natural dialogue in isolation, but fails to synthesize these features into a dynamic tapestry of negotiating information and relation. In other words, even though AI has made impressive progress in delivering the core ingredients to natural conversation, it only excels in mimicking its surface elements. Turn-taking behavior in real life is not governed by rigid adherence to a particular role, perfect transition prediction, or a strict avoidance of gaps and overlaps. Achieving genuine naturalness will require AI systems that can interpret social contexts and scripts and respond dynamically to the fluid negotiation of interlocutor relations in conversation. To converse naturally, it is not just about what you say without interrupting the other person, but also about how you respond to the developing dynamics of a conversation within the context in which that conversation happens.
Data availability
All recordings supporting the findings of this study are available within the paper or its Supplementary Information.
Notes
The two podcasts are sperate episodes from Finding your causal Magic (with UnJaded Jade): Episode from Dec 22, 2024: “How You're Hustling Is All Wrong (with Sahil Bloom)”, available here: https://www.youtube.com/watch?v=gMhbcLTufMI and Episode from Oct 20, 2024: “Why Redefining Productivity Works (with Ruby Granger)”, available here: https://www.youtube.com/watch?v=fELbEbnD5p8
Source: “‘What’s our band? Stray Old Men?’ Ryan Reynolds & Hugh Jackman on Deadpool & Wolverine”, MTV (with Josh Horowitz as host); available here: https://www.youtube.com/watch?v=L6y8viGzWhI
An anonymous reviewer questioned the validity of these annotations given the indirect access to functional intent through the cues listed here. We note that all data were double coded by the two authors, who are well-trained in analysing knowledge asymmetry. The annotations showed an agreement of 93%; the remaining discrepancies resolved through discussion. The combination of contextual information and morphosyntactic cues are highly informative for identifying functional intent.
An anonymous reviewer pointed to the fact that backchannels are usually considered an ‘unproblematic overlap’ (see Schegloff 2000 for a more nuanced distinction of overlaps) since the addressee does not claim the turn but signals that the speaker should continue. While we fully agree with this interpretation of backchannels, our annotations were purely descriptive with respect to the rhythm of turn taking, not the possible functions over various overlaps. We therefore continue considering backchannels as overlaps, even though they are indeed ‘unproblematic’ or co-operative just as other forms of overlapping can be (Murata 1994; Samrose et al 2018).
References
Arora S, Lu Z, Chiu CC, Pang R, Watanabe S (2025) Talking turns: benchmarking audio foundation models on turn-taking dynamics. arXiv:2503.01174. https://doi.org/10.48550/arXiv.2503.01174
Auer P (2005) Projectionin interaction and projection in grammar. Text 25: 7–36.
Barr DJ, Levy R, Scheepers C, Tily, HJ (2013) Random eff ects structure for confirmatory hypothesis testing: Keep it maximal. Journal of memory and language 68(3): 255-278.
Bates D, Mächler M, Bolker B, Walker S (2015) Fitting linear mixed-effects models using lme4. J Stat Softw 67(1):1–48. https://doi.org/10.18637/jss.v067.i01
Bavelas JB, Jong PD, Korman H, Jordan SS. (2012). Beyond back-channels: A three-step model of grounding in face-to-face dialogue. In Proc. FBID 2012: 5-6.
Bender EM, Gebru T, McMillan-Major A, Shmitchell S (2021) On the dangers of stochastic parrots: can language models be too big? In: Proceedings of the 2021 ACM conference on fairness, accountability, and transparency (FAccT ’21), pp 610–623
Berg FSA (2025) Analysing podcast intimacy: four parameters. Convergence 31(4):1423–1438
Boersma P, Weenink D (2025) Praat: doing phonetics by computer. http://www.praat.org Accessed 9 Sep 2025.
Bögels S, Torreira F (2015) Listeners use intonational phrase boundaries to project turn ends in spoken interaction. J Phonetics 52:46–57
Brabra H, Báez M, Benatallah B, Gaaloul W, Bouguelia S, Zamanirad S (2022) Dialogue management in conversational systems: a review of approaches, challenges, and opportunities. IEEE Trans Cogn Dev Syst 14(3):783–794
Calhoun S, Carletta J, Brenier JM, Mayo N, Jurafsky D, Steedman M, Beaver D (2010) The NXT-format Switchboard Corpus: a rich resource for investigating the syntax, semantics, pragmatics and prosody of dialogue. Language resources and evaluation 44(4): 387-419.
Casillas M, Frank MC (2017) The development of children’s ability to track and predict turn structure in conversation. Journal of memory and language 92: 234-253.
Cassell J, Nakano YI, Bickmore TW, Sidner CL, Rich C (2001) Non-verbal cues for discourse structure. In: Proceedings of the 39th annual meeting of the Association for Computational Linguistics, pp 114–123. https://www.aclweb.org/anthology/P01-1016.pdf
Chang S, Li B, Sainath TN, Zhang C, Strohman T, Liang Q, He Y (2022) Turn-taking prediction for natural conversational speech. arXiv:2208.13321. https://doi.org/10.48550/arXiv.2208.13321
Clark HH, Brennan SE (1991) Grounding in communication. In: Resnick LB, Levine JM, Teasley SD (eds) Perspectives on socially shared cognition. American Psychological Association, Washington DC, pp 127–149
Clark HH (1996) Using language. Cambridge University Press.
Clark HH, Wilkes-Gibbs D (1986) Referring as a collaborative process. Cognition 22(1):1–39
Clayman SE (2012) Turn-constructional units and the transition-relevance place. The handbook of conversation analysis. Malden, Wiley-Blackwell, pp 151–166
Darapaneni N, Paduri AR, Tank U, KanthaSamy B, Ranjan A, Krisnakumar R (2023) Conversational AI: a study on capabilities and limitations of dialogue-based systems. In: Morusupalli R et al (eds) Multi-disciplinary trends in artificial intelligence (MIWAI 2023), vol 14078. LNCS. Springer, pp 503–512
Dethlefs N, Hastie H, Cuayáhuitl H, Yu Y, Rieser V, Lemon O (2016) Information density and overlap in spoken dialogue. Comput Speech Lang 37:82–97. https://doi.org/10.1016/j.csl.2015.11.001
Dominguez S, Devouche E, Apter G, Gratier M (2016) The roots of turn-taking in the neonatal period. Infant Child Dev 25(2):101–115
Drew P, & Sorjonen ML (1997) Institutional dialogue. Discourse as social interaction 2: 92-118.
Duncan S (1972) Some signals and rules for taking speaking turns in conversations. J Pers Soc Psychol 23(2):283–292
Ehret K, Bosman L, Babayode A, Chan N, Fong I, Harris N, Taboada M (2024) Podcasts as an emerging register of computer-mediated communication. Reg Stud 6(2):128–174
Farkas D (2024) Canonical and non-canonical questions in discourse. In: Eckardt R, Walkden G, Dehé N (eds) The Oxford handbook of non-canonical questions. OUP, Oxford, pp 1–25
Farkas D, Bruce KB (2010) On reacting to assertions and polar questions. J Semant 27(1):81–118
Garrod S, Pickering MJ (2015) The use of content and timing to predict turn transitions. Front Psychol 6:751
Google (2023) NotebookLM: an AI-first notebook from Google DeepMind. Google Blog. Accessed 3 Mar 2025.
Google (2024a) NotebookLM’s new audio overviews help you quickly understand documents. Google Blog. Accessed 3 Mar 2025.
Google (2024b) New features in NotebookLM: December 2024 update. Google Blog. Accessed 3 Mar 2025.
Gratier M, Devouche E, Guellai B, Infanti R, Yilmaz E, Parlato-Oliveira E (2015) Early development of turn-taking in vocal interaction between mothers and infants. Front Psychol 6:1167
Grice HP (1957) Meaning. The philosophical review 66(3): 377-388.
Grice HP (1975) Logic and conversation. In: Cole P, Morgan JL (eds) Speech acts. Brill, Leiden, pp 41–58
Heim JM (2019) Commitment and engagement: the role of intonation in deriving speech acts. PhD dissertation, University of British Columbia
Heim JM (2025) Quantifying division of labour: effects of clause type on intonational meaning. Speech Commun 172:103265. https://doi.org/10.1016/j.specom.2025.103265
Heim JM, Wiltschko ME (2020) Deconstructing questions: reanalysing a heterogeneous class of speech acts via commitment and engagement. Scand Stud Lang 11(1):56–82
Heim JM, Keupdjio H, Lam ZW, Osa-Gómez A, Thoma S, Wiltschko ME (2016) Intonation and particles as speech act modifiers: a syntactic analysis. Stud Chin Linguist 37(2):109–129
Heim JM, Marí JR, Wiltschko M (2025) Three stages in the acquisition of English response tokens: a window into the development of common ground. Cah Praxém. https://doi.org/10.4000/15b3u
Heiselberg L, Have I (2023) Host qualities: conceptualising listeners’ expectations for podcast hosts. J Stud 24(5):631–649. https://doi.org/10.1080/1461670X.2023.2178245
Heldner M, Edlund J (2010) Pauses, gaps, and overlaps in conversations. J Phonetics 38(4):555–568
Heritage J (2012) Epistemics in action: action formation and territories of knowledge. Res Lang Soc Interact 45(1):1–29
Hjalmarsson A (2011) The additive effect of turn-taking cues in human and synthetic voice. Speech Commun 53(1):23–35
Hoogland D (2025) Conversational turn timing: the effects of prosody and pragmatic context in production and perception. PhD thesis, Newcastle University
Hoogland D, White L, Knight S (2023) Speech rate and turn-transition pause duration in Dutch and English spontaneous question–answer sequences. Languages 8(2):115
Hung V, Elvir M, Gonzalez A, DeMara R (2009) Towards a method for evaluating naturalness in conversational dialog systems. In: Proceedings of the IEEE international conference on systems, man, and cybernetics, pp 1236–1240
Jacobsen UC (2014) Knowledge asymmetry in action. Hermes J Lang Commun Bus 53:57–74
Jarrett K (2009) Private talk in the public sphere: podcasting as broadcast talk. Commun Polit Cult 42(2):116–135
Ji Z et al (2023) Survey of hallucination in natural language generation. ACM Comput Surv 55(12):1–38
Jokinen K (2010) Non-verbal signals for turn-taking and feedback. In: LREC
Kangulu S (2025) The role of podcasts in shaping public discourse. Res Invent J Res Educ 5(1):60–66. https://doi.org/10.59298/RIJRE/2025/516066
Kendrick KH, Holler J, Levinson SC (2023) Turn-taking in human face-to-face interaction is multimodal: gaze direction and manual gestures aid the coordination of turn transitions. Philos Trans R Soc B Biol Sci 378(1875):20210473
Kim Y (2019) What is story-structure type?: knowledge asymmetry, intersubjectivity, and learning opportunities in conversation-for-learning. Appl Linguist 40(2):307–328
Kitzinger C, Mandelbaum J (2013) Word selection and social identities in talk-in-interaction. Commun Monogr 80:176–198
Kokala A (2024) Revolutionising content creation: leveraging AI-driven podcast generation with NotebookLM and personalized insights. Int Res J Mod Eng Technol Sci 6:1798–1803
Kolekar SS, Richter DJ, Bappi MI, Kim K. (2024). Advancing AI voice synthesis: Integrating emotional expression in multi-speaker voice generation. Proceedings of the International Conference on Artificial Intelligence in Information and Communication (ICAIIC)
Kulkarni P, Mahabaleshwarkar A, Kulkarni MH, Sirsikar N, Gadgil KD (2019) Conversational AI: an overview of methodologies, applications and future scope. In: Proceedings of the 5th international conference on computing, communication, control and automation (ICCUBEA), pp 1–7
Lala D, Inoue K, Kawahara T (2018) Evaluation of real-time deep learning turn-taking models for multiple dialogue scenarios. In: Proceedings of the 2018 international conference on multimodal interaction (ICMI’18), pp 76–83
Le Bigot L, Terrier P, Amiel V, Poulain G, Jamet E, Rouet JF (2007) Effect of modality on collaboration with a dialogue system. Int J Hum Comput Stud 65(11):983–991
Legerstee M, Corter C, Kienapple K (1990) Hand, arm, and facial actions of young infants to a social and non-social stimulus. Child Dev 61(3):774–784
Lenth R (2025) emmeans: estimated marginal means, aka least-squares means. R package version 1.11.2–8. https://doi.org/10.32614/CRAN.package.emmeans
Levinson SC (2016) Turn-taking in human communication: origins and implications for language processing. Trends Cogn Sci 20(1):6–14
Levinson SC, Torreira F (2015) Timing in turn-taking and its implications for processing models of language. Front Psychol 6:731
Lewis DK (1979) Scorekeeping in a language game. J Philos Logic 8:339–359
Liu Y (2024) Human-like natural language generation in spoken dialogue systems. Doctoral dissertation, Ulm University
Magyari L, De Ruiter JP (2008) Timing in conversation: the anticipation of turn endings. In: Proceedings of the 12th workshop on the semantics and pragmatics of dialogue, pp 107–114
Mahlasela O, Baloyi E, Siphambili N, Khan ZC (2024) Artificial Intelligence Impact on the realism and prevalence of deepfakes. Available online: https://www.researchgate.net/profile/Errol-Baloyi-2/publication/384056978_Artificial_Intelligence_Impact_on_the_realism_and_prevalence_of_deepfakes/links/66e7ec96dde50b3258772247/Artificial-Intelligence-Impact-on-the-realism-and-prevalence-of-deepfakes.pdf Accessed 14 Apr 2025.
Malamud SA, Stephenson T (2015) Three ways to avoid commitments: declarative force modifiers in the conversational scoreboard. J Semant 32(2):275–311
McTear M (2020) Conversational AI: dialogue systems, conversational agents, and chatbots. Springer
Meyer AS (2023) Timing in conversation. J Cogn 6(1):20
Mitsui K, Nakamura A, Takayama K (2023) Towards a method for evaluating naturalness in conversational dialogue systems. In: Proceedings of the annual meeting of the association for computational linguistics, pp 1121–1133
Murata K (1994) Intrusive or co-operative? A cross-cultural study of interruption. J Pragmat 21(4):385–400
Murray SE (2014) Varieties of update. Semant Pragmat 7(2):1–53
Murray SE, Starr WB (2021) The structure of communicative acts. Linguistics and Philosophy 44(2): 425-474.
Murray SE, Starr WB (2021) The structure of communicative acts. Linguist Philos 44(2):425–474
Oller DK (2000) The emergence of the speech capacity. Psychology Press
Osa Gomez del Campo A (2020) Epistemic (mis)alignment in discourse: what Spanish discourse markers reveal. Doctoral dissertation, University of British Columbia
R Core Team (2025) R: a language and environment for statistical computing. R Foundation for Statistical Computing, Vienna
Reddy V, Hay D, Murray L, Trevarthen C (1997) Communication in infancy: mutual regulation of affect and attention. In: Bremner G, Slater A, Butterworth G (eds) Infant development. Psychology Press, pp 247–274
Riest C, Jorschick AB, de Ruiter JP (2015) Anticipation in turn-taking: mechanisms and information sources. Frontiers in psychology 6:89.
Roberts C (2012) Information structure: towards an integrated formal theory of pragmatics. Semant Pragmat 5:1–69. https://doi.org/10.3765/sp.5.6
Roberts F, Margutti P, Takano S (2011) Judgments concerning the valence of inter-turn silence across speakers of American English, Italian, and Japanese. Discourse Processes 48(5): 331–354. https://doi.org/10.1080/0163853X.2011.558002
Sacks H, Schegloff EA, Jefferson G (1974) A simplest systematics for the organization of turn-taking for conversation. Language 50(4):696–735
Samrose S, Zhao R, White J, Li V, Nova L, Lu Y, Ali MR, Hoque ME (2018) CoCo: collaboration coach for understanding team dynamics during video conferencing. Proc ACM Interact Mob Wearable Ubiquitous Technol. https://doi.org/10.1145/3161186
Schaefer MY, Lowe RJ (2024) Critical co-presentership: podcasting as reflective practice. In: Verla Uchida A, Roloff Rothman J (eds) Cultivating critical friendships through reflective practice. Candlin & Mynard
Schegloff EA (2000) Overlapping talk and the organization of turn-taking for conversation. Lang Soc 29(1):1–63. https://doi.org/10.1017/S0047404500001019
Searle JR (1975) Indirect speech acts. Syntax and semantics 3: 59-82.
Searle JR (1980) Minds, brains, and programs. Behav Brain Sci 3(3):417–424
Sidnell J (2012) Who knows best? Evidentiality and epistemic asymmetry in conversation. Pragmat Soc 3(2):294–320
Skantze G (2021) Turn-taking in conversational systems and human–robot interaction: a review. Comput Speech Lang 67:101178
Skantze G, Hjalmarsson A (2013) Towards incremental speech generation in conversational systems. Comput Speech Lang 27(1):243–262
Stalnaker R (1978) Assertion. In: Cole P (ed) Syntax and semantics, vol 9. Pragmatics. Academic Press, New York, pp 315–332
Stalnaker R (2014) Context. Oxford University Press, Oxford
Stivers T et al (2009) Universals and cultural variation in turn-taking in conversation. Proc Natl Acad Sci USA 106(26):10587–10592
Ten Bosch L, Oostdijk N, Boves L (2005) On temporal aspects of turn taking in conversational dialogues. Speech Commun 47(1–2):80–86
Vaswani A et al (2017) Attention is all you need. In: Advances in neural information processing systems
Ward NG, DeVault D (2017) Challenges in building highly-interactive dialog systems. AI Mag 37(4):7–18
Westera M (2017) Exhaustivity and intonation: a unified theory. Doctoral dissertation, University of Amsterdam
Wilks Y, Catizone R, Worgan S, Turunen M (2011) Some background on dialogue management and conversational speech for dialogue systems. Comput Speech Lang 25(1):128–139
Wilson TP, Wiemann JM, Zimmerman DH (1984) Models of turn-taking in conversational interaction. J Lang Soc Psychol 3(3):159–183
Wiltschko M (2021) The grammar of interactional language. Cambridge University Press, Cambridge
Xiao Y, He L, Guo H, Xie FL, Lee, T (2025, July) PodAgent: A comprehensive framework for podcast generation. In Findings of the Association for Computational Linguistics: ACL 2025: pp. 23923-23937.
Yngve V (1970) On getting a word in edgewise. In: Papers of the sixth regional meeting of the Chicago Linguistic Society. Chicago Linguistic Society, Chicago, pp 567–577
Acknowledgements
We are grateful to Dr Judith Schlenter for her advice on some of the statistical modelling.
Author information
Authors and Affiliations
Contributions
All authors contributed to the study conception and design. Material preparation, data collection and initial analysis were performed by YAC under the supervision of JMH. Secondary data annotation and detailed analysis were performed by JMH. All figures were prepared by JMH. The first draft of the manuscript was written by YAC and substantially expanded by JMH. All authors commented on previous versions of the manuscript. All authors read and approved the final manuscript.
Corresponding author
Ethics declarations
Conflict of interest
The authors declare no competing interests.
Additional information
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/
About this article
Cite this article
Carruthers, Y.A., Heim, J.M. Not your average podcaster: turn-taking timings and conversational role behavior in AI-generated dialogues. AI & Soc (2026). https://doi.org/10.1007/s00146-026-03365-3
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s00146-026-03365-3
