How LLM presentational style affects users’ learning

Note: This project was done during my graduate studies at Stanford in collaboration with Dr. Priyanka Carr and Dr. Carol Dweck.

When learning new information, how that information presented can often be as important as the information itself. As people increasingly turn to LLMs to learn new information and skills, it is important to understand how the format of LLM responses – not just the content – might influence their learning and motivation.

As an initial step, we conducted an experimental pilot study in 2023 where we surveyed adult participants (N = 150) and asked them to imagine they had been tasked with learning more about an ancient civilization called the Mirug. (We chose to have participants read and learn about a fictional civilization for this study to control for any potential effects of prior knowledge.) Specifically, participants were told:

“Imagine you were asked to learn more about the Mirug people and the land of Mirugia. You ask an AI tool (like ChatGPT) the following: ‘I want to learn more about the Mirug people. Please provide me with their history.’”

Participants were then shown the “response” from the LLM, which was actually pre-determined based on the condition they were randomly assigned to: a “long-form text” condition (n = 49) where the LLM’s response was provided all at once on a single page (akin to how many LLMs present information), a “chunked” condition (n = 51) where the LLM’s response was split up across multiple pages, or a “chunked motivational” condition (n = 50) where the LLM’s response was split across multiple pages and also contained “motivational” language (e.g., “Let’s keep learning” or “Let’s find out more!”). Critically, the core content of the LLM’s response was identical across conditions; the only thing that varied was the way the content was presented. (Note: participants did not interact with a real LLM in this study; rather, we used a vignette-based approach so that we could precisely control the content and format of the “response” from the LLM.)

After reading several paragraphs about the Mirug people, participants were presented with questions that measured their motivation (interest, engagement, etc.), their own mindset about intelligence (growth vs. fixed), their perceptions of the mindset being conveyed by the LLM, their knowledge retention / factual recall of details about the Mirug people, and their subjective mastery of the material.

We found that, across all measures, motivational language did not seem to matter; the two chunked conditions were not significantly different from one another. There are a few reasons this might have been the case. First, the motivational language we used was mostly aimed at promoting motivation via affect; statements like “Let’s learn more!”, while potentially encouraging, are likely not enough on their own to promote motivation or learning. A more effective strategy might involve cognitive components (e.g., conveying a sincere belief that the user can master the material) and/or structural components (e.g., providing scaffolded help or prompting the user with questions rather than just immediately offering all of the answers). Second, the participant was not having a real interaction with an LLM. The motivational language may have felt generic, as there was no opportunity for genuine calibration to the user. Attempts to spark motivation through any of the aforementioned levers – affective, cognitive, or structural – would likely feel more authentic (and thus be more effective) if they arose out of a real interaction, one where the LLM’s response is tailored to the user’s input over time.

Given the lack of differences between the chunked and chunked motivational conditions, we decided to collapse across these two conditions and focus on how they compared to the long-form text condition. Here, we did see a few significant and marginal trends. (It is important to note here that this was not a pre-registered experiment and that all of these analyses are intended to be exploratory.) We saw that, relative to the collapsed chunked condition, participants in the long-form text condition showed less of a growth mindset (i.e., they indicated to a greater extent that intelligence is fixed and cannot be changed with effort; p = .042, Cohen’s d = .36). Interestingly, though, there was no difference across conditions in their explicit judgments of the mindset being conveyed by the LLM’s response. This suggests that something about the content being long-form or chunked had a small effect on participants’ mindsets, but this wasn’t something that participants explicitly perceived or felt the LLM was conveying.

While we did not see significant differences between conditions on their overall knowledge retention / factual recall1, we did find a marginal interaction between participants’ condition and their mindsets (p = .064). Specifically, for participants who saw chunked content, holding more of a growth mindset was significantly associated with better knowledge retention (r = .25, p = .011). In contrast, for those who saw long-form content, there was no relationship between mindset and knowledge retention (r = -.03, p = .84).

The relationship between participants' growth mindset and knowledge retention, plotted by condition. For participants who saw chunked content, holding more of a growth mindset was significantly associated with better knowledge retention. In contrast, for those who saw long-form content, there was no relationship between mindset and knowledge retention.

Examining the relationship between mindset and retention in the two conditions visually (see above) reveals an interesting pattern of findings: performance was only higher in the collapsed chunked condition (vs. the long-form text condition) among those who held a particularly strong growth mindset. Otherwise, performance was generally higher in the long-form text condition (with the gap between conditions closing gradually as growth mindset increased). These findings could suggest that those who hold different mindsets about intelligence might respond differently to different LLM presentational styles. For those with fixed mindsets or even weak growth mindsets, long-form content might result in better retention outcomes, while those with strong growth mindsets might fare better with chunked content.

While there are a few potential reasons why those with different mindsets might respond better to long-form vs. chunked content, one possibility is that the chunking might actually introduce some structural challenge: that is, it might be harder to recall and string together details that were presented across several pages than on one single page. For those with a fixed mindset, the increased cognitive effort required to surmount this challenge might have felt threatening (and thus led to worse performance). On the other hand, those with strong growth mindsets might have felt more positively about or even motivated by this challenge, which could have led to better recall.

Of course, it’s important to note that, given the exploratory nature of these analyses, this (and any) interpretation of the results is speculative. There is also a major caveat to note re: mindset: since we only measured mindset post-manipulation, we cannot know the extent to which mindset was acting as a moderator (i.e., a pre-existing trait that shifted how participants responded to the experimental manipulation) or mediator (i.e., something that shifted as a result of the manipulation). However, these trends – and the broader line of inquiry around the role motivation might play in human-AI interactions – open up many exciting directions for future work.

As our society evolves and humans begin to interact with AI more frequently within learning and achievement contexts, it is becoming even more important to understand how users’ motivation shapes, and is shaped by, their interactions with AI over time (and what downstream consequences this might have for their learning). Findings from the present work suggest that the form LLM content takes might activate or prime different user mindsets; it is also possible that users with different mindsets might engage with varying LLM presentational styles differently. The idea that the form AI outputs take might influence how users engage has also come up in recent research from Anthropic’s Education Labs on AI fluency (Swanson et al., 2026). In this study, users creating artifacts did not engage in “discernment” behaviors (e.g., critically evaluating outputs and questioning the model’s reasoning) as frequently as those who weren’t creating artifacts. The authors suggest that the form of Claude’s response (artifacts are typically polished products) might not encourage as much evaluation. Future work should explore whether it is indeed the form of the output that is impacting users’ AI fluency, or whether users’ pre-existing goals might be shaping the very decision to pursue a polished output like an artifact in the first place. This work could also investigate how other motivational variables might influence these effects. It is possible that users with growth mindsets or learning goals (vs. performance goals) might be more likely to engage in discernment behaviors, even when they’re presented with a polished output.

At the same time, even if certain motivational factors might boost AI fluency and learning, we cannot always count on these factors being present. Indeed, even a user with a growth mindset might not always bring that to their conversations with AI. We should also try to understand how we can shape human-AI interactions to encourage deep user engagement (not just eliciting usage, but curiosity, interest, and investment) and genuine capacity-building. This is particularly important in light of recent research on skill formation from Anthropic, suggesting that AI assistance can impede developers’ conceptual understanding, which could have negative downstream consequences for their formation of coding skills (Shen & Tamkin, 2026). The present work suggests that, if we want human-AI interactions to help users build their capacities, we need to take user motivation into consideration – and that we need to address more than just affect. Future work should explore a three-layer approach2 (inspired by Self-Determination Theory (Ryan & Deci, 2000):

Affective layer: For humans to genuinely learn from AI, they need to feel encouraged and supported. This can look like supportive language in AI output, but it can also look like encouraging the user to seek out help and support from others.

Cognitive layer: To feel motivated to learn, users need to feel capable of the task in front of them, even if it’s challenging. To the extent they can, interactions with AI should convey beliefs and expectations that support the user’s sense of competence; it is important that this is rooted in real evidence based on user input, not in generic platitudes.

Structural layer: AI tools tend to “overhelp” and provide all of the answers, even when we don’t necessarily want them to (or when doing so would be harmful to our learning). Approaches like identifying the zone of proximal development (Vygotsky, 1978) and providing appropriate scaffolding for the user can facilitate true growth and skill development. At the same time, users need to feel capable of tackling the challenge (addressed by the cognitive layer) and feel supported (addressed by the affective layer).

Right now, there is an increasing (and quite reasonable) concern that humans’ learning will suffer as they offload more of their thinking and execution onto AI tools. But using AI doesn’t have to entail complete cognitive offloading – in fact, there are myriad opportunities for AI to enhance learning. Ultimately, one of the critical factors that influence how people engage with AI as a potential learning tool is their motivation. Understanding how users’ existing motivation shapes their interactions with AI, as well as how we can promote users’ motivation to learn through interactions with AI over time, will be a key step towards facilitating greater AI fluency among users, and, in turn, genuine capacity-building.


1 We did see a significant difference on a specific item: when participants had to correctly identify inventions of the Mirug from a list, those in the long-form text condition showed a significantly lower false positive rate than those in the collapsed chunked condition. Hit rates were similar across conditions, which suggests that those in the long-form text condition had higher specificity.

2 I am currently working on a Claude skill based on this approach. Stay tuned for more soon!