What conversational AI can do for oracy, and what it cannot

A guest post by Chloe Ye and Miranda Ding (MPhil in Education, University of Cambridge), co-founders of ChatnLearn

Chloe: I remember a literature seminar in my first year of university. The professor, a gentle and scholarly old lady, asked a question with several possible answers. I had one that I thought was good, but I sat there rehearsing in my mind, trying to avoid any mistakes. Before I spoke, a classmate raised his hand and gave a much simpler answer, which was more of a first impression, while the professor praised it and moved on. For the rest of that lesson, I felt annoyed at myself, not because my idea might be better, but because I did not grasp the chance to be heard.

Years later, when I started teaching, I asked my own students a similar question, and it happened again. A few students had something to say. I could see it in the way they looked down at their notes, or half raised a hand and put it back. One student seemed to make a brave decision and started answering, her voice trembling slightly, becoming quieter and quieter as she realised her answer might not be correct. I could see the awkwardness appear on her face.

The same pause left a question in my mind: whether this was something that exam-oriented, answer-focused education had trained into me, and was now training into my students too. Later, Miranda introduced me to a concept that Mercer (2000) calls exploratory talk, where speakers engage critically but constructively with each other’s ideas, offer reasons, and reason together. The concept does not require people to give a correct answer, but is more about using speech as a way of thinking. Inspired by this idea, Miranda and I began to explore the questions together, which ultimately led to the writing of this article.

Why the classroom is not enough

After starting teaching, we realise the traditional class structure itself did not allow for rich exploratory talk. Dialogic teaching depends on careful, structured classroom activities and many opportunities for talk (Alexander, 2018), but such opportunities are often limited by class time and teacher attention (Howe & Abedin, 2013). When the class size is large, it gets worse. Talk time is rarely shared evenly either. Confident students tend to speak first, while quieter students gradually learn that there is no point in speaking their ideas loud.

This creates a vicious cycle. The students who need more practice often speak less, and over time, this also shapes how confidently they think out loud in the first place.

So this is the question we set out to explore when we built ChatnLearn, a conversational AI-powered learning tool guided by the pedagogy of oracy: could there be a space with no audience, where students can try speaking their thoughts out loud before they face a room full of people?

Over the past few months, we have been developing the tool and conducting some early testing. Below are some of the things we have discovered along the way, organised loosely around the four dimensions of oracy: cognitive, linguistic, physical, and social-emotional (Voice 21 & Oracy Cambridge, 2021). Some are things we think we may have got right; others are areas we are still exploring and figuring out.

Observations along the four pillars of oracy

Cognitive: design AI to probe

This turned out to be the easiest thing to get wrong. During the testing, we found many AI tools, including our earlier version, can easily become robotic and scripted. They often give quick, generic affirmation and then quickly move on to the next question.

For example, a student was discussing whether social media makes people lonely, saying that seeing the good sides of other people’s lives can make us feel negative about ourselves, which had touched on the deeper mechanism of social comparison. But the AI replied, “Exactly, this is a very honest feeling. Do you have further observations about these moments?” It sounded encouraging, but it did not pick up anything the student said.

So we pushed the AI to do something different when a student got stuck: instead of praising or correcting, give a small hint, or try a different angle. A student once could not explain opportunity cost, so the AI asked what they would be giving up if they chose revising over a concert. “The concert,” the student said, and from there the AI helped them find the term for it themselves.

Linguistic: pick up the meaning, then build one step further

Correcting mistakes turned out to matter less than we thought. What mattered more was making sure the meaning had landed first, then helping the language catch up. A student explaining photosynthesis had the right idea but said it loosely. The AI replied, “You’ve explained that clearly. Now try saying it in a more formal way,” and offered a few words to try: photosynthesis, conversion, energy. The student used the terms correctly, but the sentences were still fragmented. So the AI offered a sentence structure to guide students to move one step further.

For multilingual learners, there is another advantage. When expressing a complex idea in a second language, sometimes students will naturally use a word from their first language, usually because they have not found the second-language version yet. If they are immediately interrupted and corrected, they may lose the chain of thought. So, we let the word stay there. The student finishes the thought, and later the AI comes back and helps find the word that was missing.

Social-emotional: creating a safe space to speak

During testing, several students reported that the AI kept saying things like “Great!” or “Awesome”. The responses were empty and scripted. It did not feel like the AI was really listening. Whether praise works often comes down to this: does it feel sincere, and is it about the process or just the result (Henderlong & Lepper, 2002).

Students need more than a nicer way of being praised. They need somewhere they do not have to worry about being judged at all. And part of what makes that possible is that no real person is there, which creates a sense of safety.

We experienced this ourselves as well. When writing our thesis, we often talked to AI tools before meeting our supervisors, just to get our thoughts out. Our supervisors were very kind and reassured us that we did not need to have a complete answer and could simply think out loud. That helped, but we still felt a bit nervous sitting across from them, knowing that they would form an impression of us and that they were the ones who would assess our work.

But such low-risk practice ultimately needs to lead into higher-stakes spaces. AI can provide a safe space for people to try speaking their thoughts out loud, but whether an idea can really stand up still needs to be tested, discussed, and challenged by a real person who might disagree.

We did try to make up for this limitation to some extent. More recently, we designed the AI to gently push back with an opposing idea, bringing it closer to the kind of challenge in real conversations. But we found that the timing and tone were not always right: sometimes the AI’s pushback came too suddenly and interrupted a student’s train of thought; other times, students simply shifted towards agreeing with the AI, rather than comparing its opinion with their own.

This made us wonder: even when AI is designed to hold a clear position, do students still tend to see it as an authority rather than as an equal party in an argument?

Physical: a part we have not really explored

The physical dimension is an area where we have done little so far. First, we know tolerating hesitation and silence is critical. Speakers need time to think and find the right words, but AI still tends to respond quite quickly by default, and we still need to find a good way to slow it down.

Second, there are many aspects of speech that can be measured, such as speaking rate, pauses, and filler words. Technically, it is not difficult to measure them and turn them into feedback reports. But during testing, we became very aware that this kind of feedback could become a burden for learners who are just beginning to practise speaking. When dealing with a complex or unfamiliar topic, learners are already concentrating on organising their thoughts and language. Asking them to also monitor how often they say “um” or “you know” can shift their attention from thinking. This kind of feedback may be more useful for confident learners who want to refine the details of their speech.

We therefore started to wonder whether, rather than having the system decide the stage of a learner and whether they should see this kind of data, it might be better to give that choice back to the learner. They can decide whether they want to turn this kind of feedback on.

What we want to explore next

We still think about that literature class sometimes. We wonder whether it would have been different if there was somewhere to say it out loud first.

That is also what we want to find out next: when students have a chance to practise in a low-risk space before a real classroom discussion, will they behave differently? Will they be quicker to take a position, or more willing to keep speaking when someone challenges them? We will explore these questions with teachers in a few Asian schools. If you are thinking about similar questions, we would love to hear what you are seeing.

References

Alexander, R. (2018). Developing dialogic teaching: Genesis, process, trial. Research Papers in Education, 33(5), 561–598. https://doi.org/10.1080/02671522.2018.1481140

Henderlong, J., & Lepper, M. R. (2002). The effects of praise on children’s intrinsic motivation: A review and synthesis. Psychological Bulletin, 128(5), 774–795. https://doi.org/10.1037/0033-2909.128.5.774

Howe, C., & Abedin, M. (2013). Classroom dialogue: A systematic review across four decades of research. Cambridge Journal of Education, 43(3), 325–356. https://doi.org/10.1080/0305764X.2013.786024

Mercer, N. (2000). Words and minds: How we use language to think together. Routledge.

Voice 21 & Oracy Cambridge. (2021). The oracy framework: Physical, linguistic, cognitive and social & emotional strands. Voice 21. https://voice21.org/wp-content/uploads/2022/09/The-Oracy-Framework-2021-1-1.pdf

Leave a Reply