Unlocking Human-Robot Interaction through Large Language Models

Wednesday 12 March 2025


Researchers have been exploring ways to improve human-robot interaction, and a recent study has shed new light on the potential benefits of using large language models (LLMs) in this context.


The study, which involved participants interacting with an industrial robot, found that when LLMs were used to generate responses, participants rated the interaction as more engaging and trustworthy than when pre-programmed scripts were used. The researchers also observed distinct visual attention patterns in both conditions, with participants focusing on different aspects of the robot and task.


One of the key findings was that the LLM-enhanced condition led to longer fixation durations, indicating a sustained engagement with the robot and task elements. In contrast, the pre-programmed script condition resulted in shorter fixations and faster saccades, suggesting a more focused attention on specific aspects of the task.


The study also highlighted the importance of aligning interaction modalities with task requirements. The authors found that pre-programmed scripts were better suited for straightforward tasks due to their predictability and efficiency, while LLMs were more effective in complex, dynamic scenarios.


The researchers used a custom framework combining GPT-4o-mini, mobile eye-tracking, and YOLOv8-based object detection to guide interactions in real-time. They also measured energy consumption across both conditions, finding that the pre-programmed script condition delivered deterministic interactions with minimal redundancy, while the LLM-enhanced condition occasionally produced redundant instructions.


The study’s findings have implications for the development of more advanced human-robot interaction systems. By leveraging LLMs’ ability to adapt to dynamic scenarios and provide context-aware responses, researchers may be able to create more engaging and trustworthy interactions that improve overall system performance.


The authors’ approach also highlights the importance of considering the specific demands and complexity of tasks when designing interaction modalities. By aligning interaction styles with task requirements, developers can optimize system performance and user experience.


As research continues to explore the potential benefits of LLMs in human-robot interaction, this study provides valuable insights into the role of adaptability and context-awareness in enhancing engagement and trust. By refining these approaches, researchers may be able to create more sophisticated systems that improve collaboration between humans and robots.


Cite this article: “Unlocking Human-Robot Interaction through Large Language Models”, The Science Archive, 2025.


Human-Robot Interaction, Large Language Models, Engaging, Trustworthy, Industrial Robot, Visual Attention, Fixation Durations, Saccades, Task Requirements, Adaptability


Reference: Tim Schreiter, Jens V. Rüppel, Rishi Hazra, Andrey Rudenko, Martin Magnusson, Achim J. Lilienthal, “Evaluating Efficiency and Engagement in Scripted and LLM-Enhanced Human-Robot Interactions” (2025).


Leave a Reply