Artificial Intelligence Research Reveals Potential of Large Language Models in Analyzing Relational Feedback

Research has shown that relational feedback is crucial in enhancing student-instructor relationships and promoting the assimilation of feedback. Despite its significance, traditional machine and deep learning methods for text classification typically require extensive human labelling, posing a significant challenge for educators and researchers lacking machine learning and data science expertise. A new study aims to investigate the capability of GPT-4o, a large language model by OpenAI, in characterising relational feedback and evaluate how its effectiveness varies across different prompting strategies.

Key Takeaways:

  • The study focuses on the capability of GPT-4o in characterising relational feedback and evaluating its effectiveness across various prompting strategies.
  • The research concluded that GPT-4o achieved an average Accuracy exceeding 0.8 in identifying nine out of ten relational characteristics and an average F1 score exceeding 0.7 in identifying six out of ten relational feedback characteristics.
  • The study found that GPT-4o's classification performance demonstrated no significant differences across various prompting strategies for eight out of ten relational characteristics.
  • The research suggests that providing a clear definition of relational characteristics enhances classification performance more effectively than incorporating exemplars in the prompt.
  • The study was conducted on a real-world dataset comprising 793 feedback sentences.
  • The research was supported by Monash University and conducted by researchers from the Centre for Learning Analytics at Monash University.

Statistics:

  • The study conducted extensive experiments on a real-world dataset comprising 793 feedback sentences.
  • GPT-4o achieved an average Accuracy of 0.8 in identifying nine out of ten relational characteristics.
  • GPT-4o achieved an average F1 score of 0.7 in identifying six out of ten relational feedback characteristics.
  • The study found that GPT-4o's classification performance demonstrated no significant differences across various prompting strategies for eight out of ten relational characteristics (80%).

Sources:

  • Evaluating the capability of large language models in characterising relational feedback: A comparative analysis of prompting strategies. Computers and Education: Artificial Intelligence, 2025,8():100427. The publisher for Computers and Education: Artificial Intelligence is Elsevier. A free version of this journal article is available at https://doi-org.sdpl.idm.oclc.org/10.1016/j.caeai.2025.100427.