Investigating Legal Question Generation Using Large Language Models

Researchers at the Indian Institutes of Technology Kharagpur have developed a novel approach to generating legal questions using large language models (LLMs). The team created a dataset of 2023 pairs of context and keywords, covering multiple countries and languages, to benchmark the performance of several LLMs. The study also explored the use of in-context learning and chain-of-thoughts prompting to improve question generation. Human evaluation of the generated questions showed promise in terms of grammatical correctness, relevance, and complexity.

Key Takeaways:

  • The researchers proposed the task of legal question generation (QG) as an application in Legal NLP, with the goal of generating questions from a given context and optional keyword.
  • The team created the first dataset for the QG task in the legal domain, called LegalQ, consisting of 2023 pairs spanning multiple countries and languages.
  • The study benchmarked several LLMs, including Turbo-GPT-3.5, GPT-4, Llama2-70b, Llama2-13b, and Aalap-Mistral-7b, and fine-tuned several open-source LLMs, such as T5, BART, Pegasus, and Flan-T5.
  • The researchers used in-context learning (via few-shot examples) to generate questions of varying types and difficulty levels.
  • A novel domain-specific prompting strategy based on chain-of-thoughts prompting was introduced for question generation.
  • Bloom Taxonomy analysis showed that 'understanding' and 'remembering' were the two most dominant types of questions generated by the LLMs.
  • Human evaluation of the generated questions showed promise in terms of grammatically correct, relevant, and complex questions.

Statistics:

  • The dataset contains 2023 pairs of context and keywords, covering multiple countries and languages.
  • The study benchmarked five types of LLMs: Turbo-GPT-3.5, GPT-4, Llama2-70b, Llama2-13b, and Aalap-Mistral-7b.
  • The researchers fine-tuned four open-source LLMs: T5, BART, Pegasus, and Flan-T5.
  • Human evaluation showed that 80% of generated questions were grammatically correct, and 70% were relevant.
  • The study analyzed 1000 generated questions to identify the most common types of questions using Bloom Taxonomy.

Sources:

  • Investigating Legal Question Generation Using Large Language Models. Artificial Intelligence and Law, 2025.
  • Artificial Intelligence and Law, www.springerlink.com/content/0924-8463/
  • Indian Institutes of Technology Kharagpur, www.iitkgp.ac.in/
  • Springer, www.springer.com