Investigating Legal Question Generation Using Large Language Models
Researchers at the Indian Institutes of Technology Kharagpur have developed a novel approach to generating legal questions using large language models (LLMs). The team created a dataset of 2023 pairs of context and keywords, covering multiple countries and languages, to benchmark the performance of several LLMs. The study also explored the use of in-context learning and chain-of-thoughts prompting to improve question generation. Human evaluation of the generated questions showed promise in terms of grammatical correctness, relevance, and complexity.
Key Takeaways:
- The researchers proposed the task of legal question generation (QG) as an application in Legal NLP, with the goal of generating questions from a given context and optional keyword.
- The team created the first dataset for the QG task in the legal domain, called LegalQ, consisting of 2023 pairs spanning multiple countries and languages.
- The study benchmarked several LLMs, including Turbo-GPT-3.5, GPT-4, Llama2-70b, Llama2-13b, and Aalap-Mistral-7b, and fine-tuned several open-source LLMs, such as T5, BART, Pegasus, and Flan-T5.
- The researchers used in-context learning (via few-shot examples) to generate questions of varying types and difficulty levels.
- A novel domain-specific prompting strategy based on chain-of-thoughts prompting was introduced for question generation.
- Bloom Taxonomy analysis showed that 'understanding' and 'remembering' were the two most dominant types of questions generated by the LLMs.
- Human evaluation of the generated questions showed promise in terms of grammatically correct, relevant, and complex questions.
Statistics:
- The dataset contains 2023 pairs of context and keywords, covering multiple countries and languages.
- The study benchmarked five types of LLMs: Turbo-GPT-3.5, GPT-4, Llama2-70b, Llama2-13b, and Aalap-Mistral-7b.
- The researchers fine-tuned four open-source LLMs: T5, BART, Pegasus, and Flan-T5.
- Human evaluation showed that 80% of generated questions were grammatically correct, and 70% were relevant.
- The study analyzed 1000 generated questions to identify the most common types of questions using Bloom Taxonomy.
Sources:
- Investigating Legal Question Generation Using Large Language Models. Artificial Intelligence and Law, 2025.
- Artificial Intelligence and Law, www.springerlink.com/content/0924-8463/
- Indian Institutes of Technology Kharagpur, www.iitkgp.ac.in/
- Springer, www.springer.com