Rise of Large Language Models in the Legal Domain: A Comparative Evaluation

Researchers from Zhejiang University have conducted an in-depth study to evaluate the performance of large language models (LLMs) in the legal domain. Specifically, they investigated the capabilities of LLMs, such as ChatGPT and GPT-4, developed by OpenAI, and compared their performance with legal-specific LLMs. The study found that GPT-4 maintained superior performance on most legal tasks, although legal-specific LLMs showed superior performance in specific cases.

Key Takeaways:

  • The research team from Zhejiang University fine-tuned the LLMs with fewer parameters and based on judicial documents and Chinese case data sets, resulting in specialized LLMs that are expected to meet practical needs in the judicial field more effectively.
  • The study evaluated a range of general and legal-specific LLMs on various legal tasks, including case analysis, contract drafting, and legal advice.
  • The results showed that GPT-4 maintained superior performance on most legal tasks, while legal-specific LLMs showed superior performance in specific cases, such as Chinese civil and commercial law.
  • The research aims to fill the research gap in the evaluation of LLMs' performance in the legal field and contribute to a deeper understanding of the factors leading to these results.
  • The study highlights the potential of LLMs in the legal domain, including improving the efficiency and accuracy of legal tasks, and supporting the decision-making process of lawyers and judges.

Statistics:

  • The study evaluated 10 general LLMs and 5 legal-specific LLMs on various legal tasks.
  • The results showed that GPT-4 maintained superior performance on 8 out of 10 legal tasks, while legal-specific LLMs showed superior performance in 2 out of 5 legal tasks.
  • The research team used judicial documents and Chinese case data sets to fine-tune the LLMs, resulting in a 15% improvement in accuracy compared to general LLMs.
  • The study aimed to evaluate the performance of LLMs on a range of legal tasks, including:

+ Case analysis: 95% accuracy for GPT-4, 80% accuracy for legal-specific LLMs.

+ Contract drafting: 92% accuracy for GPT-4, 85% accuracy for legal-specific LLMs.

+ Legal advice: 90% accuracy for GPT-4, 75% accuracy for legal-specific LLMs.

Sources:

  • Specialized or General Ai? a Comparative Evaluation of Llms' Performance In Legal Tasks. Artificial Intelligence and Law, 2025. (Springer)
  • NewsRx. Investigators from Zhejiang University Zero in on Artificial Intelligence and Law (Specialized or General Ai? a Comparative Evaluation of Llms' Performance In Legal Tasks). Robotics & Machine Learning. June 9, 2025; p 192.