Benchmarking Open-Source Large Language Models for Phishing URL Detection

Researchers from Universitas Islam Riau, in collaboration with international colleagues, have conducted a comprehensive benchmarking study of 21 state-of-the-art open-source large language models (LLMs) for phishing URL detection. According to the study, large open-source LLMs (27B parameters) achieve performance exceeding 90% F1-score without fine-tuning, closely matching proprietary models. The research emphasizes the practical potential of open-source LLMs for phishing detection and provides insights for effective prompt engineering in cybersecurity applications.

Key Takeaways:

  • The study evaluated 21 state-of-the-art open-source LLMs, including Llama3, Gemma, Qwen, Phi, DeepSeek, and Mistral, for phishing URL detection.
  • The research demonstrated that large open-source LLMs (27B parameters) achieve performance exceeding 90% F1-score without fine-tuning.
  • Few-shot prompting consistently delivers the highest accuracy (91.24% F1 with Llama3.3_70b) in phishing URL detection.
  • Chain-of-thought prompting significantly lowers accuracy and increases inference time in phishing URL detection.
  • The study highlights smaller models (7B-27B parameters) offering strong performance with substantially reduced computational costs in phishing URL detection.

Statistics:

  • 90% F1-score achieved by large open-source LLMs (27B parameters) without fine-tuning for phishing URL detection.
  • 91.24% F1-score achieved by few-shot prompting with Llama3.3_70b in phishing URL detection.
  • 7B-27B parameters of models offering strong performance with substantially reduced computational costs in phishing URL detection.

Sources:

  • Benchmarking 21 Open-Source Large Language Models for Phishing Link Detection with Prompt Engineering. Information, 2025,16(5):366. (Information - http://www.mdpi.com/journal/information/)
  • NewsRx. Research from Universitas Islam Riau Yields New Findings on Information Technology (Benchmarking 21 Open-Source Large Language Models for Phishing Link Detection with Prompt Engineering). Information Technology Newsweekly. June 10, 2025; p 755.