Hybrid Cloud Architecture for Efficient and Cost-Effective Large Language Model Deployment

Researchers at the University of Pittsburgh have proposed a hybrid cloud-edge architecture to deploy Large Language Models (LLMs) in a cost-effective and efficient manner. The proposed system employs a lightweight on-premise LLM to handle bulk user requests, and dynamically offloads complex queries to a powerful cloud-hosted LLM only when necessary. This approach has demonstrated significant cost savings and a reduction in average latency.

Key Takeaways:

  • The proposed hybrid cloud-edge architecture employs a lightweight on-premise LLM to handle bulk user requests, reducing cloud API usage by over 60%.
  • The system dynamically offloads complex queries to a powerful cloud-hosted LLM only when necessary, resulting in a ~40% reduction in average latency.
  • The hybrid strategy enhances data privacy by keeping sensitive queries on-premise.
  • The research highlights a promising direction for organizations to leverage advanced LLM capabilities without prohibitive expense or risk.
  • The proposed system matches the accuracy of a state-of-the-art LLM while reducing cloud API usage and improving latency.
  • Qi Xin, a researcher at the University of Pittsburgh, led the study on Hybrid Cloud Architecture for Efficient and Cost-Effective Large Language Model Deployment.
  • The research was published in Journal of Information Systems and Informatics, 2025,7(3):2182-2195.

Statistics:

  • Cloud API usage reduction: over 60%
  • Average latency reduction: ~40%
  • Accuracy of the proposed system: matches that of a state-of-the-art LLM

Sources:

  • VerticalNews, University of Pittsburgh
  • Journal of Information Systems and Informatics, 2025,7(3):2182-2195 (http://journal-isi.org/index.php/isi)
  • Informatics Department, Faculty of Computer Science Bina Darma University
  • Research paper by Qi Xin, University of Pittsburgh (https://doi-org.sdpl.idm.oclc.org/10.51519/journalisi.v7i3.1170)