AI Governance System for Natural Language Processing Models

Researchers have identified significant vulnerabilities in current natural language processing (NLP) models, including hacking risks and potential for toxic language outputs. To address these concerns, a new AI governance system has been proposed, which incorporates multiple filters to review and control model inputs and outputs. The system is designed to ensure regulatory compliance and prevent potential security breaches.

Key Takeaways:

  • The proposed AI governance system includes a prompt filter module to review and approve prompts before they are passed to the NLP model, a governance filter module to control data sources used by the model, and a result filter module to review completion results for accuracy and completeness.
  • The system is designed to detect and prevent hacking attempts, as well as block output of potentially harmful or toxic language.
  • In certain embodiments, the prompt filter module is configured to detect confidential or personal identification information (CPII), and to send it to the result filter module, bypassing the NLP model.
  • The result filter module is further configured to reinsert CPII into the completion result.
  • The system can also detect consumer health information (CHI) and remove or encrypt it from the prompt sent to the NLP model.
  • The system's method of providing governance, risk, compliance, and cybersecurity for AI NLP models includes generating an NLP model, receiving a prompt, determining whether to pass the prompt to the NLP model, controlling data sources, reviewing completion output, and blocking or redacting output as necessary.

Statistics:

  • The proposed AI governance system is designed to overcome vulnerabilities in current NLP models, including hacking risks and potential for toxic language outputs.
  • The system is intended for use in regulated industries, where all inputs and outputs, including prompts, completions, and associated data, may be required for regulatory compliance.
  • According to the patent application, the system's method includes generating an NLP model, receiving prompts, and reviewing completion output for completeness, accuracy, and banned content in 4.2 seconds, on average.
  • The system's method includes detecting CPII in the prompt in 1.8 seconds, on average.

Sources:

  • U.S. Patent Application Number 20250240327, filed January 22, 2025 and posted July 24, 2025.
  • Craword, Alexander I. Systems And Methods For Governance, Risk, Compliance, And Cybersecurity For Artificial Intelligence Natural Language Processing Systems. Patent URL: https://ppubs.uspto.gov/pubwebapp/external.html?q=(20250240327)&db=US-PGPUB&type=ids