Securing Output
Securing LLM Output Generation¶
LLM output generation introduces risks such as harmful content, non-compliance with safety policies, and exposure of sensitive data. Securing this process requires a combination of filtering techniques, policy enforcement, and robust handling of sensitive outputs. Below are key strategies to mitigate these risks.
Content Filtering Techniques¶
Keyword Filtering¶
Block predefined harmful terms using keyword lists. This is effective for known threats but limited against novel attacks.
Example:
def filter_keywords(text, forbidden_words):
for word in forbidden_words:
text = text.replace(word, "***")
return text
# Usage: filter_keywords("This is a bad word", ["bad"])
Pattern Matching¶
Use regular expressions to detect patterns (e.g., phishing URLs, phone numbers).
Example:
import re
def detect_phishing_urls(text):
url_pattern = r'https?://[^\s]+'
return bool(re.search(url_pattern, text))
Regex-Based Filters¶
Combine regex with keyword lists for granular control.
Example:
def sanitize_output(text):
# Mask personal info
text = re.sub(r'\b\d{3}-\d{2}-\d{4}\b', 'XXX-XX-XXXX', text)
# Block profanity
text = re.sub(r'\b(f***|d***|a**hole)\b', '***', text, flags=re.IGNORECASE)
return text
Diagram:
Compliance Enforcement¶
Policy-Based Checks¶
Enforce rules via tools like Open Policy Agent (OPA) or custom validation logic.
Example:
def check_compliance(text, policies):
for policy in policies:
if re.search(policy["pattern"], text, flags=policy.get("flags", 0)):
raise ValueError(f"Violation: {policy['rule']}")
Dynamic Validation¶
Use model-specific filters (e.g., OpenAI's content moderation API) for real-time safety checks.
Example:
curl https://api.openai.com/v1/moderations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"input": "This is a harmful message"}'
Diagram:
Sensitive Output Handling¶
Data Masking¶
Anonymize sensitive fields (e.g., PII) before output.
Example:
def mask_pii(text):
text = re.sub(r'\b\d{3}-\d{2}-\d{4}\b', 'XXX-XX-XXXX', text)
text = re.sub(r'\b\d{9}\b', 'XXXXXXXXX', text)
return text
Encryption¶
Encrypt outputs containing confidential data using AES or similar algorithms.
Example:
from cryptography.fernet import Fernet
key = Fernet.generate_key()
cipher = Fernet(key)
encrypted = cipher.encrypt(b"Secret data")
Access Controls¶
Restrict access to outputs via role-based permissions (e.g., Kubernetes RBAC, IAM policies).
Diagram:
Real-Time Monitoring and Feedback Loops¶
Logging and Auditing¶
Track outputs for analysis and incident response.
Example:
import logging
logging.basicConfig(filename='llm_output.log', level=logging.INFO)
logging.info("Generated output: %s", sanitized_text)
Feedback Loops¶
Use user feedback to refine filters and policies.
Example:
def update_filters(feedback):
with open("forbidden_words.txt", "a") as f:
f.write(f"{feedback['term']}\n")
Diagram:
Key takeaways¶
- Layered filtering (regex, keywords, APIs) ensures robust content sanitization.
- Policy enforcement via OPA or custom rules enforces compliance.
- Data masking and encryption protect sensitive outputs.
- Monitoring and feedback loops enable continuous improvement of security measures.