Safety Model Information
I/O Guardrails
- input guardrail: check query before processing to see if its safe to answer
- output guardrail (“Shield Reviewer”): scans the generated content chunk-by-chunk to ensure the output remains compliant
Launch Guardrails
“Safety Auto raters” ensure new features don’t compromise safety without red teaming:
How to classify?
high priority launch: policy changes, new features, business interest
low priority launch: auto-rated
User Feedback
Half a million user feedback point daily. How to filter? For both safety and quality outputs.
Safety Model Design
- Modularized the policy: the words of the policy becomes prompt blocks/semantic rules
- Orchestration: (small LLM) - route to “vertices” of policy block ^; (vertical judges) designed by scratch; (rules) - aggregate judgment
- Prompt Optimization: if judgment is wrong, prompt optimization
Factoids
“Red-Teaming is time consuming for every launch, ~2k scenarios per launch, we can’t afford to manual red-teaming for all”
