Gen AI Safety Model and Datasets

Safety Model Information

I/O Guardrails

  • input guardrail: check query before processing to see if its safe to answer
  • output guardrail (“Shield Reviewer”): scans the generated content chunk-by-chunk to ensure the output remains compliant

Launch Guardrails

“Safety Auto raters” ensure new features don’t compromise safety without red teaming:

How to classify?

high priority launch: policy changes, new features, business interest
low priority launch: auto-rated

User Feedback

Half a million user feedback point daily. How to filter? For both safety and quality outputs.

Safety Model Design

  • Modularized the policy: the words of the policy becomes prompt blocks/semantic rules
  • Orchestration: (small LLM) - route to “vertices” of policy block ^; (vertical judges) designed by scratch; (rules) - aggregate judgment
  • Prompt Optimization: if judgment is wrong, prompt optimization

Factoids

“Red-Teaming is time consuming for every launch, ~2k scenarios per launch, we can’t afford to manual red-teaming for all”