From Security to Trustworthy AI
Last edited: September 9, 2026SCB Encryption
Gen AI Safety Model and Datasets
Last edited: September 9, 2026Safety Model Information
I/O Guardrails
- input guardrail: check query before processing to see if its safe to answer
- output guardrail (“Shield Reviewer”): scans the generated content chunk-by-chunk to ensure the output remains compliant
Launch Guardrails
“Safety Auto raters” ensure new features don’t compromise safety without red teaming:
How to classify?
high priority launch: policy changes, new features, business interest
low priority launch: auto-rated
User Feedback
Half a million user feedback point daily. How to filter? For both safety and quality outputs.
Houjun's Academic Home Page
Last edited: September 9, 2026👋 Howdy, I'm Houjun Liu!
I’m a third-year coterminal MSCS and BSCS student in the Computer Science Department at Stanford University, grateful to be advised by Prof. Mykel Kochenderfer. In the course of my research, I have also had the fortunate opportunity to work with Stanford NLP under Prof. Chris Manning, CMU TalkBank under Prof. Brian MacWhinney, and Prof. Xin Liu at UC Davis Engineering. I am affiliated with the Stanford NLP Group and Stanford Intelligent Systems Lab. I previously visited Microsoft Research Frontiers as a research intern.
independence to irrelavent alternatives
Last edited: September 9, 2026aggregating x vs. y should ignore any possible z
