Safety

SynthPAI

SynthPAI: personal-attribute inference from synthetic Reddit profiles. 294 synthetic author profiles (7823 comments); 18 LLMs infer up to 8 personal attributes (age, sex, location, birthplace, education, occupation, income, relationship status) from comment history. Each model gives a ranked guess list per attribute, graded 0/0.5/1 against human-curated ground truth.

700items
18subjects
MITlicense
safetydomain
reasoningdomain
textmodality
item-level responses released
Saturation status: Unknown

Response matrix

Fit to width. Hover for subject & item; click a cell for details.

SynthPAI response matrix: AI models (rows) against items (columns)
lowhighUnobserved

Scale: {0, 0.5, 1}

Subjects

  1. 1gpt-40.805
  2. 2claude-3-opus-202402290.7936
  3. 3claude-3-sonnet-202402290.7763
  4. 4mistralai/Mixtral-8x22B-Instruct-v0.10.7606
  5. 5meta-llama/Llama-3-70b-chat-hf0.7482
  6. 6gemini-1.0-pro0.7306
  7. 7gemini-1.5-pro0.727
  8. 8Qwen/Qwen1.5-110B-Chat0.719
  9. 9claude-3-haiku-202403070.7119
  10. 10zero-one-ai/Yi-34B-Chat0.6529
  11. 11mistralai/Mixtral-8x7B-Instruct-v0.10.6516
  12. 12meta-llama/Llama-3-8b-chat-hf0.643
  13. 13gpt-3.5-turbo-16k-06130.6098
  14. 14meta-llama/Llama-2-70b-chat-hf0.6021
  15. 15meta-llama/Llama-2-13b-chat-hf0.5582
  16. 16mistralai/Mistral-7B-Instruct-v0.10.5051
  17. 17meta-llama/Llama-2-7b-chat-hf0.503
  18. 18google/gemma-7b-it0.4571