OpenAI adds an AI safety layer to detect misuse without retaining enterprise data
PolicyOpenAI is adding a new safety capability that allows enterprises to detect misuse of its AI systems across multiple interactions without retaining prompts or responses, enabling risk monitoring while preserving its Zero Data Retention (ZDR) commitments. “OpenAI does not retain…prompts or model responses after a request is processed,” the company said in a blog post, describing its ZDR approach. The new system, called Private Safety Processing, is “designed to identify patterns across related interactions without giving OpenAI personnel access to the underlying content.” The capability is being tested with eligible enterprise and API customers and is intended to address a limitation in existing safety controls that evaluate interactions individually, making it harder to detect risks that unfold over time. What does Private Safety Processing do Private Safety Processing is designed to extend existing safety systems by correlating activity across related interactions rather than analyzing each prompt in isolation, according to OpenAI. Under the model, automated systems analyze interactions and generate “a narrowly defined signal indicating the type of activity involved,” instead of exposing the underlying prompts or responses, according to OpenAI. The system can operate whether customer data remains within enterprise-controlled infrastructure or is stored by OpenAI with encryption keys controlled by the customer, the company said. “In both cases, automated systems can identify potential misuse and return limited safety signals without exposing the underlying prompts or responses to OpenAI personnel,” the post added. Why existing safety controls fall short OpenAI said the new capability is intended to address a gap in the detection of AI risks. “The most serious AI safety risks are not always visible in a single interaction,” the company said, noting that harmful intent may only become clear when multiple interactions are viewed together. Such risks include repeated attempts to probe safeguards, coordinated activity across accounts, and misuse that emerges over a sequence of interactions, according to the company. As AI systems take on longer and more complex tasks, evaluating individual prompts in isolation can limit the ability to identify such patterns, OpenAI said in the post. Diverging approaches to AI safety The introduction of Private Safety Processing highlights differing approaches to safety among AI providers. OpenAI said the system is designed to detect misuse patterns across interactions while preserving zero data retention. By contrast, some providers retain customer interaction data for a period of time to support safety monitoring, reflecting a different approach to identifying risks that span multiple requests. Sanchit Vir Gogia, chief analyst at Greyhound Research, said the difference lies in how evidence is handled rather than whether signals are used. “This is a disagreement about how much raw content you need besides a signal you are keeping regardless, rather than privacy against surveillance,” Gogia said. He added that both approaches rely on derived indicators but differ in where investigation data resides. “Anthropic wants enough content to investigate the case. OpenAI wants the customer to hold the case while the provider holds the alarm,” he said. Signal-based detection and verification The system’s reliance on signals rather than direct data access shifts how enterprises verify and investigate incidents, analysts noted. “The architecture is entirely viable. Security has worked from derived indicators for a generation. The difficulty is verification, not feasibility,” Gogia said. He added that detecting behavior across time requires retaining some form of representation. “A system cannot detect behaviour across time unless it remembers something across time,” he said. Private Safety Processing “is privacy-preserving abuse detection. It is not an enterprise forensic record, and OpenAI does not claim it is,” he said. Implications for regulated sectors According to Apeksha Kaushik, senior principal analyst at Gartner, the approach could influence AI adoption in industries with strict data requirements. “Privacy-preserving safety models, such as those employing Zero Data Retention (ZDR), represent an emerging approach that may lower barriers to AI adoption in regulated sectors like financial services and healthcare,” she said. Such models “may help organizations address certain privacy requirements and may align with frameworks such as GDPR and HIPAA, contingent on specific implementation details and regulatory guidance,” she noted. Kaushik added that organizations should evaluate such approaches against their compliance requirements. “Organizations are encouraged to consult with their compliance and legal teams to determine whether such approaches meet their specific regulatory and operational requirements,” she said. OpenAI said enterprises retain control over their data under this model and can investigate alerts using their own systems. Customers can also choose to share relevant data with the company to support investigations or appeals. Analysts say this shifts responsibility toward enterprises. “Zero Data Retention does not remove the forensic burden. It relocates it,” Gogia said.
Read full story at CSO Online →