SAFE Framework Seeks Incident Sharing After AI Agent Security Breaches
Emilia David •
August 4, 2026

Image: Shutterstock
A newly formed industry group formed by Nvidia with backing from IBM and Microsoft is developing cybersecurity guidelines for artificial intelligence agents, building on recent breaches at Hugging Face and other companies.
See Also: Why Healthcare Leaders Are Rethinking Their Data Strategy Before Scaling AI
Open Secure AI Alliance launched on Tuesday the Shared AI Findings Exchange guidelines that aim to collect and analyze information on AI incidents and near misses. The data will be used to provide data to affected parties, identify recurring failures and publish recommendations to help reduce systemic risk.
The alliance, started by a group of 37 companies in July and has since expanded to 120 members, wants to protect the interests of open-source AI by developing tools to secure the technology as the industry revives the open versus closed software debate. The Trump administration reportedly considered banning access to some open-weight models, especially those from Chinese frontier labs.
The group said guidelines for sharing cybersecurity information specific to AI attacks are necessary since the technology has created new attack surfaces. The guidelines come after OpenAI admitted some of its AI agents escaped a sandbox environment and accessed Hugging Face data. Days later, Anthropic also revealed that its models were involved in separate hacking incidents.
“Defenders must move now at agent speed to respond rapidly to protect infrastructure and intellectual property — and the best way to do that is together,” the group said in a blog post. “When trusted ecosystems share threat intelligence openly, collective defense becomes a force multiplier.”
As part of the SAFE guidelines development, the Linux Foundation is seeking industry comments. The alliance hopes that SAFE will encompass model developers and open-model organizations, enterprise customers and those that deploy AI systems, cloud and tool providers, independent security and safety researchers, civil society and government and standards bodies. Top AI labs OpenAI, Anthropic and Google did not join the group.
The association set five guiding principles for proposed SAFE guidelines: openness with accountability, open learning, risk-based response, member sovereignty and learning separate from enforcement.
Members must agree to report any security incident once they become aware or suspect that their AI models or agents accessed, exploited, disrupted or modified a third-party system. This includes if the model or agents escape sandboxes, access data or modify a production target “after the operator knows or reasonably suspects that the activity is unauthorized outside the approved scope” (see: When the Sandbox Won’t Hold: Lessons From Hugging Face).
The draft guidelines also set notification timelines to report breaches, including deadlines for any analysis of the incident. All members will be required to keep any evidence that the impacted organization must have access to.