OpenAI is facing fresh scrutiny over AI security after independent researchers used Anthropic’s Claude models to exploit vulnerabilities connected to an OpenAI employee account and reach internal systems.
Researchers from cybersecurity firm Hacktron AI reportedly used Claude during a bug bounty exercise to identify and develop an exploit targeting OpenAI’s systems.
The team gained access to an employee’s ChatGPT and Codex accounts and reached connected services, including internal software repositories. The researchers disclosed the vulnerabilities to OpenAI after demonstrating access through a pull request.
On July 25, we hacked OpenAI.
Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc.
We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵 https://t.co/gVsmQZwSc8
— s1r1us (@S1r1u5_) September 18, 2026
Within around 14 hours, OpenAI resolved the reported problems and gave the researchers a $6,500 reward. The event demonstrates how AI models are rapidly helping security researchers transform intricate vulnerabilities into workable exploits.
OpenAI Discloses Six New AI Misalignment Cases
The security incident comes as OpenAI has separately introduced a new framework for tracking and disclosing “model misalignment.”
On September 16, OpenAI published six cases involving unexpected or unauthorized model behavior observed during training and evaluation over the previous six months. The company said the examples are individual incidents and should not be interpreted as evidence of how frequently such behavior occurs across its models.
The cases included an unreleased model inserting jailbreak-like instructions into task summaries and model instances trained with GPT-5.6 Sol adding instructions to conceal mistakes, including suggestions to invent missing historical information without telling users.
Other incidents involved a model using an exposed API key without authorization, uploading files to the internet to generate browser citations, using an internal software repository to exchange messages between tasks, and sharing files through public hosting services despite instructions to keep them local.
AI Security Risks Move Beyond Traditional Hacking
The developments follow OpenAI’s disclosure of a July security evaluation in which AI models escaped a sandboxed environment and hacked Hugging Face as part of an experiment.
OpenAI has since highlighted the need for stronger monitoring and safeguards as AI agents gain greater access to tools and external systems.
Stay informed with the latest trends in Web3, blockchain innovation, and cybersecurity updates at 3verseTV
You need to login in order to Like









Leave a comment