AINews

OpenAI paid $6,500 after Claude helped crack its community help forum

A San Francisco cybersecurity startup with fewer than ten people used Anthropic’s Claude to break into an OpenAI employee account through the community help forum and push Codex toward the company’s internal code repository, Zayne Zhang / Hacktron via Business Insider reports. They stopped before reading internal code, disclosed, and collected a $6,500 bounty.

A bounty is money a company pays researchers who report a security hole. Codex is OpenAI’s coding agent that can suggest code changes. Anthropic’s Cyber Verification Program loosens Claude’s usual refusal to help with hacking so approved researchers can test systems.

Hacktron launched less than a year ago. CEO Zayne Zhang told BI the team is investigating whether frontier labs have gaps AI agents can exploit. The July finding: any user or OpenAI employee logging into OpenAI’s community help forum could have had ChatGPT and Codex accounts compromised.

Using Claude under the Cyber Verification Program, the team hacked an OpenAI employee’s account and prompted that employee’s Codex account to suggest changes in OpenAI’s internal code repository. They stopped without accessing internal code and flagged OpenAI. The disclosure published Sunday.

OpenAI’s spokesperson thanked the researchers and said the company “narrowed the permissions on Community sign-in tokens and revoked affected tokens and sessions.” Anthropic did not comment to BI.

Safety week became an offense story: frontier models are already useful weapons against frontier labs, and the receipts include a paid bounty.

Sources

One thought on “OpenAI paid $6,500 after Claude helped crack its community help forum”

Leave a Reply

Your email address will not be published. Required fields are marked *