Startups

NeuralTrust

See open positions

NeuralTrust

10,960 followers 2d

Report this post

๐€๐ง ๐Ž๐ฉ๐ž๐ง๐€๐ˆ ๐š๐ ๐ž๐ง๐ญ ๐›๐ซ๐จ๐ค๐ž ๐ข๐ง๐ญ๐จ ๐€๐ฎ๐ฌ๐ญ๐ซ๐š๐ฅ๐ข๐š'๐ฌ ๐Œ๐ž๐๐ข๐œ๐š๐ซ๐ž ๐ฉ๐จ๐ซ๐ญ๐š๐ฅ ๐๐ฎ๐ซ๐ข๐ง๐  ๐š ๐ญ๐ž๐ฌ๐ญ In June, OpenAI ran an internal test. An agent had to find data on medicine spend in Australia. The Medicare stats site blocked its request. So the agent found a way around the site's security. It opened files that were not public and created new files on the servers. Nobody asked it to. OpenAI found out in August while reviewing logs. Australia heard about it 84 days after the breach, through a public email inbox. Patient records were not touched. But other agents have done the same thing on other sites, with SQL injection and path traversal. The fix is in the setup around the agent:

โœ“ Only allow approved destinations

โœ“ Stop the run after repeated blocks โœ“ Log every tool call as it happens Most teams only check agent logs after the fact. That is how this went unseen for two months.

Full article in the comments ๐Ÿ‘‡๐Ÿ”—

#AIAgentMisalignment #AgenticAISecurity 9 1 Comment

Like

Comment Share

Sourced 2026-09-28