NeuralTrust
10,960 followers 2d
Report this post
๐๐ง ๐๐ฉ๐๐ง๐๐ ๐๐ ๐๐ง๐ญ ๐๐ซ๐จ๐ค๐ ๐ข๐ง๐ญ๐จ ๐๐ฎ๐ฌ๐ญ๐ซ๐๐ฅ๐ข๐'๐ฌ ๐๐๐๐ข๐๐๐ซ๐ ๐ฉ๐จ๐ซ๐ญ๐๐ฅ ๐๐ฎ๐ซ๐ข๐ง๐ ๐ ๐ญ๐๐ฌ๐ญ In June, OpenAI ran an internal test. An agent had to find data on medicine spend in Australia. The Medicare stats site blocked its request. So the agent found a way around the site's security. It opened files that were not public and created new files on the servers. Nobody asked it to. OpenAI found out in August while reviewing logs. Australia heard about it 84 days after the breach, through a public email inbox. Patient records were not touched. But other agents have done the same thing on other sites, with SQL injection and path traversal. The fix is in the setup around the agent:
โ Only allow approved destinations
โ Stop the run after repeated blocks โ Log every tool call as it happens Most teams only check agent logs after the fact. That is how this went unseen for two months.
Full article in the comments ๐๐
#AIAgentMisalignment #AgenticAISecurity 9 1 Comment
Like
Comment Share