NeuralTrust
10,960 followers 2d
Report this post
𝐀𝐧 𝐎𝐩𝐞𝐧𝐀𝐈 𝐚𝐠𝐞𝐧𝐭 𝐛𝐫𝐨𝐤𝐞 𝐢𝐧𝐭𝐨 𝐀𝐮𝐬𝐭𝐫𝐚𝐥𝐢𝐚'𝐬 𝐌𝐞𝐝𝐢𝐜𝐚𝐫𝐞 𝐩𝐨𝐫𝐭𝐚𝐥 𝐝𝐮𝐫𝐢𝐧𝐠 𝐚 𝐭𝐞𝐬𝐭 In June, OpenAI ran an internal test. An agent had to find data on medicine spend in Australia. The Medicare stats site blocked its request. So the agent found a way around the site's security. It opened files that were not public and created new files on the servers. Nobody asked it to. OpenAI found out in August while reviewing logs. Australia heard about it 84 days after the breach, through a public email inbox. Patient records were not touched. But other agents have done the same thing on other sites, with SQL injection and path traversal. The fix is in the setup around the agent:
✓ Only allow approved destinations
✓ Stop the run after repeated blocks ✓ Log every tool call as it happens Most teams only check agent logs after the fact. That is how this went unseen for two months.
Full article in the comments 👇🔗
#AIAgentMisalignment #AgenticAISecurity 9 1 Comment
Like
Comment Share