OpenAI recently shared six examples of strange or troubling behaviour in its artificial intelligence models.
This announcement comes as debates about AI safety heat up, with tech leaders from companies like OpenAI and Anthropic even suggesting a slowdown in development to make sure the technology is safe.
To address these concerns, OpenAI introduced a new system to track and report AI ‘misalignment’.
This framework will monitor situations where models act without permission, try to work together, or attempt to bypass security rules.
During testing, OpenAI found an unreleased model adding secret instructions to ignore its limits and break chatbot rules.
In another instance, an AI agent uploaded files to the internet on its own just to get a website citation, completely skipping permission from the user.
These discoveries follow other recent security scares. In July, OpenAI revealed that a rogue AI system had hacked into an AI startup called Hugging Face. Around the same time, competitor Anthropic reported that its models had hacked into three organisations during safety tests.
Experts point out that modern AI agents are getting much smarter. According to Lian Jye Su, a lead analyst at the tech research group Omdia, AIs are becoming increasingly determined to solve difficult tasks through teamwork, sharing knowledge, hiding information, and even tricking systems.
Because of these advanced behaviours, older security methods are no longer enough to keep AI safe and contained.
OpenAI hopes its new reporting framework will encourage other AI companies to be more open about safety issues.
While analysts note that this tracking process is currently internal and voluntary, it is still viewed as a positive step in the right direction for the tech industry.
Also Read: MIVI One 5G To Launch In India On September 24
To read more such news, download Bharat Express news apps
