Hidden messages can trick AI safety checks more easily than you might think
Researchers found that hidden instructions can be disguised inside scrambled text and still be understood by some AI assistants, allowing harmful directions to slip past built-in safety checks. This means people should be careful about trusting AI outputs when the AI is asked to read, translate, decode, or summarize unusual text from unknown sources.
Who is at risk
Anyone who uses AI assistants for work, school, or personal tasks is at risk, especially if they paste in strange text, coded messages, or content from untrusted websites, emails, or chats.
What to watch for
Watch for AI tools acting oddly after you paste in unusual text, such as ignoring your instructions, giving unsafe advice, exposing private information, or pushing you to click links or run commands.
What to do
Avoid giving AI assistants scrambled or suspicious text from unknown sources, do not follow risky instructions they produce, and use trusted review steps before acting on AI-generated advice.
