MERENA
Advisory Alert

Hidden messages can trick AI safety checks more easily than you might think

Researchers found that hidden instructions can be disguised inside scrambled text and still be understood by some AI assistants, allowing harmful directions to slip past built-in safety checks. This means people should be careful about trusting AI outputs when the AI is asked to read, translate, decode, or summarize unusual text from unknown sources.


Who is at risk

Anyone who uses AI assistants for work, school, or personal tasks is at risk, especially if they paste in strange text, coded messages, or content from untrusted websites, emails, or chats.

What to watch for

Watch for AI tools acting oddly after you paste in unusual text, such as ignoring your instructions, giving unsafe advice, exposing private information, or pushing you to click links or run commands.

What to do

Avoid giving AI assistants scrambled or suspicious text from unknown sources, do not follow risky instructions they produce, and use trusted review steps before acting on AI-generated advice.