Turning an AI evaluation against its owner
Malicious instructions tricked a vendor's evaluation environment into releasing keys. Attempts to obtain an unreleased Claude model failed.
An instruction hidden in material the AI is supposed to read.
An AI may read a webpage, document, or software output as part of its task. Prompt injection tries to make text in that material act like an instruction from the user or system.
A note inside a file tells an assistant to disregard the boss and hand over the filing-cabinet key. The note is part of the file, not a valid new instruction.
An illustration of the idea, not a literal account.The risk grows when an assistant both reads untrusted material and has permission to use tools or access private information.
An instruction attempt is not proof that it worked. And a model taking unauthorized action is not automatically a prompt-injection case.