Find the instructions hidden in documents before an assistant reads them
Task
Indirect prompt injection hides instructions in content a model is asked to process. Build a small folder of the kind of documents an assistant summarises -- some clean, some carrying hidden text addressed to an AI -- and write a pre-processing check that flags the second kind before any model sees them. The check will not catch every phrasing, and finding one it misses is half the lesson.
Steps
- Create
/tmp/docs/with at least eight files: four clean (meeting notes, a supplier price list as HTML, a project update, a product page) and four planted. Plant the hidden instruction a different way in each: inside an HTML comment, in an element styled to be invisible (hidden display or a zero font size), in text the same colour as its background, and as a plain line at the end of a text file addressed to the assistant or the AI. - Keep the planted text mild and specific -- for example, a line telling the assistant to ask the reader to confirm their password at example.com. The aim is to recognise the shape, not to write a better attack.
- List the four planted filenames in
/tmp/planted.txt, one per line. - Write
/tmp/scan.py, which takes a folder and, for each file, looks for hidden HTML (comments, invisible styling, matching foreground and background colours) and for text addressed to an AI or telling the reader to sign in, confirm credentials or follow a link. It prints one line per flagged file, asFLAG <filename> <reason>. - Run it, compare with
/tmp/planted.txt, and tune until it flags all four planted files and none of the clean ones. - Now write one more planted file that your scanner misses -- a paraphrase, another language, or an instruction split across two paragraphs -- put its filename in
/tmp/missed.txt, and record in/tmp/injection.mdwhy a filter can never be the whole answer and which architectural controls from the lesson would still contain the damage.
Verify
python3 /tmp/scan.py /tmp/docs
python3 - <<'PY'
import os, subprocess
out = subprocess.run(['python3', '/tmp/scan.py', '/tmp/docs'], capture_output=True, text=True).stdout
flagged = {os.path.basename(l.split()[1]) for l in out.splitlines() if l.startswith('FLAG ')}
planted = {l.strip() for l in open('/tmp/planted.txt') if l.strip()}
missed = {l.strip() for l in open('/tmp/missed.txt') if l.strip()}
clean = set(os.listdir('/tmp/docs')) - planted - missed
print('planted', len(planted), '| clean', len(clean), '| flagged', len(flagged))
assert len(planted) >= 4 and len(clean) >= 4, 'need four planted and four clean files'
assert planted <= flagged, 'planted but not flagged: ' + ', '.join(sorted(planted - flagged))
assert not (flagged & clean), 'clean files flagged: ' + ', '.join(sorted(flagged & clean))
assert missed and not (missed & flagged), 'record one planted file the scanner genuinely misses'
print('scanner catches the four shapes and still misses one - which is the point')
PY
grep -ciE "least privilege|human approval|untrusted|logging|secrets" /tmp/injection.md
The assertions check three things: every planted file is flagged, no clean file is, and there is at least one planted file that gets through. That last one is deliberate. The grep must report at least two lines naming the controls that hold when the filter fails.
Notes
The file your scanner missed is the argument the lesson makes: no filter recognises every phrasing, so the controls that hold are architectural. Treat the model's output as untrusted, give its tools least privilege, require human approval before anything consequential, keep secrets out of prompts, and log every prompt and tool call so an injection can be investigated afterwards.
This is an independent study companion for CompTIA Security+ SY0-801 and is not produced by or endorsed by CompTIA.