Put an AI agent on a short leash, then write the rules for the people using AI
Task
Take the manifest of a hypothetical email-triage agent -- the tools it can call, the scopes its tokens carry, whether its actions need approval -- start from the over-permissive version teams actually ship, and cut it down to least privilege with a checker that proves it. Then write the acceptable-use rules that say which data may go into which AI tools, and test them against requests a colleague might really make.
Steps
- Write
/tmp/agent-before.jsondescribing the agent as it is often first deployed: a list oftools, each withname,scopes(such asmail.read,mail.send,files.readwrite.all,db.write),approval(noneorhuman),token_ttl_minutesand, for any tool that runs code,sandbox; plus top-levelsystem_promptandaudit_logfields. Make it realistically bad: broad scopes, no approval, day-long tokens, a key pasted into the system prompt, logging off. - Write
/tmp/check_agent.py, which takes a manifest path, prints one line per violation and ends withviolations: N. Rules: no scope ending in.all; every tool whose scopes send, write or delete needshumanapproval; token lifetime at most 60 minutes; nothing insystem_promptthat looks like a credential (the word key or token followed by a value, or a long random string) -- secrets come from a vault;audit_logis true; any tool that runs code hassandboxtrue. - Run it against the before file and note the count.
- Write
/tmp/agent-after.json: the same agent with the minimum it needs to read mail and draft replies, sending gated by human approval. Run the checker until it reportsviolations: 0. - Write
/tmp/ai-use-policy.csvwith the headertool,approved,max_classification,trains_on_our_dataand rows for at least four tools -- your contracted assistant, a free public chatbot, a coding assistant and an internal model -- using the levels public, internal, confidential and restricted. - Write
/tmp/requests.csvwith the headerrequest,tool,classification,expectedholding eight requests a colleague might make, from pasting customer records into the free chatbot to sharing published marketing copy, each withallowordenyas the expected answer. Then write/tmp/policy_check.py, which decides each request from the policy file and prints one line per request, in file order, containing onlyallowordeny. Tune until every decision matches.
Verify
python3 /tmp/check_agent.py /tmp/agent-before.json | tail -1
python3 /tmp/check_agent.py /tmp/agent-after.json | tail -1
python3 - <<'PY'
import csv, json, subprocess
def count(path):
out = subprocess.run(['python3', '/tmp/check_agent.py', path], capture_output=True, text=True).stdout
return int(out.strip().splitlines()[-1].split()[-1])
before, after = count('/tmp/agent-before.json'), count('/tmp/agent-after.json')
print('violations before', before, '| after', after)
assert before >= 5, 'the starting manifest is not realistically bad'
assert after == 0, 'the trimmed agent still breaks a rule'
tools = json.load(open('/tmp/agent-after.json'))['tools']
risky = [t for t in tools if any(w in s for s in t['scopes'] for w in ('send', 'write', 'delete'))]
assert all(t['approval'] == 'human' for t in risky), 'a tool that can act has no human approval'
reqs = list(csv.DictReader(open('/tmp/requests.csv')))
got = [l.strip().lower() for l in subprocess.run(['python3', '/tmp/policy_check.py'], capture_output=True, text=True).stdout.splitlines() if l.strip()]
want = [r['expected'].strip().lower() for r in reqs]
assert len(reqs) >= 8 and got == want, 'policy decisions do not match the expected answers'
assert 'allow' in got and 'deny' in got, 'a policy that only allows or only denies decides nothing'
print('agent at zero violations; all', len(reqs), 'requests decided as expected')
PY
The first line must report at least five violations and the second 0. The assertion then re-checks the one rule that matters most -- every tool able to send, write or delete waits for a human -- and requires your usage policy to reach the expected decision on all eight requests, with at least one of each.
Notes
Count which rule removed the most risk. It is usually the approval gate: an agent with broad standing permissions turns any successful injection into a real action, and a human in front of the irreversible ones contains that even when everything else has failed. The usage policy is the other half -- the commonest AI incident is not an attack, it is somebody pasting data into the wrong tool.
This is an independent study companion for CompTIA Security+ SY0-801 and is not produced by or endorsed by CompTIA.