Put an AI agent on a short leash, then write the rules for the people using AI

applied · 70 min · Objective 2.6

Task

Take the manifest of a hypothetical email-triage agent -- the tools it can call, the scopes its tokens carry, whether its actions need approval -- start from the over-permissive version teams actually ship, and cut it down to least privilege with a checker that proves it. Then write the acceptable-use rules that say which data may go into which AI tools, and test them against requests a colleague might really make.

Steps

  1. Write /tmp/agent-before.json describing the agent as it is often first deployed: a list of tools, each with name, scopes (such as mail.read, mail.send, files.readwrite.all, db.write), approval (none or human), token_ttl_minutes and, for any tool that runs code, sandbox; plus top-level system_prompt and audit_log fields. Make it realistically bad: broad scopes, no approval, day-long tokens, a key pasted into the system prompt, logging off.
  2. Write /tmp/check_agent.py, which takes a manifest path, prints one line per violation and ends with violations: N. Rules: no scope ending in .all; every tool whose scopes send, write or delete needs human approval; token lifetime at most 60 minutes; nothing in system_prompt that looks like a credential (the word key or token followed by a value, or a long random string) -- secrets come from a vault; audit_log is true; any tool that runs code has sandbox true.
  3. Run it against the before file and note the count.
  4. Write /tmp/agent-after.json: the same agent with the minimum it needs to read mail and draft replies, sending gated by human approval. Run the checker until it reports violations: 0.
  5. Write /tmp/ai-use-policy.csv with the header tool,approved,max_classification,trains_on_our_data and rows for at least four tools -- your contracted assistant, a free public chatbot, a coding assistant and an internal model -- using the levels public, internal, confidential and restricted.
  6. Write /tmp/requests.csv with the header request,tool,classification,expected holding eight requests a colleague might make, from pasting customer records into the free chatbot to sharing published marketing copy, each with allow or deny as the expected answer. Then write /tmp/policy_check.py, which decides each request from the policy file and prints one line per request, in file order, containing only allow or deny. Tune until every decision matches.

Verify

python3 /tmp/check_agent.py /tmp/agent-before.json | tail -1
python3 /tmp/check_agent.py /tmp/agent-after.json | tail -1
python3 - <<'PY'
import csv, json, subprocess
def count(path):
    out = subprocess.run(['python3', '/tmp/check_agent.py', path], capture_output=True, text=True).stdout
    return int(out.strip().splitlines()[-1].split()[-1])
before, after = count('/tmp/agent-before.json'), count('/tmp/agent-after.json')
print('violations before', before, '| after', after)
assert before >= 5, 'the starting manifest is not realistically bad'
assert after == 0, 'the trimmed agent still breaks a rule'
tools = json.load(open('/tmp/agent-after.json'))['tools']
risky = [t for t in tools if any(w in s for s in t['scopes'] for w in ('send', 'write', 'delete'))]
assert all(t['approval'] == 'human' for t in risky), 'a tool that can act has no human approval'
reqs = list(csv.DictReader(open('/tmp/requests.csv')))
got = [l.strip().lower() for l in subprocess.run(['python3', '/tmp/policy_check.py'], capture_output=True, text=True).stdout.splitlines() if l.strip()]
want = [r['expected'].strip().lower() for r in reqs]
assert len(reqs) >= 8 and got == want, 'policy decisions do not match the expected answers'
assert 'allow' in got and 'deny' in got, 'a policy that only allows or only denies decides nothing'
print('agent at zero violations; all', len(reqs), 'requests decided as expected')
PY

The first line must report at least five violations and the second 0. The assertion then re-checks the one rule that matters most -- every tool able to send, write or delete waits for a human -- and requires your usage policy to reach the expected decision on all eight requests, with at least one of each.

Notes

Count which rule removed the most risk. It is usually the approval gate: an agent with broad standing permissions turns any successful injection into a real action, and a human in front of the irreversible ones contains that even when everything else has failed. The usage policy is the other half -- the commonest AI incident is not an attack, it is somebody pasting data into the wrong tool.

This is an independent study companion for CompTIA Security+ SY0-801 and is not produced by or endorsed by CompTIA.