AI threats and vulnerabilities, from the defender's side
This course teaches SY0-801, the Security+ exam that launches on or around 17 November 2026. If you are booked on SY0-701, which can be taken until 11 June 2027, use our SY0-701 course instead.
Objective 2.6 is new in SY0-801, and its verb is summarize. You are not asked to attack a model. You are asked to recognise what can go wrong when an organisation uses AI, name it correctly, and pick the control that limits the damage. This lesson is written exactly that way.
Why this matters
Every organisation now has AI in it, whether or not anyone approved it: a chat assistant in the browser, a summariser in the email client, a coding assistant in the developers' editor, a model scoring loan applications or flagging transactions. Each of those is a system that takes input, makes decisions and sometimes takes actions, and each one can be lied to, leaked through or trusted too much.
Two ideas carry most of this objective. First, a language model cannot reliably tell instructions from data -- anything it reads might be treated as a command. Second, a model is only as trustworthy as what it was trained on and what it is allowed to do. Nearly every question here is one of those two ideas wearing a scenario.
The lesson
Prompt injection and jailbreaking: what they are, how to spot them, what limits the damage
Prompt injection is getting a model to follow instructions that its owner did not give it. In direct injection the user types those instructions themselves. Indirect injection is the one that matters for defenders: the instructions are hidden in content the model is asked to process -- a web page it summarises, an email it triages, a document it reads, a code comment. The user who asked for a summary never sees the hidden text, but the model does.
Jailbreaking is the narrower case of persuading a model to ignore its own safety rules -- role-play framings, "ignore your previous instructions", hypotheticals -- so that it produces something it was built to refuse.
How it shows up:
- the assistant's output contains instructions, links or requests the user never asked for ("to finish, please confirm your password here");
- an AI tool takes an action -- sends a message, opens a URL, changes a record -- that nobody requested;
- logs show a document or page containing text addressed to "the assistant" or "the AI";
- outputs suddenly ignore the organisation's usage rules.
What limits the damage is not a better filter alone, because no filter recognises every phrasing. The controls that hold are architectural: treat model output as untrusted input to anything downstream; give the model's tools least privilege; require human approval before any consequential action; keep secrets out of prompts and system instructions, because anything the model can read it can be talked into repeating; and log every prompt, tool call and output so injection is investigable afterwards.
Poisoned training data, manipulated models and inputs shaped to slip past a classifier
These attack the model itself rather than one conversation.
- Data poisoning -- corrupting what a model learns from. Mislabelled or planted training examples teach a spam filter that a phishing template is legitimate, or plant a trigger that changes the model's behaviour only when a particular phrase appears.
- Model tampering -- altering a trained model or swapping it for a modified one, often through the supply chain: a downloaded pre-trained model, a shared model file, a compromised repository. It is software supply chain risk in a new container.
- Evasion -- inputs crafted so that a working model gets them wrong: malware adjusted until a machine-learning detector scores it benign, or an image altered so a classifier misreads it. The model is fine; the input is hostile.
The controls follow from that. Know where training data came from and who can change it. Treat model files like code: inventory them, take them from trusted sources, verify integrity with hashes or signatures, and control who can deploy them. Monitor a model's decisions for drift -- a fraud model whose approval rate quietly rises is telling you something. Test models against hostile input before trusting them, and for high-stakes decisions do not let a single model be the only control; keep a second signal or a human in the loop.
Data loss and privacy: what people paste into AI tools, and where it ends up
The most common AI incident is not an attack at all. It is an employee pasting customer records, source code or a contract into a public AI tool to get help with it. That data has now left the organisation, may be retained by the provider, may be used to improve the provider's models, and is outside every control the organisation has.
Models can also leak in the other direction. A model trained or fine-tuned on sensitive data can sometimes be coaxed into revealing that data -- reciting a record it learned, or confirming whether a specific person's data was in its training set. Anything put into a model's training data should be assumed extractable.
The controls are the familiar data-protection ones, aimed at a new channel:
- an acceptable use policy for AI that says which tools are approved and which data classifications may go into them;
- approved, contracted tools whose terms exclude training on your data and set retention limits, so people have a sanctioned alternative;
- DLP and web filtering that recognise AI sites and block or warn on sensitive data heading there;
- data minimisation for anything used to train or fine-tune, so the model never holds what it should not repeat.
Unapproved AI use is shadow AI -- the same problem as shadow IT, and the same answer: discover it, then offer something approved that is easier to use.
Bias, explainability, hallucination, and the ethical questions an organisation has to answer
Some AI risks are not breaches but bad decisions made at scale.
- Bias. A model trained on skewed or historical data reproduces the skew -- screening out candidates or declining customers along lines the organisation would never defend if a person did it. Because the model is consistent, so is the unfairness.
- Explainability. Many models cannot say why they decided something. That is a security problem when you must justify a decision to a customer, an auditor or a regulator, or work out why a detection model missed an attack.
- Hallucination. Models produce confident, fluent, wrong output -- invented citations, non-existent commands, fabricated policy. It becomes a security issue when output is trusted without checking: a coding assistant that suggests a software package which does not exist invites an attacker to register that name, so the next developer who trusts the suggestion installs the attacker's code.
The organisational answer is governance: an owner accountable for each AI system, a documented purpose, review before use in decisions that affect people, human oversight where the stakes are high, records good enough to explain a decision later, and attention to the laws on automated decision-making where the organisation operates. Domain 5 is where that governance lives; this objective is where you learn what it is protecting against.
AI integrations that can act: hijacked sessions, code execution and keeping an agent on a short leash
The risk grows sharply when an AI system can do things rather than just answer: read and send email, query databases, call APIs, run code. That system is called an agent, and it inherits every permission it is given.
What goes wrong:
- Excessive agency -- an agent with broad standing permissions turns any successful injection into real actions: data sent out, records changed, money moved.
- Hijacked sessions and stolen keys -- the API keys and tokens an AI integration uses are credentials. Leaked in code, logs or a prompt, they let an attacker use the integration directly.
- Unsafe code execution -- a tool that runs model-generated code, or passes model output into a shell or a database query, has handed control to whatever influenced the model.
Keeping an agent on a short leash uses controls you already know, applied deliberately: least privilege and narrowly scoped, short-lived tokens for every tool; sandboxing for anything that executes code; human approval for actions that move data or money or cannot be undone; rate limits; secrets held in a vault rather than in prompts or config; and a full audit log of every tool call. Automation in Domain 4 benefits from AI, and the same rule applies there: the more an automated system can do on its own, the more its permissions and logging matter.
Exam habits for 2.6
| Scenario described | Name it | First control |
|---|---|---|
| Assistant summarising a web page starts asking the user for credentials | Indirect prompt injection | Treat output as untrusted; restrict the tool; user awareness |
| Users get a model to ignore its content rules with a role-play | Jailbreaking | Layered guardrails, output monitoring, logging |
| Spam filter starts passing a phishing template after a retrain | Data poisoning | Training data provenance and change control |
| Detector scores slightly modified malware as clean | Evasion | Do not rely on one model; behavioural detection |
| Staff paste customer data into a free chatbot | Data loss / shadow AI | AI policy, approved tool, DLP |
| Hiring model rejects one group disproportionately | Bias | Governance, review, human oversight |
| Coding assistant recommends a package that does not exist | Hallucination | Verify before use; dependency controls |
| An AI agent emailed a file nobody asked it to send | Excessive agency | Least privilege, approval for actions, audit log |
What to take into the exam
- A language model cannot reliably separate instructions from data; indirect injection hides instructions in content the model reads.
- Contain AI with architecture -- least privilege, human approval, untrusted output, no secrets in prompts, full logging -- not with filters alone.
- Poisoning attacks training, tampering attacks the model, evasion attacks the input.
- The commonest AI incident is data pasted into an unapproved tool; the answer is policy, an approved alternative and DLP.
- Bias, explainability and hallucination are governance risks; agents with broad permissions turn any of the above into real actions.
Practise what you just read
1. An assistant asked to summarise a supplier's web page starts telling the user to confirm their password at a link. What is the most likely cause?
Select one
Show answer
C. Indirect injection hides instructions in content the model is asked to process, and a language model cannot reliably tell instructions from data. The user never saw the hidden text but the model did. Treating model output as untrusted and restricting what the assistant can do are the controls.
2. Which approach most reliably limits the damage from prompt injection against an AI tool that can take actions?
Select one
Show answer
D. No filter recognises every phrasing, and instructions in a prompt can be argued away. The controls that hold are architectural: least privilege for the model's tools, human approval before consequential actions, model output treated as untrusted, no secrets in prompts, and full logging.
3. An AI agent with standing permission to send email forwards a confidential file that nobody asked it to send. What is this, and what is the first control?
Select one
Show answer
A. An agent inherits every permission it is given, so broad standing permissions turn any successful injection into real actions. Narrowly scoped, short-lived tokens, human approval for actions that move data or cannot be undone, and an audit log of every tool call keep it on a short leash.
Hands-on labs
Part of the free CompTIA Security+ SY0-801 course — 47 lessons and 78 hands-on labs.
This is an independent study companion for CompTIA Security+ SY0-801 and is not produced by or endorsed by CompTIA.