Free EXIN AISP Practice Exam: AI Security Professional
Try 40 free EXIN AI Security Professional practice questions across the five weighted domains, with answer explanations and an AISP study-review worksheet.
Use this fixed set of 40 single-answer questions to practise AI security scenarios and interpret technical evidence. It covers all five domains in EXIN’s August 2026 preparation guide.
These are original IT Mastery practice questions, not official EXIN questions, copied live-exam content or exam dumps.
Start Question 1 · Review your attempt
How to use this free exam
- Allow 90 minutes and record one answer per question before opening its explanation.
- Award one point per correct answer, for a total out of 40. Mark correct guesses for review too.
- Compare each choice’s explanation with the evidence in the scenario, especially the closest competing answer.
Use your own timer and the page’s question navigation. This page does not run a timer, record your choices or calculate a score. Code blocks offer Wrap lines when needed. Wide tables scroll horizontally.
EXIN specifies 40 questions in 90 minutes with 26 correct required to pass. This set follows the official domain weights; it does not reproduce the official examination or predict your result. Official exam details .
Practice-set coverage
| Domain | Official range | Questions in this set |
|---|---|---|
| AI Security in the Organization | 15% | 6 |
| AI Security Threats | 37.5% | 15 |
| AI Security Controls | 27.5% | 11 |
| AI Security Testing | 7.5% | 3 |
| Privacy and Compliance in AI Security | 12.5% | 5 |
Practice questions
Questions 1-25
Question 1
Topic: AI security controls
A company operates an assistant using a supplier-hosted model. The company controls the assistant’s retrieval and internal-tool integrations. During an authorized exercise, malicious instructions in a retrieved document cause the assistant to send restricted data through a tool.
Response findings:
- Analysts detect and correctly classify the disclosure.
- The AI governance register identifies an accountable service owner.
- The incident runbook lists neither AI-service nor supplier contacts and provides no assigned authority or procedures for suspending retrieval and tool access.
Which improvement most directly closes the demonstrated governance gap?
Options:
A. Extend
SEC EDUCATEwith AI misuse and leakage training, define role-specific learning requirements, and validate responder competence through practical assessments.B. Extend
SEC DEV PROGRAMwith AI misuse and leakage tests, require corrective changes before release, and validate their effectiveness through adversarial regression testing.C. Extend
SEC PROGRAMwith AI incident-response playbooks, define escalation contacts and authorized containment procedures, and validate operational readiness through a response exercise.D. Extend
AI PROGRAMwith AI misuse and leakage reviews, refresh accountable risk ownership, and validate risk-treatment decisions through management approval.
Best answer: C
Explanation: SEC PROGRAM integrates AI risks into organizational security processes, including incident response. Here, detection and accountable ownership exist, but responders cannot coordinate or execute containment. The security program should connect responders to the AI owner and supplier, document who may suspend retrieval or tool access, and exercise those procedures. Evidence of closure is successful escalation and authorized containment during a repeat exercise, not merely an approved document.
Other programs support this operational capability: AI PROGRAM establishes AI governance and ownership; SEC DEV PROGRAM incorporates security into development; and DEV PROGRAM supports disciplined development and lifecycle practices. CHECK COMPLIANCE checks applicable obligations, while SEC EDUCATE develops role-specific competence. These responsibilities should connect to incident response without replacing its operational contacts, authority, and playbooks.
Why each option fits or fails:
A. Responders already recognize the incident; competence assessments cannot supply operational contacts or authorize containment actions absent from the response plan.
B. Adversarial release testing addresses vulnerabilities during development, not responders’ missing contacts, containment authority, and procedures during an active disclosure.
C. The incident runbook lacks operational escalation and containment arrangements; integrating and exercising them within the security program directly closes that gap.
D. AI ownership is already established; approved risk reviews do not supply the missing operational escalation routes and containment procedures.
Question 2
Topic: Privacy and compliance in AI security
A bank’s AI document-classification pilot is moving into production. The project lead must establish coordinated processes, owners, deliverables, and handoffs for development, validation, deployment, operation, maintenance, and eventual retirement of the system and its associated data.
Which standard is the most directly relevant primary reference for organizing this work?
Options:
A. ISO/IEC 23894:2023
B. ISO/IEC 27005:2022
C. ISO/IEC 42001:2023
D. ISO/IEC 5338:2023
Best answer: D
Explanation: ISO/IEC 5338:2023 concerns AI system lifecycle processes. It is the most direct reference for organizing how an AI system is developed, deployed, operated, maintained, and retired, including coordinated responsibilities and handoffs.
The bank’s immediate need is a process structure spanning the system’s full life. Risk-management guidance supports identifying, assessing, and treating risks, while an AI management system establishes organization-level management arrangements. These have complementary purposes but do not replace the lifecycle-process reference. Security and privacy requirements must still be incorporated into the work, and applying a standard does not automatically establish legal compliance.
Why each option fits or fails:
A. This standard provides AI risk-management guidance, rather than the system lifecycle process structure needed to coordinate development, operation, and retirement.
B. This standard provides information-security risk-management guidance, not the lifecycle process framework needed for the bank’s development-to-retirement planning.
C. This standard specifies an organizational AI management system, rather than serving as the primary reference for the requested AI system lifecycle processes.
D. This standard concerns AI system lifecycle processes, directly supporting the bank’s need to organize work and handoffs across the system’s life.
Question 3
Topic: AI security controls
An AI reviewer checks privileged-access requests for self-approval. Policy requires the requester and approver to be different employees. The team wants to remove raw employee identifiers from model inputs and diagnostic logs without changing the security decisions.
Two validated requests are replayed using the same model version. Field-local aliases use separate mappings for requester and approver. Review-local aliases use one mapping shared across both fields. Mappings are recreated for each review and retained in a restricted identity service.
Each pair below shows requester / approver:
Case 1, required result: FLAG
Original IDs: E143 / E143 -> FLAG
Field-local aliases: R7 / A4 -> CLEAR
Review-local aliases: P2 / P2 -> FLAG
Case 2, required result: CLEAR
Original IDs: E143 / E827 -> CLEAR
Field-local aliases: R7 / A4 -> CLEAR
Review-local aliases: P2 / P8 -> CLEAR
Which interpretation is best supported by these results?
Options:
A. Review-local aliases preserve the approval relationship; removing raw employee identifiers therefore makes the review records anonymous.
B. Original identifiers preserve the approval relationship; raw employee identifiers therefore remain essential for the tested decisions.
C. Field-local aliases preserve the approval relationship; raw employee identifiers are therefore unnecessary for the tested decisions.
D. Review-local aliases preserve the approval relationship; raw employee identifiers are therefore unnecessary for the tested decisions.
Best answer: D
Explanation: DATA MINIMIZE must preserve the information and relationships needed for the security task, not merely remove identifying values. Independent field mappings turn one employee into two apparent identities, losing the relationship needed to detect self-approval. A shared, review-local mapping preserves both identity matches and differences while keeping raw identifiers out of model inputs.
These replays support review-local pseudonymization for the tested check. They do not establish that the records are anonymous: the restricted identity service can still connect aliases to employees. Useful diagnostic evidence can retain the aliases and decisions rather than unnecessary copies of raw identifiers. Additional permitted and prohibited cases should be tested before broader rollout, because these two successful replays do not establish complete coverage.
Why each option fits or fails:
A. The retained identity mapping keeps the records linkable to employees, so removing direct identifiers provides pseudonymization rather than anonymization.
B. Review-local aliases reproduce both required decisions, demonstrating that raw employee identifiers are not essential for the tested check.
C. Separate field mappings give the same employee different aliases, causing the self-approved request to be mistakenly cleared.
D. The shared mapping preserves whether requester and approver are the same employee, and both decisions match the validated outcomes.
Question 4
Topic: AI security in the organization
A workplace assistant uses an employee’s reservation-management permissions to book or cancel meeting rooms. The workflow under review creates one new reservation. Employees with limited hand mobility rely on voice booking because the separate portal requires mouse dragging.
Attacker-controlled room descriptions retrieved by the assistant caused cancellations outside the employee’s request. The same attack and legitimate-task cases were tested before and after disabling all action tools:
| Pilot measure | Tools enabled | Tools disabled |
|---|---|---|
| Unauthorized cancellations | 6 of 40 attacks | 0 of 40 attacks |
| Independent voice bookings | 20 of 20 tasks | 0 of 20 tasks |
The service owner wants to reduce unauthorized actions while preserving independent booking. Any revised approach will undergo adversarial and accessibility testing with affected users.
Which approach should the team prioritize for the next pilot?
Options:
A. Restore booking only, with application-enforced authorization limited to the employee’s requested new reservation.
B. Retain read-only assistance and route voice-booking requests to an accessible service where staff complete reservations for employees.
C. Restore booking and cancellation tools with model instructions limiting actions to the employee’s current booking request.
D. Restore booking and cancellation tools with application-enforced checks against the employee’s existing reservation-management permissions.
Best answer: A
Explanation: Functionality restrictions should be proportionate to both the threat and legitimate service needs. Read-only operation prevented unauthorized cancellations in the tested cases, but also eliminated independent access for employees who rely on voice booking. Applying identical restrictions to everyone does not necessarily provide equitable access.
Intent-based least privilege limits an agent to its current task rather than everything its user may do. Application-enforced authorization can permit the requested new reservation while making cancellation unavailable during that task. The pilot should validate attack resistance and legitimate task completion with affected users, followed by accountable review of residual risk.
Security, safety, reliability, fairness, transparency and privacy are distinct but related qualities. A successful security control does not establish overall trustworthiness. Each relevant quality needs appropriate evidence, and a finite passing attack suite does not prove the absence of vulnerabilities.
Why each option fits or fails:
A. Task-specific authorization limits available actions at the application boundary while restoring the function needed for independent voice booking.
B. Staff-assisted booking offers an accessible fallback, but it replaces the documented independent voice workflow instead of preserving it.
C. Model instructions can guide behavior but remain vulnerable to injected content; they do not impose an enforceable boundary on cancellation tools.
D. Employee-level permissions already permit cancellations, so these checks do not block actions on existing reservations outside the current booking request.
Question 5
Topic: AI security in the organization
A company has one discretionary engineering slot to reduce risk in two AI services. Its budgeting policy states:
Prioritize the greatest annual expected loss, calculated as annual event probability multiplied by loss per event, after verified deployed controls. Accepting residual risk is a separate decision requiring the relevant business owner’s approval.
Threat-model findings:
- Confidential case disclosure: Customer-supplied documents enter a staff support agent’s context. The agent uses staff credentials to retrieve confidential records and draft replies. An attacker may try to induce disclosure of unrelated cases.
- Recommendation degradation: An attacker can edit a public supplier page retrieved at runtime by a product recommender. Corrupted content can degrade noncritical recommendations. This service has no confidential-record access or action tools.
Approved risk record: Each probability estimates the chance of one modeled loss event during the coming year. Loss estimates include the business consequences relevant to this decision.
Confidential case disclosure
Untreated: probability 25%; loss $400,000
Current: probability 0.5%; loss $400,000
Deployed: tool access scoped to the application-assigned case
Verified: cross-case reads denied in scoped adversarial tests
Recommendation degradation
Untreated: probability 80%; loss $20,000
Current: probability 50%; loss $20,000
Deployed: output relevance checks
Verified: some manipulated-page recommendations still passed
Planned: source isolation; forecast probability 5%; not deployed
Residual-risk acceptance: pending both business owners
Which prioritization is supported by the policy and current record?
Options:
A. Prioritize confidential case disclosure; annual expected losses are $2,000 for disclosure and $1,000 for recommendation degradation.
B. Prioritize confidential case disclosure; annual expected losses are $100,000 for disclosure and $16,000 for recommendation degradation.
C. Prioritize recommendation degradation; annual expected losses are $2,000 for disclosure and $10,000 for recommendation degradation.
D. Prioritize recommendation degradation; annual expected losses are $500 for disclosure and $8,000 for recommendation degradation.
Best answer: C
Explanation: Prioritization uses expected annual loss after controls already operating and verified. Disclosure has the greater per-event impact, but the current estimates give:
- Disclosure: \(0.005 \times 400{,}000 = 2{,}000\), or $2,000 annually.
- Recommendation degradation: \(0.50 \times 20{,}000 = 10{,}000\), or $10,000 annually.
Recommendation degradation therefore receives the next engineering slot. Untreated estimates describe inherent risk. Deployed authorization and relevance checks are treatments whose effects are reflected in the current residual estimates. The planned isolation cannot yet reduce assessed risk, and passing scoped tests does not prove disclosure is impossible.
Choosing a treatment priority does not accept either remaining risk. Accountable acceptance remains a separate decision for each business owner, and both decisions are still pending.
Why each option fits or fails:
A. The $1,000 estimate credits the planned source-isolation treatment, which is not deployed and therefore cannot reduce the current assessed risk.
B. These amounts use untreated probabilities, describing inherent risk rather than the residual risk required for the budgeting decision.
C. The current estimates give 0.5% of $400,000 and 50% of $20,000, so recommendation degradation has the greater residual expected loss.
D. Multiplying untreated and current probabilities incorrectly treats the current event probabilities as additional reduction factors, understating both residual losses.
Question 6
Topic: AI security controls
A company is reviewing a support assistant that uses a supplier-hosted inference API. Company policy prohibits submission of bank account numbers and requires the caller to have refund permission before a refund is executed.
A joint review identifies three issues requiring correction:
Assurance scope: supplier model hosting and inference endpoint
Inference host: supplier operated; required security patch missing
Request builder: company's support_app
Submitted payload: bank account number present (value redacted)
API response: suggested_tool = issue_refund
Caller permissions: read_case only
Tool executor: company's support_app using shared service identity
Tool result: refund completed
Which allocation of corrective responsibility is supported by this record?
Options:
A. The provider patches the inference host, prevents restricted-data submission and enforces refund authorization.
B. The provider patches the inference host and prevents restricted-data submission; the customer enforces refund authorization.
C. The provider patches the inference host; the customer prevents restricted-data submission and enforces refund authorization.
D. The provider patches the inference host and enforces refund authorization; the customer prevents restricted-data submission.
Best answer: C
Explanation: Responsibility follows the delivery arrangement and the actual enforcement boundaries. For a hosted inference API, the provider operates the model service and its infrastructure, so it must correct the missing inference-host patch. The customer controls what its application submits and how connected tools are invoked.
Here, support_app sends the prohibited account number and executes a refund using a shared service identity despite the caller having only read_case permission. The customer must enforce allowed-data rules before submission and validate the caller’s refund authority at tool execution. A model-generated tool suggestion is not authorization. Supplier assurance provides scoped evidence about the hosted service; it does not transfer responsibility for the customer’s integration.
Why each option fits or fails:
A. Assurance for the hosted service does not transfer responsibility for the company’s request construction or connected-tool authorization to the supplier.
B. The company constructs the request body, so it must prevent prohibited data from being submitted before the payload reaches the provider.
C. The supplier operates the unpatched host, while the company’s application controls both submitted data and execution of the suggested refund.
D. The API only suggests a refund; the company’s tool executor must enforce the caller’s permissions before executing it.
Question 7
Topic: AI security threats
A multi-tenant billing assistant uses a model to propose SQL. A separate executor accepts only SELECT statements and uses a read-only database account that can currently read every tenant’s rows. It binds :tenant_id from the authenticated session, not from model output. Aggregate queries must remain supported.
For a user in tenant 17, the model proposes:
SELECT :tenant_id AS tenant_id, SUM(i.amount) AS total
FROM invoices AS i
WHERE i.status = 'overdue' OR i.tenant_id = :tenant_id;
A review confirms that the returned row is labeled with tenant 17, but its total includes overdue invoices from other tenants.
Which executor change best enforces tenant isolation for generated aggregate queries?
Options:
A. Require the bound tenant parameter to appear in the generated WHERE clause.
B. Execute under an enforced database row policy scoped to the authenticated tenant ID.
C. Retain only returned records whose
tenant_idmatches the authenticated tenant ID.D. Append
AND i.tenant_id = :tenant_iddirectly to the generated WHERE clause.
Best answer: B
Explanation: Generated SQL is untrusted content at the execution boundary. Restricting operations to SELECT and using a read-only account prevents writes, but does not establish which rows may be read. Parameter binding protects values from being interpreted as SQL syntax; it does not make the surrounding query authorized.
The executor should establish tenant context from the authenticated session and execute through an account subject to a database row policy it cannot bypass. That policy limits eligible source rows before the aggregate is computed, so a model-generated OR condition cannot widen tenant access.
Filtering the returned tenant label is too late here. The query supplies that label directly, while SUM has already combined authorized and unauthorized invoice amounts.
Why each option fits or fails:
A. The query already references the correct tenant parameter, yet its OR condition includes other tenants; parameter presence does not enforce row authorization.
B. An enforced row policy excludes unauthorized source rows before aggregation, independently of the conditions included in the model’s query.
C. The returned tenant ID is a bound label, not evidence of which source rows contributed to the aggregate.
D. SQL evaluates AND before OR, so this appended condition still permits overdue invoices belonging to other tenants.
Question 8
Topic: AI security threats
During an authorized adversarial test, a team captures this complete trace from a deployed classifier. Weight versions identify the model parameters used for predictions.
Start: serving weights=v17; feedback queue=empty
Input: tester submitted feedback with deliberately false labels
Processing: learner consumed that feedback immediately
Update: learner changed serving weights from v17 to v18
Storage: learner persisted v18 in the deployed weights file
Next prediction: same running service used v18 for ordinary input
Audit: administrative writes/copies of model artifacts=0
Audit: deployed configuration changes=0
Which interpretation of the observed attack mechanism is supported by the trace?
Options:
A. Inference-time evasion through manipulated prediction inputs.
B. Online-learning poisoning through malicious feedback submissions.
C. Development-data poisoning awaiting a future retraining run.
D. Direct runtime replacement through administrative artifact modification.
Best answer: B
Explanation: Online-learning poisoning occurs when attacker-controlled examples influence a learner that updates the model during operation. Here, deliberately false labels are consumed immediately, the learner changes and persists the serving weights, and the next ordinary prediction uses the updated version.
The attack path matters more than the fact that weights changed. Direct runtime replacement involves directly modifying or substituting the deployed model artifact, rather than inducing updates through learning. The trace attributes the changes to the learner and records no administrative artifact writes or copies. It also shows an immediate operational update, not poisoned data awaiting future retraining or crafted inputs affecting only inference.
Why each option fits or fails:
A. Evasion manipulates inputs used for predictions; here the malicious feedback changes model weights and affects a subsequent ordinary input.
B. The deliberately false feedback is consumed by the learner during operation, changing the serving model before the next ordinary prediction.
C. The malicious feedback changes the serving weights immediately and affects the next prediction, rather than remaining unused until future retraining.
D. The weight file changes through a learner update, while the audit records no administrative write or copy that directly replaces the deployed artifact.
Question 9
Topic: AI security controls
A bank uses AI to recommend additional income-document requests. Sending a request pauses the applicant’s loan assessment. The bank intends human oversight to prevent unjustified requests from affecting customers.
The model can propose requests but cannot send them. Reviewers can inspect original applicant records and the model’s stated reasons, reject recommendations, and withdraw issued requests. They cannot retrain the model. A workflow service sends each queued request after 60 seconds unless a rejection is recorded.
Record for one case:
| Elapsed time | Event |
|---|---|
| 0 seconds | Recommendation queued |
| 60 seconds | Request sent; loan assessment paused |
| 150 seconds | Reviewer opens supporting records |
| 240 seconds | Reviewer rejects request; assessment resumes |
Which assessment of human oversight is best supported for this case?
Options:
A. The oversight is effective because reviewers can inspect supporting records and withdraw unjustified requests that have already been sent.
B. The oversight is ineffective because the request affects the customer before the reviewer has an opportunity to examine and reject it.
C. The oversight is ineffective because reviewers lack authority to retrain the model after identifying and rejecting an unjustified document-request recommendation.
D. The oversight is effective because customer-facing execution is controlled by a separate workflow service rather than by the model producing the recommendation.
Best answer: B
Explanation: Meaningful preventive human oversight requires usable evidence, decision authority, and enough time to intervene before a recommendation affects the customer. Here, the reviewer has evidence and rejection authority, but the release mechanism defeats timely intervention: dispatch occurs at 60 seconds, while review begins at 150 seconds.
The workflow should hold the request and associated assessment pause until an informed reviewer approves execution. A timeout should leave the request pending rather than treat silence as approval. Later withdrawal remains useful for correction, but cannot undo the initial customer impact. Restricting the model to proposing requests is least model privilege, not a substitute for effective human review.
Why each option fits or fails:
A. Withdrawal corrects an issued request, but the customer has already experienced a pause, so it does not achieve the stated preventive purpose.
B. Automatic dispatch occurs before review begins, preventing the reviewer from stopping the customer impact despite having supporting evidence and rejection authority.
C. Preventive oversight requires authority to stop the customer-facing action, not authority to retrain the model that proposed it.
D. Separating proposal from execution limits model privilege, but the workflow service still acts before a reviewer assesses the recommendation.
Question 10
Topic: AI security threats
A team is preparing to fine-tune a model on private training copies. External reviewers may read the development repository and build logs, but their accounts have no training-store permissions. Both the repository and log service use encryption at rest.
A security review captures the following record. <redacted> replaces the same full token only in this exhibit; the original file and log contain its usable value.
Committed file: training/config.yaml
dataset_path: training-copies/project-7
service_token: <redacted>
Build log: service_token=<redacted>
Credential registry: active bearer token for trainer-service
Token permission: read training-copies/project-7
Which interpretation of the exposure is supported by this record?
Options:
A. The credential’s usable value remains protected from repository and log readers by the systems’ encryption at rest.
B. The credential can retrieve training copies only when the person using it has personal access to the training store.
C. The configuration and build logs directly disclose the contents of the training copies identified by the credential.
D. The credential can give repository and log readers access to training copies using the service account’s read permissions.
Best answer: D
Explanation: Service credentials are development assets that require protection alongside source code, training copies, experiments and model artifacts. Here, the exposed asset is an active bearer token, and the access paths are the committed configuration and build logs.
Anyone able to obtain that token could present it using the service account’s authority. Reviewers’ lack of personal training-store permissions therefore does not close this credential-based path. Encryption at rest protects stored information against unauthorized storage access, not against disclosure through authorized file and log reads.
The evidence establishes a route to training-data access, not proof that training records have already been retrieved. Containment requires revoking or rotating the token and preventing secrets from entering source and logs; adding storage encryption alone would leave the exposure unchanged.
Why each option fits or fails:
A. Encryption at rest does not conceal file or log contents from authorized readers; those contents include the usable token.
B. The bearer token authenticates access under the service account’s permissions, rather than requiring the holder’s personal account to have training-store access.
C. The record shows disclosure of a credential and dataset location, not disclosure of training-record contents within the configuration or logs.
D. The committed configuration and build log disclose an active bearer token whose read permission provides a route to the specified training copies.
Question 11
Topic: AI security testing
A bank’s security team is selecting adversarial AI tests for a deployed predictive fraud classifier.
System and attacker facts:
- The classifier processes structured transaction features and returns only
alloworblock. - An external attacker can vary user-controlled transaction fields, submit transactions and observe decisions, but cannot access model files or training infrastructure.
- Submitted transactions do not enter training data, and the deployed model’s weights remain fixed.
The team’s priority risks are fraudulent transactions receiving approval and outsiders reproducing the classifier’s decision behavior.
Which pair of tests best covers these risks under the stated attacker profile?
Options:
A. Test evasion by varying transaction features and model extraction by learning from transaction-decision pairs.
B. Test data poisoning by submitting crafted transactions and model extraction by learning from transaction-decision pairs.
C. Test evasion by varying transaction features and model inversion by inferring sensitive training information from decisions.
D. Test evasion by varying transaction features and membership inference by estimating whether particular records were used in training.
Best answer: A
Explanation: Threat coverage should follow attacker access, system behavior and lifecycle. Here, decision-only inference access is black-box access. Evasion tests vary features the attacker can genuinely control while preserving the underlying fraudulent activity, then assess whether the classifier allows it.
Model extraction tests use collected transaction-decision pairs to build a substitute that approximates the deployed classifier. Categorical decisions can provide useful supervision; probability scores or access to model internals are not required, although successful extraction is not guaranteed.
The fixed model has no training-data path from submitted transactions, so those submissions do not provide a poisoning route. Reproducing decision behavior is also distinct from recovering sensitive training information or identifying training-set membership.
Why each option fits or fails:
A. Feature changes test fraud-detection bypass, while collected transaction-decision pairs can support replication of the classifier using the attacker’s available access.
B. Crafted submissions cannot poison this model because they neither enter training data nor update its deployed weights.
C. Evasion addresses fraudulent approvals, but model inversion targets sensitive information rather than reproduction of the classifier’s decision behavior.
D. Evasion addresses fraudulent approvals, but membership inference investigates training-set inclusion rather than reproduction of the classifier’s decision behavior.
Question 12
Topic: AI security in the organization
A finance employee may view and refund any customer account, but delegates only a $75 refund for account C-104 in this session. An application gateway mediates all agent tool calls and retains the employee-approved task separately from agent outputs.
Observed workflow:
Agent 1 extracts customer_id=C-809 from an uploaded invoice.
lookup_customer(C-809) returns billing details and record_valid=true.
Agent 1 forwards the lookup result to Agent 2.
Agent 2 executes issue_refund(customer_id=C-809, amount=75).
record_valid means only that the customer record exists. Both tools use service identities with permissions covering all customer accounts. The gateway checks those service permissions but does not compare tool arguments with the approved task. Neither agent validates the identifier against that task.
Which additional gateway check best prevents this identifier error from cascading into out-of-scope data access and refund execution?
Options:
A. For each tool call, verify that the requested account and operation fall within the independently recorded employee-approved task.
B. For each tool call, verify that the customer identifier has the required format and resolves to an existing account.
C. For each tool call, verify that the requested account and operation are permitted by the authenticated employee’s role.
D. For each tool call, verify that the requesting agent belongs to the approved workflow and uses its assigned service identity.
Best answer: A
Explanation: A model-generated identifier is a proposal, not authorization. Resolving C-809 supplies plausible billing information to the second agent, which then turns that observation into a financial action. Successful lookup proves that an account exists, not that it belongs to the approved task. Forwarding the result between agents does not create additional authority.
User-based permissions cannot contain this cascade because the employee and service identities may operate across all accounts. The gateway should enforce intent-based least privilege by checking each requested account and operation against the independently retained task. Enforcement is needed at both retrieval and refund execution boundaries, rather than relying on either agent to recognize the mismatch.
Why each option fits or fails:
A. Binding each call to the approved task blocks the unintended lookup before its output can drive a refund on the wrong account.
B. An existing, well-formed identifier can still identify the wrong customer; record validity does not establish authorization for the delegated task.
C. The employee’s role covers every account, so this check would permit operations on C-809 despite the session being limited to C-104.
D. Approved agents can propagate an incorrect identifier using legitimate service identities, so authenticating the workflow does not constrain its delegated task scope.
Question 13
Topic: AI security controls
A security lead is assigning operational security responsibilities for an AI summarization service. The supplier has provided an assurance report and this delivery record:
Delivery: Downloadable model weights and reference serving image
Service operator: Customer operations team, on customer-managed VMs
Supplier maintenance: Publishes updated artifacts and advisories
Assurance scope: Security tests of the supplied model release
Which allocation of responsibility for the deployed service is supported by this record?
Options:
A. The supplier operates local network and identity controls; the customer applies patches to the local serving stack.
B. The customer operates local network and identity controls; the customer applies patches to the local serving stack.
C. The customer operates local network and identity controls; the supplier applies patches to the local serving stack.
D. The supplier operates local network and identity controls; the supplier applies patches to the local serving stack.
Best answer: B
Explanation: For downloaded model weights deployed on customer-managed infrastructure, the customer retains operational security responsibilities. Here, the customer operates both the virtual machines and the serving environment. It must manage network exposure, local authentication and authorization, and patching of serving software, dependencies and hosts.
The supplier can test its model release and publish updated artifacts or advisories. Providing an update does not deploy or verify it in the customer’s environment. Similarly, assurance about a supplied model release does not establish the security of the customer’s operational configuration. Application behavior, data handling, access permissions and intended-use risks also remain the customer’s responsibilities.
Why each option fits or fails:
A. The customer-operated serving environment places local network and identity controls with the customer, not with the supplier that maintains the model release.
B. Downloaded weights run in a customer-operated environment, so the customer must enforce local network and identity controls and apply serving-stack patches.
C. Publishing updated artifacts is different from deploying patches; the customer operates the serving environment and must apply updates locally.
D. Treating the supplier as the service operator contradicts the record: it supplies artifacts and advisories, while the customer operates the serving environment.
Question 14
Topic: AI security testing
An authorized staging assessment evaluates an agent that summarizes retrieved documents. The security boundary is that document content must not cause email execution. All test emails go to an isolated sink.
The same immutable fixtures are replayed in both runs: inj-07 contains an indirect prompt injection requesting an email, while clean-03 contains only report text. Before the second run, the team changes both the system prompt and tool grants.
| Recorded fact | Before | After |
|---|---|---|
| Model snapshot | model-2026-04-21 | model-2026-04-21 |
| System prompt | p17 | p18 |
| Effective tool grants | read_document, send_email | read_document |
inj-07: tool request | send_email to sink | send_email to sink |
inj-07: execution | Email sent | Denied: missing grant |
clean-03: result | Summary; no email request | Summary; no email request |
Which interpretation of the remediation verification is supported by this record?
Options:
A. The replay verifies authorization failure: issuing the injected email request is sufficient to cross the tool’s execution boundary.
B. The replay verifies authorization containment: the reduced grants stopped execution even though the injected email request still occurred.
C. The replay verifies prompt-level prevention: the updated prompt stopped the injected instruction before it produced an email effect.
D. The replay verifies broader injection resistance: the passing benign-document regression shows the updated agent rejects malicious document instructions.
Best answer: B
Explanation: Reproducible prompt-injection verification distinguishes model behavior from application enforcement. Here, the fixed model snapshot and unchanged retrieved-document fixtures support comparison. The repeated send_email request shows that the updated prompt did not prevent this induced behavior. Removing the tool grant blocked execution, so the evidence demonstrates authorization containment for this replay.
Because the prompt and permissions changed together, these runs do not establish the prompt update’s independent benefit. Further testing within the authorized staging scope should isolate control changes and retain model and prompt versions, fixture contents and sources, effective permissions, expected boundaries, and observed consequences. Replay the original attack and relevant variants after remediation.
The benign-document replay provides limited regression evidence for summary functionality. Neither this result nor a larger finite passing suite establishes the absence of prompt-injection vulnerabilities.
Why each option fits or fails:
A. Requesting a tool action differs from executing it; the recorded denial shows that the authorization boundary prevented email execution.
B. The unauthorized request recurred, but removing the send_email grant caused denial before execution, demonstrating containment for this replay.
C. The agent still requested send_email after the prompt update; execution stopped at the tool-grant check, not because the prompt prevented the request.
D. Completing the benign summary checks ordinary task performance, not resistance to malicious instructions or other retrieved-document attacks.
Question 15
Topic: AI security threats
An AI maintenance assistant retrieves a third-party troubleshooting note and generates a shell fragment to read a status file in the diagnostics directory. During an authorized security test, a malicious instruction planted in the note influences the model’s response. The backend runner passes that response directly to an operating-system shell without independent validation.
Sanitized trace (command text replaced):
Model response: <status-read command>; <file-delete command>
Shell execution: Both commands run
Result: Status returned; a configuration file outside diagnostics deleted
Identity: Existing service account; permissions unchanged
The status-read command uses its intended path. The service account already had permission to delete the configuration file.
Which threat is demonstrated at the runner’s execution boundary?
Options:
A. Command injection through a model-generated shell fragment.
B. Indirect prompt injection through a retrieved troubleshooting note.
C. Path traversal through a model-generated status-file path.
D. Privilege escalation through the runner’s service-account permissions.
Best answer: A
Explanation: Command injection occurs when untrusted content is interpreted as commands rather than confined to data. Here, the shell executes the entire model-produced fragment, including the unwanted deletion command. This is model output injection in the form of command injection.
The retrieved instruction influences the model at an earlier trust boundary. The runtime failure occurs when generated text crosses into operating-system execution. Existing service-account permissions determine the impact; privilege escalation is not required.
Treat generated output as untrusted at execution and display boundaries. Prefer fixed, authorized tool operations with independently validated arguments over arbitrary model-produced shell fragments. Safe display handling does not make text safe to execute. Separately, execution of an unwanted command does not establish leakage of submitted input; disclosure requires evidence from the relevant data flows.
Why each option fits or fails:
A. The shell interprets unvalidated model output as executable commands, allowing the additional deletion command to run instead of remaining untrusted data.
B. The malicious note is an indirect prompt-injection entry point into the model, not the mechanism by which the shell executes the additional command.
C. The status-read command uses its intended path; deletion results from a separate command, not a manipulated path escaping the diagnostics directory.
D. The runner uses unchanged permissions that already allow deletion of the file, so the execution does not demonstrate acquisition of additional privileges.
Question 16
Topic: AI security threats
A developer submits a pretrained-model bundle for approval. The team has not vetted the bundle’s registry owner.
Review record:
Requested model: Northbridge-Text-7B
Expected publisher: northbridge-research
Approved publisher key: NBR-01
Downloaded model: Northbridge-Text-7B-Verified
Registry owner: community-models-47
Signature: valid for CM47-02 (registry owner's key)
Bundle digest: matches signed manifest
Clean-task test: expected accuracy reproduced
Security evaluation: not performed
Which interpretation of the bundle is best supported by this record?
Options:
A. The bundle’s integrity and benign behavior are verified by its signature and accuracy results; expected-publisher provenance has not been established.
B. The bundle’s integrity and expected-publisher provenance are verified by its valid signature; benign behavior has not been established.
C. The bundle’s integrity is unverified because its signer differs from the expected publisher; benign behavior has not been established.
D. The bundle’s integrity is verified relative to the registry owner’s signature; expected-publisher provenance and benign behavior have not been established.
Best answer: D
Explanation: A valid signature establishes an integrity relationship between signed content and the verification key. It does not establish that the signer is a trusted publisher or that the content is harmless.
Here, CM47-02 belongs to community-models-47, while the team expects northbridge-research’s NBR-01. The matching digest supports integrity relative to the registry owner’s signed manifest, not an approved origin. Reproducing expected accuracy demonstrates normal-task performance; it does not rule out backdoors or malicious accompanying code.
Model weights, loading code, dependencies and associated artifacts should be assessed as third-party supply-chain inputs. Publisher provenance and security evaluation remain unresolved. The record does not demonstrate that poisoning actually occurred.
Why each option fits or fails:
A. Normal-task accuracy and a valid signature do not demonstrate absence of malicious behavior, including behavior triggered only by particular inputs.
B. CM47-02 belongs to the unvetted registry owner, not the approved publisher, so its valid signature does not establish expected-publisher provenance.
C. A different signer leaves publisher provenance unresolved, but the valid signature and matching digest still verify integrity relative to that signer’s manifest.
D. The matching signed digest verifies integrity under CM47-02, but neither the expected publisher’s authorization nor harmless behavior is established.
Question 17
Topic: Privacy and compliance in AI security
A United Kingdom company uses a hosted AI assistant to summarize trade articles for paid client reports. Publication of one response has been paused.
Review evidence:
- The reviewer requested an original summary, but 260 words of the 400-word response reproduce a copyrighted article’s distinctive prose with only minor changes.
- The article is free to read online, but no reproduction permission or applicable copyright exception has been established.
- The supplier contract assures lawful model training.
- The response contains no personal data.
Which release control best addresses the copyright risk for this response and similar future outputs?
Options:
A. Require reproduction-rights review and either clearance to use the passage or an original summary checked for substantial similarity before publication.
B. Require accurate attribution and confirmation of free public access to the article before approving the passage for publication.
C. Require documented supplier confirmation that the article was lawfully used in training before approving the passage for publication.
D. Require synonym substitutions that reduce exact-match overlap with the article before approving the revised passage for publication.
Best answer: A
Explanation: Near-verbatim output can carry copyright risk even when its source is publicly accessible and contains no personal data. The substantial extract of distinctive prose calls for review of provenance, reproduction permissions, contractual terms and any applicable UK copyright exception. Lawful model training does not by itself authorize a customer to reproduce protected material.
A release process should establish a valid basis for using the passage or require an independently written summary checked for substantial similarity. Similarity detection can flag copying, but synonym replacement alone does not establish independent expression. Uncertain permissions or exceptions should be escalated to an appropriate legal reviewer. Personal-data compliance is a separate assessment and does not resolve copyright obligations.
Why each option fits or fails:
A. This addresses permission for the actual output and provides a checked rewrite when reproduction rights cannot be established.
B. Attribution and public accessibility do not themselves grant permission to reproduce a substantial protected passage in a commercial report.
C. Permission to use material in training does not establish the customer’s right to reproduce protected passages in published reports.
D. Replacing words can lower an exact-match score while retaining protected expression, so it does not establish that the revised passage is cleared.
Question 18
Topic: AI security controls
A security team reviews a proposed minimization change to an applicant-screening model’s evaluation export. Routine exports go to all testers; only designated evaluators can access a separate age-band store. The privacy team has confirmed a lawful basis for time-limited age-band evaluation.
Evaluation review record:
Required check: false-negative rates by age band
Model inputs: qualifications, relevant experience
Rows: same applicants, predictions and verified outcomes
Current export: applicant_id, age_band, prediction, verified_outcome
Proposed export: applicant_id, prediction, verified_outcome
Overall false-negative rate: 12% for both exports
Age-band source: restricted evaluation store, join on applicant_id
applicant_id: links to identifiable application records
Which interpretation of the proposed change is best supported by the record?
Options:
A. The export preserves age-band evaluation because age is excluded from model inputs; predictions and verified outcomes are sufficient for the required comparison.
B. The export leaves data exposure unchanged because applicant IDs remain linkable; retaining age bands in routine exports would create no additional disclosure risk.
C. The export narrows age-data exposure by removing age bands; the required comparisons still need a governed join with the separate evaluation store.
D. The export preserves age-band evaluation because its overall false-negative rate is unchanged; the aggregate result can replace the required subgroup comparison.
Best answer: C
Explanation: Data minimization should limit unnecessary exposure while preserving information needed for a justified evaluation. Predictions and verified outcomes support an overall false-negative rate, but age-band rates also require group membership. Identical aggregate performance cannot rule out different error rates across age groups. An attribute can therefore be necessary for evaluation even when it is not a model input.
Here, routine exports can omit age bands while designated evaluators perform a controlled join with the restricted store for the approved, time-limited analysis. This separates general testing access from sensitive evaluation access without abandoning the required check for potentially unfair outcomes.
The remaining linkable identifiers mean the export is not anonymous. Minimization can reduce disclosure impact without eliminating personal-data obligations or making every attack impossible.
Why each option fits or fails:
A. Age bands need not be model inputs to be necessary for evaluation; predictions and outcomes alone cannot assign errors to age groups.
B. Linkable IDs preserve privacy obligations, but removing age bands still reduces the information disclosed if a routine export is exposed.
C. Removing age bands limits their disclosure through routine exports, while approved access to group membership enables the required subgroup error measurements.
D. An unchanged overall false-negative rate can conceal differences between age groups, so it cannot replace the required age-band comparison.
Question 19
Topic: AI security threats
An AI engineering team evaluates a third-party model bundle containing model weights and a Python loader.
Evidence:
- The supplier provides the bundle and its SHA-256 checksum through the same unauthenticated download location. The team’s calculated bundle checksum matches.
- A detached signature covers only the weights. It verifies using a supplier public key authenticated separately through a trusted channel.
- No signature covers the loader, and neither component has undergone behavioral security testing.
Which assessment of the provenance evidence is best supported?
Options:
A. The checksum establishes supplier attribution and integrity for the entire bundle; behavioral safety remains unverified for both components.
B. The signature establishes supplier attribution and integrity for the weights only; behavioral safety remains unverified for both components.
C. The signature establishes supplier attribution and integrity for the entire bundle; behavioral safety remains unverified for both components.
D. The signature establishes supplier attribution and integrity for the weights only; behavioral safety is established for the signed component.
Best answer: B
Explanation: Provenance verification distinguishes an authenticated supplier relationship from a matching digest. A checksum comparison shows consistency with the supplied checksum, but it does not authenticate that checksum. Someone able to replace both the download and its checksum can preserve the match.
A valid signature checked against an independently authenticated supplier key binds the signed content to that signing identity and supports integrity since signing. Here, that assurance covers the weights, not the loader.
Signing does not establish harmless behavior. A supplier can sign an artifact containing malicious behavior or vulnerabilities. Treat model weights and accompanying executable code as supply-chain assets: obtain appropriate provenance evidence for each and assess their security separately. Behavioral testing informs that assessment but cannot prove the absence of every vulnerability.
Why each option fits or fails:
A. An attacker could replace both the bundle and its checksum at the unauthenticated location, preserving the match without establishing supplier attribution.
B. The independently authenticated key validates the supplier’s signature over the weights only; neither the signature nor the checksum establishes benign behavior.
C. The signature covers only the weights, so its authenticity and integrity assurances do not extend to the unsigned loader.
D. A signature establishes signing identity and integrity, not behavioral safety; legitimately signed weights can still contain harmful behavior.
Question 20
Topic: AI security threats
A supplier-inspection classifier labels product photographs as damaged or undamaged. An attacker seeks undamaged outputs for damaged products. The attacker can upload photographs and receive class labels, but cannot access weights, gradients, or training data.
Evaluation record:
- Adversarial training used bounded additive pixel noise.
- A held-out test found no successful evasions on 400 images using that perturbation family and bound.
- Small rotations and JPEG recompression preserve the correct damage label and are feasible for the attacker, but have not been evaluated adversarially.
Which follow-up test most directly addresses the remaining gap in adversarial-training coverage?
Options:
A. Use label-only feedback to search over class-preserving rotations and JPEG recompression settings for inputs that trigger misclassifications.
B. Generate additive pixel-noise attacks on a surrogate classifier and transfer them to the deployed classifier within the evaluated bound.
C. Use label-only feedback to search over more additive pixel-noise examples within the already evaluated perturbation bound.
D. Measure classification accuracy on randomly sampled, class-preserving rotations and JPEG recompression settings applied to representative test images.
Best answer: A
Explanation: Adversarial training can improve resistance to the perturbation family used during development without providing resistance to other transformations. The 400 passing trials support resistance to the evaluated additive-noise attacks only.
The coverage gap concerns whether an attacker can select rotations or JPEG recompression settings that preserve the true damage label but change the classifier’s output. A feedback-guided search using production class labels directly tests that threat. This is black-box evasion: the attacker manipulates inference inputs and does not need access to the deployed model’s internals.
Testing should verify that transformed images retain their correct labels and record the permitted access, search limits, and reproducible results. Passing this additional evaluation would extend the evidence, not prove universal robustness.
Why each option fits or fails:
A. Searching untested, feasible transformations with the attacker’s actual feedback directly evaluates whether resistance extends beyond the evaluated additive-noise attacks.
B. Transfer changes the model used to construct attacks, not their perturbation family; these additive-noise attacks leave rotation and recompression coverage unassessed.
C. More additive-noise tests strengthen evidence within the trained family but leave the feasible rotation and recompression attacks unassessed.
D. Random sampling measures typical transformed-input accuracy rather than using available feedback to seek settings deliberately selected to trigger misclassification.
Question 21
Topic: AI security testing
A security team plans to test whether a ticket-summarizing agent can be induced to request refunds through adversarial instructions embedded in retrieved tickets. Success means the application’s authorization layer blocks those requests before a business action occurs.
Approved scope:
Only the test environment and synthetic tickets are authorized. Refund attempts may target the simulator; production tool calls are prohibited.
Preflight routing record:
retrieval_index: synthetic_ticket_fixtures
refund_adapter.mode: replay
refund_adapter.on_replay_miss: forward
refund_adapter.forward_target: production_refund_api
refund_adapter.forward_credentials: production_refund_write
refund_adapter.per_call_limit: EUR 5
smoke_test.refund_result: replay_hit
smoke_test.production_calls: 0
Which interpretation of this record is supported?
Options:
A. The planned run is outside scope because replay misses can reach a write-enabled production refund API.
B. The planned run is within scope because synthetic tickets and the per-call cap keep refund effects within authorized bounds.
C. The planned run is within scope because the replay hit confirms that its refund calls remain simulated.
D. The planned run is outside scope because the replay hit demonstrates that the preflight already executed a production refund.
Best answer: A
Explanation: Authorization must be enforced across every reachable path, not inferred from a successful smoke test. The replay hit shows only that one request received a recorded response. On a replay miss, this adapter forwards the request to a production refund API using write-capable credentials, so the planned adversarial run is not isolated as authorized.
The live fallback should be disabled or replaced with the approved simulator, with unmatched requests failing closed. Containment should then be verified before testing. Injecting instructions through retrieved tickets is adversarial AI security testing; the replay smoke test is a functional preflight check, not proof that adversarial requests cannot trigger business actions.
Why each option fits or fails:
A. Unmatched calls are forwarded to the production refund API with write permissions, contrary to the prohibition on production tool calls.
B. Synthetic tickets bound input data, and the €5 cap limits transaction value, but neither prevents or authorizes production tool calls.
C. The replay hit establishes that one request used a recorded response; other requests can follow the configured live fallback.
D. The record explicitly reports zero production calls; receiving a replayed response does not demonstrate that a real refund occurred.
Question 22
Topic: AI security threats
A security reviewer examines the handoff from pretraining to fine-tuning. The fine-tuning job loads an intermediate checkpoint before continuing training. Experiment users may manage shared files but are not authorized to change approved model weights.
Handoff record:
Shared directory: /shared/stage1/
checkpoint.bin writers: training-service, experiment-users
checkpoint.bin signature: none
checkpoint.sha256 writers: training-service, experiment-users
Checkpoint gate: computed SHA-256 matches checkpoint.sha256 value
Independent trusted checkpoint reference: none
Training data/code/config writers: release-service only
Data/code/config gate: hashes match protected release manifest
Release manifest writers: release-service only
Which interpretation of the handoff’s integrity protection is supported by this record?
Options:
A. Direct model poisoning is prevented because the loader verifies training code and configuration against a protected release manifest.
B. Data poisoning is possible because experiment users can replace a checkpoint that is consumed during the next training stage.
C. Direct model poisoning is possible because experiment users can replace both the checkpoint weights and the checksum used for verification.
D. Direct model poisoning is prevented because the loader verifies checkpoint contents against a SHA-256 checksum before loading the weights.
Best answer: C
Explanation: Checkpoint weights are model artifacts, not training examples. Maliciously replacing weights between training stages is direct model poisoning during development, distinct from poisoning the data used to learn them.
Here, experiment users can replace both the unsigned checkpoint and its checksum. SHA-256 verification establishes that the file matches the supplied digest, but the digest is under the same writers’ control. Protecting and verifying training data, code and configuration does not close this separate integrity gap.
Checkpoint verification needs an independent trust anchor, such as an approved digest in a protected manifest or a signature checked against a trusted signing key, alongside restricted write access. Such verification establishes integrity and provenance, not benign model behavior. The record demonstrates an exposure, not that malicious alteration has occurred.
Why each option fits or fails:
A. Verified code and configuration do not authenticate separately writable checkpoint weights, which have no independent trusted reference.
B. Checkpoint weights initialize the model; changing them alters the model artifact rather than the training examples used for learning.
C. Replacing weights directly alters the model, and replacing the equally writable checksum allows the substituted checkpoint to pass the load check.
D. The checksum comparison cannot detect malicious substitution when experiment users can replace both the checkpoint and its expected digest.
Question 23
Topic: Privacy and compliance in AI security
A scheduling assistant uses a shared document index. Medical notes in the index have employee names replaced with employee codes. A shift supervisor may view names and approved working hours, but is not authorized to access medical records.
An authorized tester captures this record. The replay uses the same request, model version, and settings, with retrieval disabled.
Request: "What are Priya Shah's approved working hours?"
Retrieved roster: W62 | Priya Shah | 09:00-17:00
Retrieved medical note: W62 | Treatment for depression began in 2024.
Response: "Priya Shah works 09:00-17:00."
Response continued: "Her treatment for depression began in 2024."
Replay without retrieval: "I do not have Priya Shah's working-hours records."
Which interpretation of the personal-data exposure is best supported by this record?
Options:
A. Sensitive inference produced new personal health data by interpreting Priya’s approved working hours.
B. Retrieval disclosed anonymous health data by using a medical note with no employee name.
C. Training memorization disclosed personal health data by reproducing a medical record associated with Priya.
D. Retrieval disclosed personal health data by linking the coded note to Priya’s roster entry.
Best answer: D
Explanation: The evidence supports augmentation-data leakage through retrieval. The medical detail appears in a retrieved note, matches the generated disclosure, and is absent from the replay with retrieval disabled. This supports retrieval as the likely source rather than memorization or a new inference about working hours.
The shared code W62 links the medical note to Priya Shah in the roster. Replacing names with codes does not make information anonymous when accessible records permit identification.
Priya is the affected person, and the supervisor is the unauthorized recipient. The exposure flows from the medical-note index through retrieved context into generated content. Permission to view working hours does not authorize access to health history; retrieval authorization must enforce that boundary.
Why each option fits or fails:
A. The medical history is explicitly recorded in retrieved content, rather than newly inferred from Priya’s working hours.
B. Omitting the name does not establish anonymity because W62 can be linked to Priya through the retrieved roster.
C. The retrieved note directly supplies the disclosed detail; the replay without retrieval does not support training memorization as the likely source.
D. The shared code W62 links the medical note to Priya’s name, and the assistant passes that sensitive information to an unauthorized recipient.
Question 24
Topic: AI security in the organization
An AI security lead reviews a roadmap checkpoint. Ownership and policies are approved, integration-specific development and supplier procedures are in use, and implemented controls have passed scoped checks. API requests and retrieved documents do not update model weights.
Complete current threat register:
| Component | Recorded operation | Threat entry |
|---|---|---|
| Classifier | Local training records -> model weights | Training-data poisoning |
| API assistant | Customer text -> hosted inference | Training-data poisoning |
| Retrieval assistant | Internal documents -> model context | Training-data poisoning |
Checkpoint note:
Understand is complete because every inventoried AI component has a threat entry.
Which interpretation of this checkpoint is best supported?
Options:
A. Understand is complete because every component in the asset inventory is covered by the common threat entry.
B. Understand is incomplete because inventory coverage does not establish applicable threats for the different components and their trust boundaries.
C. Understand is complete because successful tests of implemented controls establish threat coverage across the different AI components.
D. Understand is incomplete because implementing and testing controls before finalizing threat analysis violates GUARD’s required activity sequence.
Best answer: B
Explanation: Understand identifies applicable threats from intended use, assets, data flows, trust boundaries, attacker capabilities, and potential impact. The inventory supports this work but does not replace it.
The classifier’s training process can expose learning data to poisoning. Hosted inference requires analysis of submitted data, supplier handling, and attackable inputs. Retrieval requires analysis of document trust and access permissions, including indirect prompt injection and unauthorized disclosure. Requests and retrieved documents do not update weights here, so the common training-poisoning entry does not establish coverage. Supplier-side poisoning could still matter, but it requires analysis of that supplier path.
Govern, Understand, Adapt, Reduce, and Demonstrate are interacting activities. Existing procedures, controls, and checks can continue while the missing threat analysis is developed. Passing scoped checks informs assessment; it does not establish complete threat coverage.
Why each option fits or fails:
A. An inventory establishes which assets exist, but repeating a poisoning label does not analyze how each component can be attacked.
B. Applying one training-related entry across all three operations leaves external processing and retrieval exposures without a differentiated threat analysis.
C. Scoped control checks demonstrate what was tested, not that the threat register identifies every applicable attack path.
D. GUARD activities interact, so existing adaptations, controls, and checks need not wait for a finalized threat analysis.
Question 25
Topic: AI security in the organization
A procurement assistant sends staff requests and externally supplied vendor PDFs to a supplier-hosted model. Its customer-operated tool gateway grants the assistant the procurement manager’s permissions, including editing vendor bank details. An attacker can alter a submitted PDF but cannot access staff accounts or the model. Following malicious instructions in that PDF could cause payment diversion.
The finance risk owner, who is authorized to accept this risk, signed the following deployment record.
Deployment record:
Supplier clause: "The service provides input filtering."
Filter coverage: Not specified
Supplier adversarial test evidence: Not supplied
Incident notification: No commitment
Customer test: 24 ordinary quote summaries passed
Inherent risk: High
Residual risk: Low, credited to supplier filtering
Acceptance: Signed by finance risk owner
Which interpretation of the recorded risk reduction is best supported?
Options:
A. The contract demonstrates filtering effectiveness for staff requests; the owner’s signature accepts uncertainty about vendor-PDF coverage.
B. The supplier’s filter ownership transfers the injection risk; the owner’s signature accepts the remaining exposure from customer tool permissions.
C. The normal-summary tests demonstrate filtering effectiveness for vendor PDFs; the owner’s signature accepts the remaining financial risk.
D. The contract records a treatment commitment; the owner’s signature records acceptance, not evidence supporting the residual-risk reduction.
Best answer: D
Explanation: Inherent risk describes exposure before treatment. Here, hostile vendor content can cross the document-to-model trust boundary and influence actions affecting vendor bank details. The customer’s tool permissions make payment diversion a credible impact.
Supplier input filtering is a treatment commitment, not demonstrated effectiveness. The record supplies neither coverage nor adversarial results. Successful ordinary summaries establish task performance, not resistance to indirect prompt injection. The credited reduction from High to Low is therefore unsupported by this evidence.
The finance owner’s signature records acceptance; it does not validate the risk estimate. The owner may knowingly accept uncertainty, but that uncertainty must be visible. Obtain filtering scope, relevant attack-test evidence and incident-reporting commitments, then reassess residual risk and confirm informed acceptance. Supplier operation does not remove customer responsibility for authorization at the tool boundary.
Why each option fits or fails:
A. The clause identifies neither covered input routes nor tested effectiveness, so it cannot substantiate reduced risk even for staff requests.
B. Operating the filter does not transfer the customer’s financial-loss risk from model-driven actions using customer tool permissions.
C. Successful ordinary summaries measure clean-task performance; they do not show that malicious instructions embedded in vendor PDFs are blocked.
D. The filtering promise and acceptance signature document treatment and a decision, but neither demonstrates the effectiveness needed to justify the reduced rating.
Questions 26-40
Question 26
Topic: AI security controls
A company is reviewing a supplier-hosted model API for an internal document assistant. The supplier operates inference, stores service logs and deploys model updates. The company controls its application and access rules.
Review evidence:
- Signed contract: 99.9% monthly availability, with no terms covering submitted-data use, security-incident notification or material model changes.
- Supplier assurance: an organization-wide security certification, with no evidence linking those responsibilities to this API service.
- Customer controls: an AI risk owner is assigned, and integration security testing, compliance review and user training are complete.
Which next action would most directly close the demonstrated supplier-governance gap?
Options:
A. Obtain approval of an internal AI policy assigning data handling, incident response and model-change oversight to customer teams, and document their acceptance.
B. Obtain an independent compliance assessment of the supplier’s published data, incident and model-update policies, and record approval of the intended use.
C. Obtain service-specific contractual commitments for data handling, incident notification and material model changes, and verify the supplier’s supporting operational evidence.
D. Obtain release-test evidence for integration logging, error handling and model-version compatibility, and incorporate it into the customer’s secure-development approval.
Best answer: C
Explanation: Supplier assurance must follow the operating boundary. Here, the supplier controls inference, service logs and model updates. Availability terms and an organization-wide certification do not establish how submitted data may be used, how incidents will be reported or how material changes will be managed. The customer needs agreed service-specific responsibilities and evidence that the supplier fulfills them.
Supplier oversight within SEC PROGRAM addresses this gap. AI PROGRAM establishes accountability and risk ownership, but customer policy alone cannot bind the supplier. SEC DEV PROGRAM supports secure development, while DEV PROGRAM organizes development practices; customer integration evidence does not demonstrate supplier-controlled operations. CHECK COMPLIANCE evaluates applicable requirements, and SEC EDUCATE develops staff capability. These activities complement supplier assurance rather than replace missing commitments and operational evidence.
Why each option fits or fails:
A. AI PROGRAM can assign customer accountability, but an internal policy does not create supplier obligations or demonstrate the hosted service’s controls.
B. CHECK COMPLIANCE evaluates applicable requirements, but reviewing general policies does not establish service-specific contractual duties or evidence of their operation.
C. The supplier controls model operations, so agreed service-specific obligations and operational evidence directly address the missing commitments and assurance.
D. Integration testing supports SEC DEV PROGRAM, but it does not establish the supplier’s obligations for submitted data, incident notification or material changes.
Question 27
Topic: AI security threats
An expense assistant’s application instructions limit it to explaining reimbursement rules and displaying claim status. Approving claims is outside its permitted role.
A manager can approve claims through a separate portal. In a message sent directly to the assistant, the manager instructs it to override its role restriction and approve a pending claim. The instruction is submitted for execution, not as text to analyze.
How should the manager’s message be classified?
Options:
A. An indirect prompt-injection attempt.
B. An authorized change to the assistant’s task.
C. A direct prompt-injection attempt.
D. A model-output-injection attempt.
Best answer: C
Explanation: Direct prompt injection introduces conflicting instructions through a user’s direct interaction with a model or assistant. Here, the manager explicitly asks the assistant to override its application-defined role and perform claim approval. The manager’s authority in a separate portal does not expand the assistant’s authorization.
Whether the assistant complies is a separate issue: an unsuccessful attempt is still an attempt. An ordinary task change remains within the service’s permitted scope, while a benign quoted instruction is treated as material to analyze rather than an instruction to execute. Classification therefore depends on the instruction’s source, its intended use, and its conflict with the service’s constraints.
Why each option fits or fails:
A. Indirect injection arrives through lower-trust content such as retrieved documents or tool responses, rather than the manager’s direct interaction with the assistant.
B. The manager’s approval authority in the separate portal does not authorize overriding this assistant’s restricted role.
C. The manager directly supplies an instruction to override the application’s role restriction, making this a direct prompt-injection attempt.
D. Model output injection concerns untrusted generated content reaching an unsafe execution or rendering sink, not conflicting instructions submitted to an assistant.
Question 28
Topic: AI security threats
A multi-tenant conversational application keeps conversation context in a shared cache. Conversation identifiers are unique only within each tenant. Two testers request summaries of their own conversations. The tenant and user fields below come from verified authentication.
Test trace:
Stored (Cedar, u14, c7): "Renewal ceiling: EUR 42,000."
Stored (Maple, u91, c7): "Renewal ceiling: EUR 9,000."
Request 1: tenant=Cedar user=u14 conversation=c7
Cache GET key=ctx:c7 result=MISS
Cache PUT key=ctx:c7 context="Renewal ceiling: EUR 42,000."
Request 2: tenant=Maple user=u91 conversation=c7
Cache GET key=ctx:c7 result=HIT context="Renewal ceiling: EUR 42,000."
Model input 2 message="Summarize my conversation."
Model input 2 context="Renewal ceiling: EUR 42,000."
Model output 2: "Your renewal ceiling is EUR 42,000."
Which interpretation is best supported by this trace?
Options:
A. A session-authentication failure causes the application to associate the second request with the first tenant’s authenticated identity.
B. A model-memorization failure causes the model to reproduce another tenant’s conversation from information retained during training.
C. A cache-key collision causes the application to supply one tenant’s conversation context to another tenant’s model request.
D. An indirect prompt injection causes the model to treat cached conversation content as instructions to disclose another tenant’s context.
Best answer: C
Explanation: Session isolation must be enforced by the application components that store and retrieve conversation state. Tenant-local conversation identifiers can repeat, so a shared cache keyed only by ctx:c7 treats different tenants’ conversations as the same entry. Maple’s request retrieves Cedar’s context despite having the correct authenticated identity. The boundary failure occurs before model generation.
Cache keys should incorporate trusted tenant identity and the conversation identifier, with conversation-owner authorization enforced on retrieval. Verification should cover both cold and warm caches using identical local identifiers across tenants.
This is a conventional access-control failure in an AI application’s runtime. Model instructions or prompt-injection filtering would not repair the missing cache isolation.
Why each option fits or fails:
A. The second request retains Maple’s verified identity; the wrong conversation context appears when the application reads the shared cache.
B. The disclosed amount is already present in the second request’s model input; the trace does not demonstrate leakage from learned parameters.
C. Both tenants use conversation identifier c7, and the shared key ctx:c7 omits tenant identity, causing Maple’s request to retrieve Cedar’s context.
D. The cached note contains ordinary conversation data, not malicious instructions, and the application supplies the foreign context before model generation.
Question 29
Topic: AI security threats
During an authorized security test, a standard account requests lengthy repetitions and redundant searches of unchanged public material from a company-funded AI research assistant.
Review evidence:
- Limits: 10 requests per minute and 4 concurrent tasks per account.
- Observed use: 2 requests per minute, up to 4 concurrent tasks, approximately 30,000 billed tokens and 40 paid searches per task.
- Results: Mostly repetitive output with little task value; some responses contain confidential-looking document titles.
- Data trace: Those titles originated in the submitted requests, and searches accessed only public content.
- Impact: Charges from the account exceeded the operator’s $100 daily budget; response latency and availability remained normal.
Which threat is most directly demonstrated by these results?
Options:
A. Cost abuse through disproportionate paid model and tool consumption.
B. Resource exhaustion through saturation that denies service to other users.
C. Data exfiltration through unauthorized disclosure of restricted retrieved material.
D. Model extraction through reconstruction of model behavior from collected responses.
Best answer: A
Explanation: Cost abuse drives the operator’s expenditure through paid processing or downstream actions that provide little useful value. Here, each task consumes many billed tokens and paid searches, producing repetitive results and exceeding the budget.
Request-rate limits constrain how often tasks start, not how expensive each task becomes. Request complexity, generation length, tool calls and concurrent execution all affect resource consumption. Per-task token and tool-use budgets, combined with aggregate spending limits, address this financial exposure more directly.
Data exfiltration concerns unauthorized transfer of data. Confidential-looking titles alone do not establish such a transfer: the trace attributes them to submitted text rather than restricted retrieved material.
Why each option fits or fails:
A. The low-value tasks consume substantial paid tokens and searches while exceeding the budget, even though request and concurrency limits are respected.
B. Latency and availability remained normal, so the observed resource consumption does not establish saturation or denial of service.
C. The titles came from the submitted requests, and retrieval accessed only public material, so the trace does not demonstrate disclosure of restricted data.
D. Repeated responses do not demonstrate reconstruction of model behavior; the observed harm is excessive expenditure rather than acquisition of a functional model approximation.
Question 30
Topic: AI security threats
A failed fine-tuning job automatically attaches a crash report to a public issue tracker. The security team confirms:
- Training files are encrypted at rest.
- The reporter’s service identity has permission to read training files and environment variables and to publish public issues.
- The attachment contains two complete confidential customer-support messages, experiment settings, and an active repository access token copied from an environment variable.
- No model weights or checkpoints are attached.
- Anonymous readers can download the attachment.
Which assessment best describes the demonstrated exposure?
Options:
A. Development-time data leakage and credential exposure through unauthorized access to the training-data store.
B. Configuration and credential exposure, with the attached training inputs protected by storage encryption.
C. Model leakage and credential exposure through the reporter’s permitted diagnostic publication.
D. Development-time data leakage and credential exposure through the reporter’s permitted diagnostic publication.
Best answer: D
Explanation: Crash diagnostics create a separate data flow and can expose sensitive copies outside the training environment. The customer-support messages are development-time training data, not model parameters. Publishing these diagnostic copies causes development-time data leakage, while publishing the active repository token creates credential exposure. The reporter’s technical permissions enabled the disclosure; a storage break-in was unnecessary.
Encryption at rest protects stored files, but an authorized reader receives usable data that can subsequently be overshared. Protection must therefore cover diagnostic bundles, experiment records, source, configuration, credentials, and model artifacts. Limit diagnostic exports to necessary nonsensitive fields and control publication access. An exposed live token also requires revocation or rotation; encrypting the original storage does not resolve that exposure.
Why each option fits or fails:
A. The reporter used its assigned permissions; the evidence shows public access to exported copies, not unauthorized access to the training-data store.
B. Encryption at rest does not protect plaintext samples after an authorized process copies them into a publicly readable attachment.
C. The attachment contains raw training samples copied directly from development data, not a model artifact. No model-inference output demonstrates disclosure of memorized information.
D. Confidential training messages and a live credential became anonymously downloadable when the reporter published their copies using its assigned permissions.
Question 31
Topic: AI security controls
An expense-reimbursement portal uses AI to flag possible duplicate claims. Employees can view the matched claims and a plain-language reason for each flag. The portal displays:
AI flags possible duplicates; a finance reviewer decides reimbursement. Similar claim details can refer to separate purchases.
The service operates as follows:
- A finance team reviews the original receipts and can correct flags before deciding reimbursement.
- The only help route is an automated FAQ that accepts no messages and provides no staff contact details.
- Model confidence scores are not displayed.
Which assessment of the service’s transparency is best supported?
Options:
A. Transparency is sufficient because an authorized finance team reviews each flag before deciding reimbursement.
B. Transparency is sufficient because employees can inspect the matched claims and understand the flag’s stated reason.
C. Transparency is incomplete because employees lack a contact route to the team responsible for decisions.
D. Transparency is incomplete because each AI flag is presented without its model confidence score.
Best answer: C
Explanation: Transparency for an AI-supported service includes an understandable account of the AI’s role, relevant limitations, and how to reach a responsible person. The notice distinguishes an automated flag from a human reimbursement decision and warns that similar records may not be duplicates. Showing matched claims also helps employees understand why a flag appeared.
The missing element is a usable contact route to the finance team. Internal human review is an oversight control; its existence does not establish user-facing transparency. Confidence scores may provide additional information, but displaying them is not essential in every service and does not replace explaining limitations or enabling human contact.
Why each option fits or fails:
A. Authorized review supports human oversight, but it does not let employees raise questions or corrections with a responsible person.
B. Visible records and a stated reason aid understanding, but they do not provide access to a responsible person when a flag is disputed.
C. Employees are told the AI’s role and limitation, but the FAQ provides no way to contact anyone accountable for the reimbursement decision.
D. Confidence scores are not a universal transparency requirement; their absence does not establish the gap demonstrated by the inaccessible human contact route.
Question 32
Topic: AI security threats
An accounts-payable assistant can use an invoice-approval tool with the signed-in employee’s permissions. The employee may approve invoices, but the application does not check tool actions against the current task. A developer claims that wrapping retrieved text in <supplier_note> tags creates a trust boundary that prevents external instructions from controlling the assistant.
An authorized adversarial tester produces this sanitized record:
User task: Summarize INV-42; do not approve it.
Tester control: Retrieved supplier document only.
Document use: Inference context, not training.
Wrapper: <supplier_note> ... </supplier_note> in both runs.
Original text: Invoice facts; manual quote, "Check totals before approval."
Original result: Summary; no tool call.
Changed text: Original content plus a direction to approve before summarizing.
Changed result: approve_invoice(INV-42) accepted; invoice marked approved.
Which interpretation of the changed result is supported?
Options:
A. The result is indirect prompt injection: tagged supplier content redirected the assistant beyond the user’s task, so the claimed trust boundary failed.
B. The result is data poisoning: inserting supplier instructions into retrieved context changed the model’s learned behavior governing its decisions about invoice approval.
C. The result is authorized workflow execution: the employee’s invoice-approval permissions allowed the supplier instruction to extend the assistant’s assigned task.
D. The result is direct prompt injection: combining supplier content with the user request made the added instruction part of the user’s direct interaction.
Best answer: A
Explanation: Delimiters are formatting cues, not an enforced trust or authorization boundary. In the changed run, an instruction from a retrieved supplier document redirected the assistant from summarization to approval. The tester controlled that lower-trust document, making the entry point indirect prompt injection. The unchanged tags and successful approval demonstrate that the claimed isolation did not hold in this test.
The original manual quotation was task data, not automatically an attack merely because it used imperative language. Likewise, an ordinary user task change is not automatically malicious. The employee’s approval permissions do not authorize every agent action. Tool authorization should enforce the user’s current delegated intent rather than depend on the model respecting tag boundaries.
Why each option fits or fails:
A. The tester changed lower-trust retrieved content, and the resulting approval contradicted the explicit user task despite the unchanged tags.
B. The document was used for inference rather than training, so this modification influenced the current interaction rather than poisoning learning data.
C. The employee’s ability to approve invoices does not authorize the assistant to approve one during a task explicitly limited to summarization.
D. Combining inputs does not change their origin; the tester controlled a retrieved document, not the user’s direct interaction with the model.
Question 33
Topic: Privacy and compliance in AI security
A German employer licenses a vendor-hosted AI assistant originally supplied for drafting routine internal announcements. The employer changes its intended purpose and configures it to score CVs and rank job applicants.
- Recruitment use: A recruiter makes final interview decisions using the ranking as the main selection evidence.
- Technical change: The underlying model is unchanged. The supplier did not design the recruitment use.
- Data responsibilities: The employer determines recruitment purposes and retention. The supplier processes applicant data only on the employer’s documented instructions.
Under the EU AI Act and GDPR, which classification should govern the employer’s compliance review for the recruitment use?
Options:
A. The recruitment system is high-risk; the employer is an AI Act provider and deployer and a GDPR controller.
B. The recruitment system is not high-risk; the employer is an AI Act deployer only and a GDPR controller.
C. The recruitment system is high-risk; the employer is an AI Act provider and deployer and a GDPR processor.
D. The recruitment system is high-risk; the employer is an AI Act deployer only and a GDPR controller.
Best answer: A
Explanation: The EU AI Act assesses an AI system’s intended use, not merely its underlying model or hosting arrangement. Ranking applicants for recruitment is a high-risk employment use here because it materially influences selection. Human involvement in the final decision does not automatically change that classification.
When a deployer changes the intended purpose of a previously non-high-risk system so that it becomes high-risk, it assumes provider obligations for that system. The employer also operates the recruitment system, so it is a deployer.
The GDPR analysis is separate. Determining the purposes and means of applicant-data processing makes the employer a controller; the supplier processing data on its instructions acts as a processor. The changed use therefore requires reassessment of high-risk AI controls and data-protection duties, including lawful basis, purpose limitation, transparency, minimization and retention.
Why each option fits or fails:
A. The employer establishes and operates the high-risk recruitment use, while determining the purposes of applicant-data processing establishes its separate GDPR controller role.
B. Recruitment ranking materially influences access to interviews, so final human decision-making does not by itself remove this use from the high-risk category.
C. The employer determines why applicant data is processed and its retention, making it the controller; the supplier’s instructed processing does not transfer that role.
D. Changing a previously non-high-risk system’s intended purpose into high-risk recruitment use creates provider responsibilities for the employer, even when the supplier continues hosting it.
Question 34
Topic: AI security controls
An HR assistant retrieves employee accommodation records at runtime. To minimize sensitive-data exposure, the team reduces the retrieval cap from 20 employee records per response to one. Each record concerns a different employee and contains a highly sensitive medical diagnosis.
A validation team repeats identical before-and-after workloads. Unauthorized probes use an account without access to the requested records. Utility tests use appropriately authorized accounts.
Validation report:
| Measure | Before | After |
|---|---|---|
| Unauthorized probes returning records | 10 of 10 | 10 of 10 |
| Employee records per probe response | 20 | 1 |
| Fields disclosed | Name, diagnosis | Name, diagnosis |
| Authorized task tests passed | 48 of 50 | 48 of 50 |
Which interpretation of the change is supported by this report?
Options:
A. Disclosure volume fell, unauthorized retrieval became less frequent in the tests, and potential harm to an exposed employee decreased.
B. Disclosure volume fell, unauthorized retrieval stayed equally frequent in the tests, and potential harm to an exposed employee decreased.
C. Disclosure volume fell, unauthorized retrieval became less frequent in the tests, and potential harm to an exposed employee remained unchanged.
D. Disclosure volume fell, unauthorized retrieval stayed equally frequent in the tests, and potential harm to an exposed employee remained unchanged.
Best answer: D
Explanation: Data minimization reduces the amount of sensitive information exposed when a control fails. Here, the cap reduces the number of affected employees per response from 20 to one, while authorized-task success remains 48 out of 50. This is a useful reduction in exposure without observed utility loss in the tested workload.
It does not demonstrate improved retrieval authorization: unauthorized records were returned in all ten probes before and after. The remaining record still identifies an employee and reveals a highly sensitive diagnosis, so serious individual harm remains possible. Minimization complements access control; it does not repair unauthorized retrieval, make the retained data anonymous, or remove applicable privacy obligations.
Why each option fits or fails:
A. All ten probes still succeeded, and the retained name and diagnosis preserve the potential for serious harm to an affected employee.
B. The same name and sensitive diagnosis remain exposed for each affected employee; reducing the number of employees does not reduce individual harm.
C. All ten probes disclosed unauthorized records both before and after the change, so the observed disclosure frequency did not decrease.
D. The cap reduced records per response from 20 to one, while every unauthorized probe still disclosed identifying information and a highly sensitive diagnosis.
Question 35
Topic: AI security threats
A travel-policy assistant answers from an internal knowledge base. The approved nightly hotel cap is €150 and has not changed. A security team compares a recorded baseline with a repeated test.
User request (both tests): "What is the nightly hotel cap?"
Model artifact (both tests): identical verified digest
Training, fine-tuning, online learning: no updates
Baseline retrieved record: travel/hotel, revision 12
Baseline full excerpt: "The nightly hotel cap is EUR 150."
Baseline response: "EUR 150."
Audit: attacker used a compromised editor account to replace travel/hotel
Processing: document index refreshed
Repeat retrieved record: travel/hotel, revision 13
Repeat full excerpt: "The nightly hotel cap is EUR 450."
Repeat response: "EUR 450."
Which threat interpretation is best supported by this record?
Options:
A. Indirect prompt injection through insertion of an instruction that redirects the assistant’s behavior.
B. Development data poisoning through replacement of an example used to train the assistant’s model.
C. Augmentation data manipulation through replacement of a fact in the retrieved knowledge-base record.
D. Direct model poisoning through modification of the parameters in the deployed model artifact.
Best answer: C
Explanation: Augmentation data manipulation targets external information supplied to a model during operation, such as retrieved knowledge-base content. Here, the attacker replaced the hotel cap with EUR 450. Reindexing made the altered record available to retrieval, and the assistant repeated its false value.
The user request and model artifact were unchanged, and no learning updates occurred. The evidence therefore identifies manipulation of retrieved context rather than poisoning of training data or model parameters. Because the complete excerpt contains no behavioral instruction, it does not demonstrate indirect prompt injection. Controls should protect source integrity and editing permissions, preserve access restrictions through indexing and retrieval, and account for the trustworthiness of retrieved information during response generation.
Why each option fits or fails:
A. The complete retrieved excerpt contains a false policy fact, not an instruction directing the assistant to change its behavior.
B. The index refresh changed retrieval content, while the record explicitly shows no training, fine-tuning, or online-learning updates.
C. The attacker changed the retrieved hotel-cap fact, and the response followed that change while the model artifact remained identical.
D. The unchanged verified model-artifact digest supports unchanged model parameters, rather than modification of the deployed model.
Question 36
Topic: Privacy and compliance in AI security
A privacy reviewer at an insurer examines records from an internal drafting assistant. Its prompt-monitoring service is used to detect prompt injection.
Monitoring store: retained prompt event
request_id: r-731
actor_alias: u-26
prompt: "Customer c-49 is receiving chemotherapy. Draft a claims update."
result: no attack detected
Alias directory: encrypted at rest
u-26 -> employee Maya Ortiz
c-49 -> customer Daniel Reed
investigator_access: monitoring store and alias directory
Which interpretation of the privacy exposure is supported by these records?
Options:
A. The telemetry is personal data about the customer only; the employee alias records technical activity rather than identifiable employee activity.
B. The telemetry is personal data about the employee only; customer health information is anonymous once the customer’s name is replaced.
C. The telemetry is anonymous for both people; holding identities in a separately encrypted directory breaks the link to retained prompt content.
D. The telemetry is pseudonymized personal data about both people; employee activity and customer health information remain linkable to their identities.
Best answer: D
Explanation: Prompt monitoring creates an additional collection of user content, not merely a count of attacks. Here, investigators can resolve the retained aliases through the directory, so the telemetry is pseudonymized rather than anonymous. It connects an employee to assistant use and a customer to health information. The customer is affected even though the employee submitted the prompt.
The no attack detected result does not remove these privacy implications. Monitoring needs a defined purpose, proportionate collection, restricted access and suitable retention. Separating and encrypting the directory can reduce exposure, but does not eliminate the available identity links.
Why each option fits or fails:
A. The directory links the employee alias to a named person, making the associated usage record personal data rather than anonymous technical metadata.
B. The customer alias remains resolvable, so removing the customer’s name does not anonymize the retained health information.
C. Investigators can access the directory and resolve both aliases, so separation and encryption do not make the telemetry anonymous.
D. Resolving the aliases ties assistant use to the employee and the chemotherapy statement to the customer.
Question 37
Topic: AI security controls
An AI service approves employee access to sensitive engineering files. Its fallback holds uncertain requests for manual review; no access is granted while a request is pending. Every request must receive a final decision within one working day.
Completed offline replay: The same 1,000 requests were evaluated with and without the fallback, and every manual review was completed.
- Without the fallback: 18 unauthorized approvals.
- With the fallback: 4 unauthorized approvals, all from high-confidence automatic decisions.
- The fallback referred 120 requests for manual review.
Production conditions: Each working day brings the same volume and uncertainty mix. Reviewers can complete at most 80 requests per day, with no additional capacity available.
Which assessment is best supported?
Options:
A. The fallback improves observed authorization safety and maintains the required availability, although high-confidence automatic errors remain.
B. The fallback improves observed authorization safety but cannot sustain the required availability, and high-confidence automatic errors remain.
C. The fallback improves observed authorization safety but cannot sustain the required availability, with manual review controlling the remaining authorization errors.
D. The fallback leaves observed authorization safety unchanged and cannot sustain the required availability, with high-confidence automatic errors remaining.
Best answer: B
Explanation: A manual fallback can reduce unsafe decisions while creating an availability bottleneck. The completed replay shows a safety benefit: unauthorized approvals decreased from 18 to 4. This is evidence for the tested workload, not a guarantee of safe behavior in wider use. The remaining errors came from confident automatic decisions that never reached reviewers.
Production generates 120 referrals per day, but reviewers can complete only 80. The backlog therefore grows by 40 requests per day, making the one-working-day decision requirement unsustainable. Holding requests protects against access while review is pending, but does not preserve timely service. Additional review capacity would address the bottleneck; separate controls and validation are needed for confident automatic errors.
Why each option fits or fails:
A. Holding requests prevents access while they wait, but retaining them in a queue does not satisfy the one-working-day decision requirement.
B. Unauthorized approvals fell from 18 to 4, but daily referrals exceed review capacity by 40, and the remaining errors bypass manual review.
C. The four remaining unauthorized approvals were high-confidence automatic decisions, so routing uncertain requests to reviewers does not control those failures.
D. The reduction from 18 to 4 unauthorized approvals demonstrates a measured safety benefit, even though authorization risk has not been eliminated.
Question 38
Topic: AI security in the organization
After an access-policy change, a company’s model registry becomes publicly accessible. It contains proprietary model weights used by the production prediction service. The confidential training records remain in a separate private repository.
An authorized reviewer checks anonymous registry access:
Download model weights: succeeds; complete file retrieved
Replace model weights: denied
Which asset and security property are most directly at risk from this change?
Options:
A. Confidentiality of the training dataset.
B. Confidentiality of the model weights.
C. Integrity of the model weights.
D. Availability of the prediction service.
Best answer: B
Explanation: Model weights are an asset distinct from the training dataset and the running prediction service. Public download access directly threatens their confidentiality because outsiders can copy the proprietary model artifact. This enables direct model theft, unlike model extraction through collected prediction inputs and outputs.
The denied replacement request provides no evidence of unauthorized weight modification, and the observations do not establish a service outage. Training records remain separately protected. Exposed weights may increase the risk of inferring sensitive information from the model, but downloading weights does not itself demonstrate disclosure of the underlying training dataset.
Why each option fits or fails:
A. The training records remain in a separate private repository; downloading weights does not by itself establish direct disclosure of that dataset.
B. Anonymous download permits outsiders to copy the proprietary weight files, exposing the model artifact even though replacement remains restricted.
C. Replacement is denied, so the observed exposure permits reading the weight files rather than altering them.
D. The successful registry download demonstrates artifact exposure, not resource exhaustion or interruption of the production prediction service.
Question 39
Topic: AI security controls
A housing provider uses AI to recommend repair-request priority. Residents may submit text, photographs, or both. A security reviewer inspects the following service record.
Service review record:
Resident-facing screen
AI recommends repair priority. It may make mistakes.
Priority: Routine
Reason: No urgent risk indicators were detected.
Model confidence: High
Request review: Contact the responsible repair coordinator.
Internal validation note
The priority model receives text only, not uploaded photographs.
Weekly checks: Photo-only urgent reports can receive Routine.
The review button reaches a coordinator who can inspect photos
and change the priority.
Which interpretation of the service’s transparency is best supported by this record?
Options:
A. Transparency is adequate because weekly validation and a confidence indicator communicate the recommendation’s reliability alongside the available human review route.
B. Transparency is incomplete because users are not told that uploaded photographs are excluded from the information used to recommend repair priority.
C. Transparency is adequate because a general error warning, case-level reasons, and access to a responsible coordinator cover the service’s limitations.
D. Transparency is incomplete because users are not given the model weights and technical processing details needed to reproduce the priority recommendation.
Best answer: B
Explanation: Transparency should help users understand relevant limitations and reach a responsible person. The review button provides that contact here: the coordinator can inspect the original photographs and change priority.
The missing disclosure is that the AI assesses text only. A resident submitting photographs could interpret a routine recommendation and high confidence as evidence that the photographed damage was assessed, although those images never reached the model. A general warning that AI can make mistakes does not explain this specific, consequential limitation.
Displayed reasons relate to explainability, while weekly checks support continuous validation. Neither replaces communicating relevant limitations to users. The service should clearly disclose its text-only assessment and explain how residents can obtain human review of photographic evidence.
Why each option fits or fails:
A. Internal validation does not communicate its findings to residents, and a confidence indicator does not reveal that uploaded photographs were excluded.
B. The known text-only restriction directly affects residents submitting photographs, but it is absent from the resident-facing information.
C. The general warning and displayed reason do not disclose that photographic evidence is excluded from assessment, even though human review is available.
D. Meaningful transparency requires relevant limitations and responsible contact, not publication of model weights or sufficient technical detail to reproduce predictions.
Question 40
Topic: AI security threats
A clinical classifier trained on 60 confidential patient records predicts treatment outcomes from age band, routine measurements, and genetic-marker status. Its API exposes full outcome probabilities to six decimal places and cannot retrieve patient records.
An authorized reviewer tests 20 patients from the clinic. A custodian checks the results against confidential reference values only after the reviewer locks the estimates.
Test report:
Known: age band, routine measurements, observed outcome
Withheld: genetic-marker status and training-set inclusion
Queries: keep known inputs fixed; try both marker statuses
Estimate: marker status maximizing probability of observed outcome
No-API baseline: 10 of 20 marker estimates correct
API-assisted result: 18 of 20 marker estimates correct
API returned: probabilities only
Which interpretation is most directly supported by this report?
Options:
A. The reviewer reproduced the classifier’s behavior through model extraction.
B. The reviewer inferred training-set membership through membership inference.
C. The reviewer obtained original patient records through memorized-data disclosure.
D. The reviewer inferred a confidential attribute through model inversion.
Best answer: D
Explanation: Model inversion uses model responses, often combined with auxiliary information, to infer sensitive information. Here, the reviewer combines known patient inputs and treatment outcomes with prediction probabilities to estimate a withheld genetic-marker status. The estimates match the confidential reference for 18 of 20 patients, compared with 10 of 20 without API feedback.
Successful inversion need not reconstruct a complete or exact original training record. The demonstrated outcome is sensitive-attribute inference even though the API returns only probabilities. The report does not establish whether memorization contributed internally, and it does not identify which patients were included in training. Detailed prediction feedback and repeated query access can make inference more feasible, but these results do not guarantee success for every patient or another population.
Why each option fits or fails:
A. The measured result concerns patient attributes, not the reproduction of the classifier’s input-output behavior by a substitute model.
B. The report measures marker-status accuracy, not predictions about training-set inclusion, so it does not demonstrate membership inference.
C. Only probabilities were returned; matching marker estimates to reference values does not show that original patient records were reproduced.
D. Varying the hidden marker input and comparing probabilities of the known outcome produced independently confirmed estimates of a confidential attribute.
Review your attempt
| Domain | Correct | Missed or guessed question numbers |
|---|---|---|
| AI security in the organization | ___ / 6 | ___ |
| AI security threats | ___ / 15 | ___ |
| AI security controls | ___ / 11 | ___ |
| AI security testing | ___ / 3 | ___ |
| Privacy and compliance in AI security | ___ / 5 | ___ |
Use your result to choose the next topic to practise. An immediate repeat may test answer memory. Work through fresh questions in IT Mastery, check the cheat sheet and return to mixed practice after you can explain the distinctions without seeing the choices.
If anything seems incorrect or unclear, send private AISP feedback with this page’s question number and the detail you want reviewed.
Continue in the web app
Use IT Mastery for interactive EXIN AISP practice with mixed sets, timed mocks, topic drills, explanations, and progress tracking.