AAIA Practice Questions: 8 Worked Examples, Fully Explained
The fastest way to understand the ISACA Advanced in AI Audit (AAIA) exam is to work through questions built the way it builds them. Here are eight, spread across the three domains in roughly the proportions the exam uses. Each comes with the answer, the reason it wins, and a line on every wrong option explaining why it loses. That last part matters most, because on this exam the wrong options are rarely wrong in isolation.
These questions are original. They are not taken from the live exam or from ISACA's own practice materials, and they are not items from our paid question bank. Try each one before you open the explanation: the habit you are building is committing to an answer under uncertainty, not recognising one after the fact.
How AAIA questions are built
The AAIA exam has 90 multiple-choice questions in 150 minutes, scored on ISACA's 200–800 scale with 450 to pass. Every question has one stem and four options, and you pick the single best answer. There is no penalty for a wrong answer, so never leave one blank.
The questions come in a few forms. Many are short and direct ("Which sampling method BEST ensures…?"); others are scenarios of a few sentences; and some stems are incomplete statements that the options finish. Almost all carry a qualifier: BEST, MOST likely, FIRST, GREATEST concern, PRIMARY benefit. The qualifier is the question. Usually two or three options describe something a competent auditor might genuinely do or worry about, and the qualifier, read against the details, is what separates the best answer from the merely reasonable ones. A few questions are inverted: LEAST or EXCEPT asks you to find the one option that is not appropriate, so three of the four are correct practice.
When the options feel equally good, go back to the stem: one detail in it almost always decides the ranking.
The questions below follow the exam's domain weights: four from AI Operations, two from AI Governance and Risk and two from AI Auditing Tools and Techniques. They include short direct questions, longer scenarios and one inverted question, as the exam does.
AI Operations (46% of the exam)
AI Operations covers the life of a model after it is built: data, change, monitoring, threats and incidents. It carries nearly half the marks, and it starts where many audit programmes stop: the day a model goes live.
Question 1: A pipeline that retrains itself
A lender's credit-scoring model is retrained every month on the latest application data. Whenever a retrained version scores higher than the current model on a holdout dataset, a pipeline promotes it to production automatically. Which of the following is the IS auditor's GREATEST concern?
- A. Promotion decided on holdout accuracy
- B. The monthly retraining cadence
- C. Promotion without change approval
- D. Retraining on unvetted application data
Show the answer and explanation
Answer: C. Promotion without change approval
Promotion without change approval is the greatest concern. A retrained model is a new model, and putting it into production is a change. The pipeline bypasses the controlled process (testing, approval, a rollback plan, informing stakeholders) that AI change management shares with conventional change management, and that gate is where every other weakness in this scenario would be caught. The closest runner-up is promotion decided on holdout accuracy, a real weakness, but a narrower one: the change process is where the test criteria for a release are set and enforced, so fixing the metric alone would leave releases unreviewed.
Why the other options lose
- A. Promotion decided on holdout accuracy. Promotion decided on holdout accuracy is a real weakness for a credit model, where fairness and stability also matter. However, it is one test criterion. An approval gate is where criteria like these get defined and enforced, and fixing the metric alone would still let unreviewed models into production.
- B. The monthly retraining cadence. The monthly retraining cadence is a legitimate tuning question. However, nothing in the scenario suggests drift is outpacing it, and retraining more often through the same unapproved path would increase the exposure.
- D. Retraining on unvetted application data. Retraining on unvetted application data is a genuine data-quality risk, because bad or manipulated data flows straight into the model. However, it is one of the risks a change gate exists to catch before release. Without that gate, even well-vetted data would still produce unreviewed releases.
Domain 2, AI Operations · 2C Change Management Specific to AI · Task 12
Question 2: A metric that stays green
A bank monitors its fraud-detection model each month using precision, which has stayed above its 95% target for a year. Over the same period, losses from fraud that the model did not flag have risen steadily. Which of the following MOST likely explains why monitoring raised no alert?
- A. Recall is not monitored
- B. Too low a precision target
- C. Monthly monitoring frequency
- D. Losses tracked by finance
Show the answer and explanation
Answer: A. Recall is not monitored
Recall is not monitored, which is the most likely explanation. Precision is the share of flagged transactions that really were fraud. It says nothing about fraud the model failed to flag (false negatives), which is what recall measures. A model can keep excellent precision while missing more and more fraud, so the monitored metric was blind to exactly the failure that occurred.
Why the other options lose
- B. Too low a precision target. Too low a precision target would matter if flagged cases were the problem. However, a stricter precision target still would not register missed fraud, because precision does not measure it. The problem is which metric is monitored, not the threshold.
- C. Monthly monitoring frequency. Monthly monitoring frequency matters for fast-moving risks. However, twelve monthly readings are plenty to show a trend in any metric that captured the problem, and more frequent precision readings would show the same healthy number.
- D. Losses tracked by finance. Losses tracked by finance rather than the model owner is a reporting gap worth closing. However, it does not explain why the model's own monitoring stayed green: that monitoring measured the wrong thing, whoever sees the losses.
Domain 2, AI Operations · 2D Supervision of AI Solutions · Task 7
Question 3: An assistant that reads customer email
A retailer's customer-service assistant uses a large language model that adds product manuals and the customer's recent emails to its prompt before answering. In testing, an email containing hidden instructions caused the assistant to ignore its rules and send the customer a link to a phishing site. Which control BEST addresses this weakness?
- A. Customer multifactor authentication
- B. Encryption of emails at rest
- C. Retraining on injection examples
- D. Input and output validation
Show the answer and explanation
Answer: D. Input and output validation
Input and output validation is the best control. This is indirect prompt injection: the malicious instructions arrived through data the model ingested, not through anything the user typed. The mitigation works at both ends. Inputs are validated and constrained, for example with prompt templates, before they reach the model, and output validation blocks responses that break the rules whatever the prompt instructed.
Why the other options lose
- A. Customer multifactor authentication. Customer multifactor authentication is good practice. However, prompt injection needs no special privilege: the attacker only had to send an email, and a fully authenticated customer would still receive the malicious output.
- B. Encryption of emails at rest. Encryption of emails at rest protects stored data from unauthorized reading. However, the assistant must decrypt the email to use it, so the hidden instructions reach the model unchanged.
- C. Retraining on injection examples. Retraining on injection examples is a recognized long-term way to make a model resist injection. However, it is costly and often impractical for large language models, and it protects nothing until it is done. Input and output validation remains the effective ongoing control.
Domain 2, AI Operations · 2F Threats and Vulnerabilities Specific to AI · Task 14
Question 4: Suspected data poisoning
When data poisoning of a production AI model is suspected, the LEAST appropriate immediate action is:
- A. revoking pipeline access to the training datasets
- B. retraining the model on its existing training data
- C. sanitizing model outputs during the investigation
- D. reviewing dataset versions and hashes for tampering
Show the answer and explanation
Answer: B. retraining the model on its existing training data
Retraining the model on its existing training data is the least appropriate action. Until the poisoned records have been identified and removed, the existing dataset still contains them, so retraining now would rebuild the poisoning into a new model version. Retraining belongs to eradication, and only after the poisoned data has been isolated and removed.
Why the other options are appropriate
- A. revoking pipeline access to the training datasets. Revoking pipeline access to the training datasets is appropriate. It is a recommended containment step for data poisoning, because it stops further tampering while the incident is assessed.
- C. sanitizing model outputs during the investigation. Sanitizing model outputs during the investigation is appropriate. It is a recognized short-term measure that limits the harm poisoned data can cause while the model remains in use.
- D. reviewing dataset versions and hashes for tampering. Reviewing dataset versions and hashes for tampering is appropriate. It is how poisoning is detected and scoped, and the rest of the response depends on knowing what changed.
Domain 2, AI Operations · 2G Incident Response Management Specific to AI · Task 11
AI Governance and Risk (33% of the exam)
Governance questions test who owns what, which assessments apply and whether the organization's structures actually bind the people using AI.
Question 5: Who builds the AI inventory
Several departments are using generative AI tools, and no inventory of AI use exists. The AI governance committee asks internal audit to lead the effort to identify and catalog every AI solution in use, because audit already has access across the organization. Which is the BEST response from the chief audit executive?
- A. Let the AI or data function lead
- B. Lead it, pledging no penalties
- C. Lead it using its broad access
- D. Await an approved AI usage policy
Show the answer and explanation
Answer: A. Let the AI or data function lead
Letting the AI or data function lead is the best response. Building an AI inventory takes governance, risk, IT operations and audit working together, but ownership belongs with the part of the business accountable for AI or data management. Audit can contribute knowledge and challenge. If it ran the effort, it would later be testing the completeness of an inventory it built, which is a self-review threat.
Why the other options lose
- B. Lead it, pledging no penalties. Leading it while pledging no penalties gets one thing right: stating that the inventory is for discovery, not punishment, encourages honest disclosure. However, audit would still own an inventory it will later rely on and test, so the independence problem remains.
- C. Lead it using its broad access. Leading it using audit's broad access plays to a real strength. However, completeness is not the deciding factor. Owning the effort would impair audit's objectivity when it later tests how complete and accurate that inventory is.
- D. Await an approved AI usage policy. Awaiting an approved AI usage policy has a logic, because such a policy is often a good starting point for an inventory. However, waiting leaves unidentified AI use unmanaged in the meantime, when audit can properly contribute now to an effort that management owns.
Domain 1, AI Governance and Risk · 1B AI Governance and Program Management · Task 8
Question 6: A DPIA for a high-risk AI system
A bank in the EU uses a vendor-supplied AI system to assess the creditworthiness of loan applicants, a use the EU AI Act classifies as high risk. The compliance team completed a data protection impact assessment (DPIA) and states that it satisfies the Act's impact assessment requirement. Which of the following is the IS auditor's GREATEST concern?
- A. Vendor excluded from the DPIA
- B. No consent to automated credit scoring
- C. No fundamental rights assessment
- D. DPIA led by compliance staff
Show the answer and explanation
Answer: C. No fundamental rights assessment
No fundamental rights assessment is the greatest concern. For high-risk uses such as creditworthiness assessment, the EU AI Act requires deployers to carry out a FRIA covering the system's impact on people and society. A DPIA is similar but does not cover the range of rights a FRIA must address. It can complement a FRIA but cannot replace it, so the bank is relying on the wrong assessment to meet a legal obligation.
Why the other options lose
- A. Vendor excluded from the DPIA. A vendor excluded from the DPIA is a real third-party gap. However, widening the DPIA would still leave a DPIA, not the assessment the Act requires.
- B. No consent to automated credit scoring. No consent to automated credit scoring sounds serious. However, consent is not the only lawful basis for processing, and nothing in the scenario shows a defect there. The established gap is the missing FRIA.
- D. DPIA led by compliance staff. A DPIA led by compliance staff rather than the model owner is a question of who does the work. However, it is secondary to the fact that the organization has completed the wrong type of assessment.
Domain 1, AI Governance and Risk · 1E Leading Practices, Ethics, Regulations, and Standards for AI · Task 6
AI Auditing Tools and Techniques (21% of the exam)
The third domain works in both directions: auditing AI, and using AI to audit. The two questions below cover one of each.
Question 7: Sampling training data
Which sampling method BEST ensures that an audit sample of AI training data represents its important subgroups?
- A. Random sampling across the whole dataset
- B. Haphazard selection of records by the auditor
- C. Judgmental sampling of the highest-risk records
- D. Stratified sampling based on key parameters
Show the answer and explanation
Answer: D. Stratified sampling based on key parameters
Stratified sampling based on key parameters is the best method. Dividing the data into strata by the parameters that matter (for example, demographic group or data source) and sampling within each one guarantees that every subgroup appears in the sample, which is what a test of representativeness and bias needs.
Why the other options lose
- A. Random sampling across the whole dataset. Random sampling across the whole dataset is statistically sound and avoids selection bias. However, small subgroups can be missed or underrepresented by chance, so it does not ensure they are represented.
- B. Haphazard selection of records by the auditor. Haphazard selection of records by the auditor is a recognized sampling approach. However, it follows no structure that guarantees coverage of subgroups, and it is open to the auditor's unconscious bias.
- C. Judgmental sampling of the highest-risk records. Judgmental sampling of the highest-risk records is useful for targeting likely problems. However, it deliberately skews the sample toward risk, so it cannot show whether the data as a whole represents its subgroups.
Domain 3, AI Auditing Tools and Techniques · 3B Audit Testing and Sampling Methodologies · Task 18
Question 8: A generative AI summary as evidence
An IS auditor uses a generative AI tool to summarize 60 stakeholder interview transcripts and highlight statements that contradict each other. The summary identifies several contradictions that would support an audit finding. What should the auditor do FIRST?
- A. Document the prompts used
- B. Verify against transcripts
- C. Obtain vendor assurance
- D. Ask interviewees about them
Show the answer and explanation
Answer: B. Verify against transcripts
Verifying the contradictions against the transcripts comes first. A generative AI output is not audit evidence until the auditor has validated it, and generative models can produce statements that are not grounded in their input. A finding built on a hallucinated contradiction would be unsupported, so each flagged contradiction must be checked against the source transcripts before anything else.
Why the other options lose
- A. Document the prompts used. Documenting the prompts used supports reperformance and belongs in the audit file. However, it documents how the output was produced and does not show that the output is accurate.
- C. Obtain vendor assurance. Vendor assurance can support reliance on the tool in general. However, it cannot confirm that this particular summary of these transcripts is correct.
- D. Ask interviewees about them. Asking interviewees about the contradictions may be appropriate later. However, doing it before confirming the contradictions exist risks confronting people with statements the transcripts do not contain.
Domain 3, AI Auditing Tools and Techniques · 3C Audit Evidence Collection Techniques · Task 1
What these questions have in common
Look back over the eight and the same few ideas decide most of them.
- A retrained model is a change. Anything that alters what runs in production needs the change process, however automated the pipeline (question 1).
- The control has to be able to see the failure. A metric that cannot register missed fraud, or encryption that the application must undo to work, looks like a control without doing the job (questions 2 and 3).
- Contain and investigate before you rebuild. Retraining on data that may still be poisoned rebuilds the problem; removing the bad data comes first (question 4).
- Independence decides who should own the work. Audit contributes to management's AI inventory; it does not build what it will later test (question 5).
- A similar assessment is not the required one. A DPIA and a fundamental rights impact assessment overlap, but one cannot stand in for the other (question 6).
- Evidence has to be fit for the conclusion. A sample built to cover subgroups can show representativeness where a random one might miss them, and a generative AI summary is not evidence until it has been checked against its source (questions 7 and 8).
None of these is a fact you could memorise from a single question. Each is a way of ranking options that are all defensible, which is what the exam is actually measuring.
Why memorised answers do not carry over
Recognising a question you have already seen is a different skill from ranking four plausible options you have not. Change one scenario detail (the control already in place, the stage of the incident, who owns the process) and the best answer moves. That is why collections of remembered live questions, often sold as exam dumps, prepare candidates so badly for a scenario-based exam, besides breaching ISACA's candidate agreement. The longer case is in our AAIA explainer.
What does carry over is the reasoning in the explanations above: naming the qualifier, finding the detail that decides it, and being able to say in one sentence why the runner-up loses.
How to practise for the AAIA
Work through questions in timed sets, and weight your practice toward AI Operations the way the exam does. For every question, including the ones you got right, check that you could have explained why each other option loses; a lucky pick and a reasoned one score the same in practice and very differently on exam day. When you are consistently comfortable, move to full-length, timed mocks, because pacing 90 questions in 150 minutes is its own skill.
CertPrepX has 900+ AAIA practice questions and 1,200+ flashcards across all three domains, plus full-length mocks scored on the 200–800 scale. See the AAIA exam prep guide or start free practice.
Sources
- ISACA, Advanced in AI Audit (AAIA) credential
- ISACA, AAIA Exam Content Outline
- ISACA, Exam Candidate Guides
- EU AI Act, Regulation (EU) 2024/1689, Article 27 and Annex III
Frequently asked questions
Are these real AAIA exam questions? No. They are original questions written to follow the exam's published structure and domain weights. ISACA keeps live exam content confidential, and reproducing it breaches the candidate agreement.
Does ISACA publish official AAIA practice questions? Yes. ISACA sells an official Questions, Answers and Explanations (QAE) database for the AAIA, and it is worth using alongside the AAIA Review Manual.
How many questions are on the AAIA exam? 90 multiple-choice questions in 150 minutes. Some are unscored research items, and ISACA does not identify which.
What score do I need to pass the AAIA? 450 on the 200–800 scaled score. ISACA does not publish how many correct answers that takes on a given exam form.