Test of Design and Test of Effectiveness workflows are repetitive, but they are not low-stakes. The reviewer needs to know whether the control is designed to address the stated risk and whether it operated consistently during the review period.
AI can support this work when it is treated as a drafting layer. It should assemble the control context, read the client documentation, inspect the evidence, draft assessment notes, and flag exceptions. It should not replace the reviewer’s conclusion.
What AI should read for ToD
- Control objective and risk statement.
- Control owner and frequency.
- Policy and procedure language.
- System configuration or workflow documentation.
- Prior-year or prior-period workpaper notes when available.
What AI should read for ToE
- Evidence samples and populations.
- Approval records, tickets, logs, exports, or screenshots.
- Date ranges and period coverage.
- Exception criteria and firm review guidance.
The right output is a reviewer queue
A useful ToD / ToE workflow does not just generate paragraphs. It creates a reviewer queue with draft assessment text, evidence references, potential exceptions, missing support, and confidence notes. The reviewer can then approve, edit, or reject the draft.
How to evaluate quality
Quality should be measured by reviewer edit rate, exception catch rate, source traceability, and time saved per control. Accuracy matters, but in advisory delivery the practical question is whether senior reviewers spend less time checking mechanical work and more time applying judgment.
Security matters because evidence is sensitive
ToD / ToE evidence often includes access exports, tickets, system configurations, and screenshots. That is why the workflow needs role-based access, data isolation, audit logs, and client-approved deployment choices.
A reviewer-ready ToD / ToE output
The useful artifact is not a long AI-written paragraph. It is a structured note a reviewer can inspect quickly: control objective, control owner, frequency, evidence received, period tested, design assessment, operating effectiveness assessment, exceptions, missing support, source references, and reviewer decision.
For ToD, the system should show whether the described control addresses the risk, whether the owner and frequency are clear, whether the procedure is specific enough to operate, and whether the design language matches the evidence standard. For ToE, the system should show whether the evidence covers the period, sample, control activity, performer, reviewer, and exception criteria.
Common failure modes to catch
- The evidence proves an activity happened, but not that it happened for the required population.
- The screenshot is dated outside the review period or does not show the relevant user, system, ticket, or approval field.
- The control narrative says one person reviews the activity, while the evidence shows a different owner or no clear reviewer.
- The system generated a confident note even though the evidence is partial, stale, or missing the fields needed for sign-off.
A good workflow should surface those issues before senior review. That is the practical value: fewer avoidable review loops and a clearer basis for asking the client for better evidence.
What the first pilot should prove
Start with a narrow set of controls and representative evidence. Measure whether the AI-assisted queue reduces preparation time, preserves source traceability, catches obvious exceptions, and produces notes reviewers are willing to edit rather than rewrite. If reviewers rewrite most drafts, the workflow is not ready to expand.
Bring the real review template into the pilot. The system should learn the firm's language for design, operating effectiveness, exception rationale, and follow-up requests. That is what turns AI support into advisory delivery support instead of generic summarization.
Research notes and sources
- KPMG’s Clara announcement emphasizes AI agents for auditors while maintaining a human-in-the-loop audit experience: https://kpmg.com/us/en/media/news/kpmg-clara-smart-audit-platform.html
- DataSnipper describes audit AI agents as supporting repetitive work while keeping auditors in control with judgment and sign-off: https://www.datasnipper.com/resources/excel-agents-how-ai-agents-help-internal-audit-teams
- AICPA’s SOC 2 resources describe the criteria used to evaluate design and operating effectiveness of controls: https://www.aicpa-cima.com/resources/download/2017-trust-services-criteria-with-revised-points-of-focus-2022
Use this guide
Turn the article into a working session.
Pick one workflow from the article and map it against your own team. Write down the input sources, current manual steps, reviewer decisions, output format, and the metric that would prove the workflow is worth automating.
- What work should agents prepare before a human reviews it?
- Which documents, data sources, tools, or approved system connections would the workflow need?
- What output would make a reviewer say, this saves real time?
