Key takeaways
- Start with one bounded diligence process, not an open-ended promise to automate the complete deal lifecycle.
- Define the source set, playbook, output schema, reviewer decisions, and acceptance criteria before choosing agents or models.
- Keep materiality, escalation, and investment judgment with accountable professionals.
- Evaluate the complete system, including retrieval, tools, permissions, review, exports, and failure handling, not the model in isolation.
- Measure accepted, edited, rejected, unsupported, and missed findings during a shadow run before go-live.
M&A diligence combines repetitive preparation with consequential judgment. Teams classify files, extract terms, compare evidence, track missing support, prepare issue lists, and draft memo sections. AI can help with that preparation, but the deal team remains responsible for deciding what is material, what needs escalation, and what belongs in a final deliverable.
This guide was revised on 20 August 2026 to remove unsupported market statistics, hypothetical incidents presented as facts, categorical competitor claims, and unverified legal conclusions. It describes an implementation and evaluation method, not a guarantee that AI will identify every risk or make a diligence process legally defensible.
What diligence automation should and should not do
A useful system prepares reviewable work. It can classify documents, extract specified facts, compare clauses or schedules, flag missing material, assemble source-linked issue tables, and draft language in an approved structure. It should expose uncertainty and exceptions rather than present every generated statement as a conclusion.
The system should not determine investment materiality, provide legal advice, approve a transaction, or silently move generated findings into a client or investment-committee deliverable. Those decisions remain with the professionals accountable for the work.
Choose a bounded first process
Do not begin with the entire VDR. Choose one document family, diligence question set, or recurring output that reviewers can assess completely. Examples include change-of-control extraction from one contract class, a missing-document queue for one workstream, customer concentration support from approved schedules, or a first-pass issue table for one diligence playbook.
A bounded scope makes failure visible. The team can create a representative test set, identify authoritative sources, document known edge cases, and decide what level of reviewer correction is acceptable before the system is used on live work.
Define the output contract before the agent architecture
The output contract states what the system must produce and what evidence must accompany it. For an issue table, that may include document, clause, extracted fact, source location, risk category, reason, confidence or review flag, open question, owner, reviewer decision, and export status.
The contract should also define abstention. If the source is missing, ambiguous, contradictory, unreadable, or outside the approved scope, the correct output may be an exception for human review rather than a generated answer.
Design retrieval for the document, not only the query
Contracts and financial materials can contain definitions, amendments, schedules, tables, and cross-references that change the meaning of an isolated passage. A production system should preserve document structure and provide a way to gather connected evidence. The exact technique may involve semantic and keyword retrieval, layout-aware parsing, table extraction, metadata filters, iterative search, or task decomposition.
No retrieval design should be declared reliable from architecture alone. Test it against representative questions and known source relationships. Record whether the system found the authoritative passage, included relevant qualifications, and avoided unsupported synthesis.
Keep agents inside explicit permissions
An agent should receive only the sources and tools needed for its task. Read-only access, scoped repositories, role-aware retrieval, constrained actions, deterministic validation, and approval before consequential changes reduce the impact of mistakes or manipulated input.
This reflects the risk behind OWASP's guidance on excessive agency: excessive functionality, permissions, or autonomy can turn a model error or prompt injection into a damaging action. The safe default for diligence preparation is limited authority and visible human approval.
Make the review decision part of the product
A review screen should make it easy to inspect the proposed finding, supporting evidence, assumptions, exception reason, and relevant document context. The reviewer should be able to approve, edit, reject, escalate, or request more support without rebuilding the work outside the system.
An audit-oriented operating record can retain output versions, source references, reviewer actions, and timestamps. Whether a record is immutable, legally sufficient, or admissible depends on the actual technical controls, contracts, retention policy, and applicable law. Those properties should be verified, not assumed from a product label.
Use a shadow run before go-live
Run the new system beside representative historical or live work without allowing it to control the final deliverable. Use the same files, questions, reviewers, and deadline. Compare the output against the approved human result and inspect errors rather than relying on one headline accuracy score.
Track accepted, edited, rejected, unsupported, duplicated, and missed findings. Measure reviewer time, source-navigation time, exception quality, output-format fit, and failure recovery. OpenAI's evaluation guidance similarly emphasizes that claims depend on the task, harness, tools, scoring rules, and review procedure used to test them.
Questions to ask any diligence AI provider
- Which exact task and source set was each capability claim evaluated against?
- How are cross-references, tables, amendments, duplicate files, missing support, and conflicting evidence handled?
- What can each agent read, write, call, export, or change without approval?
- What evidence, output versions, reviewer actions, and failure events are retained, and for how long?
- Where do the application, documents, retrieval layer, model calls, logs, and backups run?
- Can the provider reproduce a representative evaluation and show the accepted, edited, rejected, and missed cases?
Where Dotnitron and Underlying fit
Dotnitron's Deal Delivery AI Production Sprint maps and builds one analyst-heavy process in 6 to 8 weeks, then uses a shadow run and agreed acceptance criteria to support a go-live decision. The engagement is implementation-led and adapts the system to the firm's sources, playbook, tools, permissions, review path, and deliverable.
Underlying is Dotnitron's early-access platform direction for source-linked document analysis, mandatory human review, and audit-oriented records. It is not included in Dotnitron's production-client count. No claim is made that it eliminates diligence risk, finds every issue, or creates legal defensibility by itself.
Primary guidance referenced
NIST AI Risk Management Framework and Generative AI Profile: https://www.nist.gov/itl/ai-risk-management-framework
OWASP guidance on excessive agency: https://owasp.org/www-project-top-10-for-large-language-model-applications/2_0_vulns/LLM06_ExcessiveAgency.html
OpenAI, trustworthy evaluation foundations: https://openai.com/index/trustworthy-third-party-evaluations-foundations/
Use this guide
Turn the article into a working session.
Pick one workflow from the article and map it against your own team. Write down the input sources, current manual steps, reviewer decisions, output format, and the metric that would prove the workflow is worth automating.
- What work should agents prepare before a human reviews it?
- Which documents, data sources, tools, or approved system connections would the workflow need?
- What output would make a reviewer say, this saves real time?