KYC and AML Software Development in 2026: What to Build, What to Buy, and What the Regulator Will Ask For
Where Compliance Software Actually Fails
KYC and AML software development is the engineering of the systems a regulated business uses to verify who its customers are, score the risk each one carries, screen them against sanctions and politically exposed person (PEP) lists, watch their transactions for suspicious patterns, and file the reports the regulator expects. KYC (know your customer) covers onboarding and ongoing due diligence; AML (anti-money laundering) covers monitoring, investigation, and reporting. The two share a customer record, a risk model, and an audit trail, which is why they get built together. A first version with identity verification, screening, and basic monitoring costs $50,000 to $150,000; a platform with a rules engine, a case manager, and regulatory filing runs $200,000 to $600,000; multi-jurisdiction enterprise builds start at $1 million.
The enforcement record shows what fails. In October 2024 TD Bank pleaded guilty to Bank Secrecy Act violations after leaving about 92% of its transaction volume, roughly $18.3 trillion, unmonitored for six years; the combined resolution was about $3.09 billion plus an asset cap. In August 2026 FinCEN assessed a $125 million penalty against UBS Financial Services, the largest ever on a broker-dealer, over more than 50,000 unmonitored foreign-currency wires. Both firms had a compliance program on paper. The software behind it did not cover the volume, and nothing in it showed the gap.
This guide follows the order a build meets the problem: what changed in the rules over the last 18 months, which parts to buy and which to build, how identity verification, screening, and transaction monitoring work at the engineering level, where machine learning is now allowed, what the audit trail has to hold, and what each stage costs. It includes the model-training pipeline we published as open source, and the reasons it stops where it stops.
What KYC and AML Software Does, Component by Component
Seven components cover almost every regulated onboarding and monitoring flow. Vendors package them differently, but these boundaries are the ones that matter for build decisions, because each has a different data source, failure mode, and regulator expectation.
- Identity verification (IDV). Document capture, a liveness check, a biometric comparison between selfie and document photo, and data checks against issuers or bureaus. In the US this satisfies the Customer Identification Program (CIP) requirement to form a reasonable belief about who the customer is.
- Customer due diligence and risk scoring. Purpose of account, source of funds, expected activity, beneficial ownership for legal entities, then a score that decides between standard, simplified, and enhanced due diligence.
- Sanctions and PEP screening. Matching of names, dates of birth, nationalities, and identifiers against the US Office of Foreign Assets Control (OFAC), EU, UN, and UK HM Treasury lists plus PEP and adverse media data, at onboarding and again whenever a list changes.
- Transaction monitoring. Rules and models that inspect payments, trades, and account events for structuring, rapid movement, unusual counterparties, and departures from the customer’s expected profile, then raise alerts.
- Case management. The investigator’s workbench: alert queue, customer view, evidence capture, escalation, four-eyes approval, and the decision record.
- Regulatory reporting. Generation and submission of suspicious activity reports (SARs) and currency transaction reports (CTRs) in each regulator’s format, with deadlines tracked from the date of detection.
- Audit trail and record retention. An immutable log of who saw what, who decided what, and which version of each list, rule, and model was in force at the time.
Trading platforms, payment gateways, neobanks, and crypto exchanges all need the same seven. The differences are volume, which regulator sits on top, and how much of the stack may be outsourced. Our guides on trading platform development and prop firm technology stacks cover the wider fintech build; this one goes deep on the compliance layer.
What Changed in the Rules in 2025 and 2026
FinCEN withdrew its 2024 proposal and published a new proposed rule on April 10, 2026. It keeps the “effective, risk-based, and reasonably designed” standard, requires a documented risk assessment, and states that institutions which responsibly experiment with machine learning, generative AI, digital identity, and APIs will not incur additional supervisory risk. No final rule yet.
FinCEN’s final rule of August 14, 2026 removed US companies and US persons from Corporate Transparency Act reporting; only foreign reporting companies still file. Onboarding flows now collect beneficial ownership for domestic entities directly under the customer due diligence (CDD) rule.
An interagency order with FinCEN concurrence lets banks obtain taxpayer identification numbers from a reliable third-party source instead of the customer, for all account types, provided the full TIN is held before account opening. One field leaves the form; one data source joins the pipeline.
The Anti-Money Laundering Authority began operations on July 1, 2025 and will select up to 40 entities for direct supervision during 2027. The AML Regulation (EU) 2024/1624 applies from July 10, 2027 with one set of due diligence rules across the EU. EU-facing products need a jurisdiction layer that can switch when it lands.
Section 199 of the Economic Crime and Corporate Transparency Act makes large organisations liable when an associated person commits fraud for their benefit, with “reasonable procedures” as the only defence. The 2026 amendment regulations add pooled-account due diligence duties and move crypto enhanced due diligence to February 2027.
The seventh targeted update on virtual assets reports that 83% of surveyed jurisdictions have Travel Rule legislation, up from 73% a year earlier, and that most identified illicit on-chain activity now involves stablecoins. Crypto transfers need originator and beneficiary data exchange built in.
NIST SP 800-63A-4 defines remote identity proofing at Identity Assurance Level 2: strong evidence, validation against the issuer, and verification through facial comparison or an automated biometric match with presentation attack detection. Vendors now quote conformance to it; ask for it by name.
The Federal Reserve, OCC, and FDIC issued SR 26-2, which replaces SR 11-7 with a proportional, risk-based approach aimed mainly at banks above $30 billion in assets. Generative and agentic AI are explicitly outside its scope.
Build, Buy, or Assemble: Where Each Component Lands
Almost nobody builds all seven components, and almost nobody should buy all seven. The pattern that works: buy where the value is in the data or the certification, build where the value is in fit to your products, your risk appetite, and your investigators.
The assembled stack. The typical result is two or three vendor APIs (identity verification, list data, sometimes a filing connector) under a custom orchestration layer that owns the customer record, the risk score, the rules, the case workflow, and the audit log. The orchestration layer is the product, and it is the part the examiner reads, because it holds the decisions.
Where a wrong choice shows up. Buying a monitoring system whose rules cannot see your product’s fields produces the TD Bank pattern: transactions flow, the system runs, and nothing it can see is wrong. Building identity verification in house produces a slower, less accurate version of a commodity with none of the certification. The failure modes point in opposite directions, which is why the split above holds across firms of very different sizes.
Identity Verification and Onboarding: Where Customers Drop Off
In Signicat’s Battle to Onboard survey of 7,600 European consumers (2022), 68% had abandoned a financial application in the previous year, and the average applicant gave up after 18 minutes 53 seconds. The two most cited reasons were the time it took and the amount of information requested. Every extra field and every second-attempt document capture has a measurable conversion cost.
Under SP 800-63A-4, remote proofing at Identity Assurance Level 2 means strong evidence (a government ID with security features), validation against the issuer where possible, and verification that the applicant is the person on the document, through facial comparison or an automated biometric match with presentation attack detection. The build decision is which path each segment gets, and when to step up.
A well-built flow collects the minimum for the lowest-risk segment and steps up on signal: a mismatch between stated and detected location, a document from a high-risk jurisdiction, an ownership structure with more than two layers, a PEP hit. Each step-up is logged with its trigger, so the examiner can see why one customer received enhanced due diligence and another did not.
The Wolfsberg Group’s guidance on digital customer lifecycle risk management describes a trigger-based approach to ongoing due diligence, with a risk profile that is continually reassessed rather than reviewed every one, three, or five years. In engineering terms the score is recomputed on events (a new beneficial owner, a screening hit, a pattern shift) and a review task is created only when it crosses a threshold.
Sanctions and PEP Screening: The Matching Problem
Screening looks simple from outside: take a name, compare it with a list, block on a match. The difficulty is that names are ambiguous, lists change daily, and the two error types cost wildly different amounts. A missed sanctioned party is a strict-liability violation; a false hit costs an investigator twenty minutes and a customer a delayed payment. The Bank Policy Institute’s 2018 study of 19 banks found a median true OFAC match rate of 0.00004% of alerts. The system exists for the one case in millions and has to stay tolerable for the rest.
How OFAC itself matches. The Treasury’s Sanctions List Search uses fuzzy logic, a combination of Jaro-Winkler string similarity and Soundex phonetic matching, with a user-set minimum score. OFAC’s FAQs state that each firm sets its own thresholds according to its own risk assessment. The regulator gives you the algorithm family and leaves the calibration to you, so the calibration has to be documented.
What goes wrong. OFAC’s Framework for Compliance Commitments lists screening software faults as a root cause in enforcement cases: lists not updated, alternative spellings not caught, SWIFT business identifier codes not screened. The 50% rule adds a structural dimension: an entity owned 50% or more in aggregate by blocked persons is blocked even if it appears on no list, so for legal-entity customers ownership has to be resolved and each owner screened.
What a good matching engine does. It normalises scripts and transliterations before comparing, since Cyrillic, Arabic, and Chinese names arrive in several Latin spellings. It scores on secondary identifiers, so a common name with a matching date of birth and nationality outranks an unusual name with nothing else in common. It applies different thresholds by list and by customer risk. It versions every list, so an investigation can show which version a customer was screened against on a given day. And it rescreens the whole base on every list update, with the delta reported, so the compliance officer knows the morning after a designation how many existing customers were affected.
Transaction Monitoring: Rules, Models, and What the Open-Source Core Looks Like
Transaction monitoring generates most of the alerts and most of the cost. McKinsey’s 2017 analysis of rule-based systems found that up to 90% of alerts can be false positives, and the Bank Policy Institute’s 19-bank sample showed roughly 16 million alerts producing about 640,000 SARs, of which a median 4% drew law-enforcement follow-up. The LexisNexis True Cost of Financial Crime Compliance study (2024) puts annual compliance spend in the US and Canada at $61 billion, with labour the largest category. Much of that labour is people clearing alerts a better system would not have raised.
Rules first. Every monitoring system starts with rules because the regulator expects specific typologies covered: structuring below reporting thresholds, rapid movement of funds in and out, round-amount transfers to high-risk jurisdictions, dormant accounts waking up, activity inconsistent with stated purpose. Rules are explainable by construction; their weakness is that they are tuned by hand, they age, and each fires independently, which is where the alert volume comes from.
Models on top. A supervised model trained on past alerts and their outcomes learns which feature combinations actually preceded a SAR, and it scores every rule alert for priority so investigators work the likely cases first. Graph features add what rules cannot see: a single account’s transactions look ordinary while the network it sits in (many senders, one collector, funds moving in a cycle) does not.
What we open-sourced. We published the training side of this as plus8soft/AML on GitHub, under Apache 2.0. It is an offline Python pipeline that takes a transaction dataset in a plain tabular schema or a graph schema with sender and receiver fields, builds features (including in-degree, out-degree, PageRank, and cumulative totals from the transaction network via NetworkX), filters correlated features, trains Random Forest, XGBoost, and LightGBM in competition, picks the winner by ROC-AUC with PR-AUC as the tiebreaker, and saves the best model with its metrics. The PR-AUC tiebreak matters: with suspicious transactions at a fraction of a percent of the data, ROC-AUC alone flatters models that are good at the easy negatives.
Where it stops, and why. The repository does not include real-time scoring, streaming ingestion, sanctions screening, case management, or SAR and CTR generation. Those parts carry the regulatory risk: they need integration with your ledger, your list vendor, your investigators, and your filing obligations, and they cannot be generic. The model is the reusable part; the controls around it are the product.
“We have posted the training pipeline on GitHub because the model is the least distinctive component of the AML stack. Anyone with labeled data can train a classifier using gradient boosting. The decisive factor in whether the system passes validation is everything surrounding the model: which features it was allowed to consider, how and by whom the threshold was set, what happened to the alerts it generated, and whether it’s possible to reproduce the assessment for a transaction made two years ago. This level is specific to each company, and it is precisely here that we focus our efforts.”
Timur Iusubaliev, COO, Plus8Soft
Security of the pipeline itself, from secrets handling to dependency scanning on the model code, follows the practices in our DevSecOps guide.
AI in AML: What Regulators Now Allow and What They Still Require
The proposed program rule states that machine learning and large language models have potential to strengthen AML programs, names generative AI, digital identity, blockchain analytics, and APIs as approaches, and says institutions that responsibly experiment with them will not incur additional risk of a significant supervisory action. It also records FinCEN’s concern that model risk management as applied to AML is overly burdensome.
Alert prioritisation is the proven use: a model that ranks rule alerts by likelihood of escalation lets a fixed investigation team clear the queue in likelihood order, and the reduction in wasted reviews is measurable within a quarter. Network detection is the second: graph features find mule networks and layering patterns no single-account rule can express. Both keep the rules in place.
Sanctions decisions remain binary and strict-liability; a model can rank a screening hit for review but cannot clear it. SAR narratives can be drafted by a language model from the case record, but the filing is a legal statement by the institution, and the investigator signs it. The Wolfsberg Group’s second statement on effective monitoring (August 2025) sets out the validation and explainability expectations for firms moving rules to models, and it assumes a human decision at the end.
Under SR 26-2 the validation burden is proportional to size and materiality, but the examiner will still ask for training data lineage, the feature list with a rationale for each, the threshold-setting record, back-testing against known cases, and a way to reproduce any historical score. Build those into the pipeline from the first version; retrofitting lineage onto a model that has run for a year is the most expensive compliance project we see. For the wider picture, including the EU position, see our AI in fintech and EU AI Act guides.
Case Management, Reporting, and the Audit Trail
The regulator examines decisions, and it examines them through the case record. FinCEN’s Year in Review for fiscal 2025 counts 4.8 million SARs and 21.5 million CTRs filed by about 335,000 registered filers. Each report has a deadline, a supporting file, and a retention clock, and each is a place a build can fail quietly.
Deadlines run from detection, not from decision. Under 31 CFR 1020.320 a SAR is due no later than 30 calendar days after initial detection, extendable by 30 days when no suspect has been identified, and never later than 60. Initial detection is an alert timestamp, so the case manager has to carry it from alert to case to filing and show it to the compliance officer as a countdown. The New York DFS action against Block’s Cash App in April 2025 cited an alert backlog; the Canaccord Genuity resolution in March 2026 involved roughly 160 SARs never filed. Both are queue-management failures before they are judgement failures.
Records are kept five years, and the version matters. 31 CFR 1010.430 requires records under the chapter to be retained for five years; SAR copies and supporting documentation are kept five years from filing. An examiner reconstructing a 2023 decision in 2027 needs the customer profile as it was, the list version screened against, the rule set in force, and the model version and score. That is an event-sourced design: append-only records of every input and decision, with current state derived from them, rather than a customer table that overwrites history.
The investigator’s screen decides throughput. A case view with the alert, the customer’s profile and risk history, related parties, prior cases, and the transactions in context, plus evidence capture and a structured decision form, turns a forty-minute review into a ten-minute one. A four-eyes step on SAR decisions, with the reviewer’s identity and timestamp in the record, is the control examiners look for first. Narrative drafting from the structured case record is the safest automation in the stack because the human decision sits after it.
What KYC and AML Software Costs to Build
The ranges below reflect what published agency estimates converge on and what we see in scoping. They assume vendor APIs for identity verification and list data, with the orchestration, rules, case management, and audit layer built. Data licences and per-check fees are on top and scale with volume.
What moves a project between rows. The number of jurisdictions is the largest multiplier: each adds a filing format, a list set, and a due diligence variant. Transaction volume is second: above a few million events a day, monitoring moves from batch to streaming and the case manager needs bulk operations. Legal-entity customers with layered ownership add beneficial ownership resolution and per-owner screening. The 20% to 30% annual maintenance figure is consistent across published agency estimates and matches our experience: rules need tuning as products change, lists and formats change on the regulator’s schedule, and models need retraining as the alert population shifts.
The cost of not building it. The LexisNexis global study estimated financial crime compliance costs at $206 billion a year, with 98% of firms reporting rising costs, most of it headcount. A monitoring and case system that halves false positives pays for its own build within the first year at any firm with more than a handful of investigators. For a scoped estimate on your team and duration, use our project cost calculator.
Choosing a Development Partner for Compliance Software
A team that has been through an examination describes the event log, the versioning of lists and rules, and the reconstruction of historical decisions before it describes screens.
Screening scores, rule parameters, and model cut-offs all have to be defensible. The partner should have a process for calibration runs, for recording who approved a threshold and on what evidence, and for re-running calibration when the customer base changes.
Identity verification and list providers get swapped. A partner who has built the orchestration layer to hold two vendors at once, and migrated a customer base between them without a rescreening gap, has done the hard version of the job.
Features that proxy for protected characteristics, data the customer did not consent to, and signals that cannot be explained to an examiner have to be excluded by design. A partner who can list what they leave out understands the regime better than one who lists what they put in. Our dedicated fintech engineering teams work under this discipline as standard, as on the compliance-facing builds for FundingPips and a top-3 global forex and CFD broker.
Common Mistakes in KYC and AML Builds
A rules engine tuned for card payments watching a new wire product, a trading platform’s internal transfers excluded from the feed, a new currency added without a threshold. Coverage gaps are the TD Bank finding at every scale. Every new product or rail needs a documented monitoring coverage decision before launch.
Alerts accumulate, investigators work the newest, and the 30-day clock runs out on the oldest. The queue needs age-ordered allocation, a countdown from detection date, and an escalation when any alert approaches the deadline.
Lists change daily. A customer clear on Monday can be designated on Tuesday. The system has to rescreen the entire base on every list update and report the delta, and it has to screen beneficial owners and counterparties, not only account holders.
Overwriting the risk score, the address, or the ownership structure in place destroys the evidence of what the firm knew at decision time. Append-only records with derived current state are the only design that survives a five-year lookback.
A model in production with no record of its training data, feature rationale, or threshold approval fails validation on the first question. The training pipeline has to produce the lineage as artefacts alongside the model, or it will never exist.
Filing formats, due diligence thresholds, list sets, and retention periods differ by regulator. A jurisdiction layer that switches these by customer and by entity costs little in the first version and a great deal to retrofit, particularly with the EU single rulebook arriving in July 2027.
KYC and AML Software Development: Frequently Asked Questions
KYC software verifies identity and assesses customer risk at onboarding and on an ongoing basis. AML software monitors transactions, manages investigations, and files regulatory reports. They share the customer record, the risk score, and the audit trail, so most builds treat them as one platform with two front ends: the customer’s and the investigator’s.
An onboarding and screening MVP on vendor APIs costs $50,000 to $150,000 over three to six months. Adding a rules engine, case management, and one-jurisdiction filing runs $200,000 to $600,000 over six to twelve months. Multi-jurisdiction enterprise platforms with machine-learning monitoring start at $1 million. Budget 20% to 30% of build cost per year for maintenance.
Buy identity verification and sanctions and PEP data, where the value is in the vendor’s coverage, certifications, and research. Build the orchestration layer, risk scoring, monitoring rules and models, case workflow, and audit trail, where the value is in fit to your products and in the decision record the examiner reads.
In the US, the Bank Secrecy Act and FinCEN’s rules (CIP, CDD, SAR and CTR filing, five-year retention), with a new AML program rule proposed in April 2026 and not yet final. In the EU, the AML Regulation applies from July 2027 under the new AMLA authority. In the UK, the Money Laundering Regulations as amended in 2026 and the failure-to-prevent-fraud offence in force since September 2025. Crypto transfers fall under the FATF Travel Rule in 83% of surveyed jurisdictions.
Models run alongside rules. Rules provide the typology coverage the examiner expects and are explainable by construction; models prioritise the alerts rules generate and detect network patterns rules cannot express. FinCEN’s April 2026 proposal states that firms which responsibly experiment with machine learning will not incur additional supervisory risk, and SR 26-2 makes validation proportional to materiality.
Under 31 CFR 1020.320, no later than 30 calendar days after initial detection, with a further 30 days when no suspect has been identified, and never more than 60 in total. The clock starts at the alert, so the case management system has to carry the detection date through to filing.
No. It is the offline model-training pipeline: data loading in plain or graph schema, feature engineering with NetworkX graph features, training of Random Forest, XGBoost, and LightGBM, and selection by ROC-AUC and PR-AUC. Real-time scoring, streaming ingestion, sanctions screening, case management, and SAR and CTR generation are built per engagement, because they depend on your ledger, your vendors, and your filing obligations.