Managing Data Quality as a First-Order Operational Risk
Data quality spent decades as a hygiene topic, funded reluctantly and noticed in post-mortems. Two things promoted it to the risk register. Straight-through processing removed the human checkpoints that once intercepted a stale instruction or a mangled counterparty name, so each remaining defect travels farther before anything catches it, and AI industrializes whatever it is fed. And the regulators have stopped taking the industry's word for its data: EMIR Refit, SFTR and MiFIR submissions are scored by ESMA against the whole population, DORA grades the register of information, and CSDR prices instruction defects on a monthly invoice.
The report gives the promoted risk the treatment other first-order risks receive. It decomposes the register entry into six data domains, each with its characteristic defect, propagation channel, loss landing, Basel event type and the KRI that sees it; anchors those KRIs to the reference points the graders now publish; maps six controls to the regimes that demand them; and shows why the loss dataset stays a supervised artifact whatever the capital multiplier does. One finding carries the argument: the EBA's new loss taxonomy has fifteen attributes a bank must tag on every loss, and none of them is a data defect.
It closes on the two adjacencies that make data quality more than the sum of its defects: the SSI database as a fraud surface, and AI as both amplifier and control.
Selected Conclusions
• The risk becomes manageable when named. Six data domains carry distinct defects, propagation channels, and loss destinations, and because the EBA's new loss taxonomy carries fifteen attributes and none for a data defect, the root-cause tag the firm adds itself is the only way the exposure becomes visible.
• Regulation grades the data daily. EMIR Refit, SFTR, and MiFIR transaction reporting are scored by ESMA against the whole population, DORA grades the register of information, CSDR prices instruction defects monthly, and the US securities lending regime joins them in 2028; the graders publish their gradebooks.
• Loss data is capital-relevant data. Under the standardized approach the loss dataset is collected, reported, and disclosed whatever the multiplier does, which makes it a regulated dataset whose own quality determines both capital calibration and the firm's ability to see this risk at all.
Subscribers to the Journal may download the full report, including 8 Conclusions and 18 Recommendations and Action Items.