September 15, 2026

Reliability as a Competitive Differentiator in Payment Infrastructure

Omar AlHalabi
Omar AlHalabi
Share Article
Reliability as a Competitive Differentiator in Payment Infrastructure

In October 2014, one of the world's major real-time settlement systems was disrupted for nine hours. The Bank of England later found that the root cause was defects introduced during earlier functionality changes to its Real-Time Gross Settlement system. Payments were delayed and operating hours had to be extended, although all submitted transactions were ultimately settled that day. Nobody was hacked. No fraud was involved. A failure within critical payment infrastructure was enough to disrupt a system expected simply to work (Bank of England, 2015). The incident mattered because it exposed how much of the financial system depends on infrastructure that is expected simply to work.

A few years later, one of the world's largest payment networks had its own version of the same story: a rare partial failure in a switch at Visa’s primary data center prevented its backup switch from activating and interfered with synchronization with the secondary site. During the disruption, several million payment transactions failed to process correctly over roughly ten hours, while shoppers and shopkeepers found out, in real time, what happens when the rails underneath “tap to pay” stop working (Visa, 2018). The incident was a reminder that resilience is not simply about having backup systems in place. Those systems also need to activate as intended, stay synchronized and support a clean recovery when something goes wrong.

Neither incident was a story about cybercriminals. They're stories about what happens when payment infrastructure, the actual wiring underneath the fraud controls and the compliance layer, stops being boring and starts being visible.

Why payments feel outages more than most industries

In many industries, a short software outage can remain largely contained without immediately affecting the end customer. Payments are less forgiving because payments run on a clock.

The Single Euro Payments Area (SEPA) Instant Credit Transfer scheme is designed to make funds available to the recipient within seconds, with round-the-clock availability (European Payments Council, 2023). The FedNow Service operates 24x7x365 and is designed for continuous processing, while The Clearing House's RTP network is similarly available around the clock (Federal Reserve; The Clearing House).

These are not simply performance targets. Continuous availability is part of what these payment infrastructures are designed to deliver.

Take the “always” out of “always-on,” even for an hour, and the failure isn't cosmetic. It's unpaid wages sitting in a queue, a merchant who can't accept a card, or a nostro position that can’t be reconciled at close of business.

That is why reliability in payments is not just an infrastructure metric. It has a direct operational consequence for every institution and customer depending on the transaction reaching its destination when expected.

Regulators have put a number on it

The Bank for International Settlements’ Principles for Financial Market Infrastructures state that a financial market infrastructure's business continuity plan should be designed so that critical information technology systems can resume operations within two hours following a disruptive event, while enabling settlement to be completed by the end of the day even in extreme circumstances (CPSS-IOSCO, 2012). That gives systemically important financial infrastructures a clear benchmark rather than a vague expectation to recover “promptly.” The point is not that every failure must be prevented. It is that the infrastructure needs to be designed and tested around what happens when prevention fails.

A 2022 Committee on Payments and Market Infrastructures (CPMI) and International Organization of Securities Commissions (IOSCO) review of 37 financial market infrastructures across 29 jurisdictions found that a small number had not fully developed cyber response and recovery plans capable of meeting that target, while others showed shortcomings when tested against extreme cyberattack scenarios (CPMI-IOSCO, 2022).

That finding matters beyond cybersecurity. It highlights the gap that can exist between having recovery plans on paper and knowing that those plans will work under the conditions in which they are needed.

In the European Union, the Digital Operational Resilience Act (DORA) has reinforced this direction with binding operational resilience requirements for financial entities within its scope. Since January 2025, banks and payment institutions have been subject to requirements covering ICT risk management, incident reporting, resilience testing and third-party risk management (European Union, Regulation (EU) 2022/2554).

For regulated financial institutions, resilience can no longer remain simply an infrastructure claim. It has to be governed, tested and evidenced.

The cost is not abstract

The financial impact of downtime varies widely by institution, transaction volume and the nature of the disruption, but the scale can be significant. Splunk's 2024 research with Oxford Economics on the financial services sector estimates that organizations lose an average of $309 million a year to unplanned outages, including revenue losses as well as contractual and legal costs (Splunk & Oxford Economics, 2024).

The figure is an industry average rather than a prediction of what any single payment outage will cost. But it puts a useful scale around a problem that is often discussed only in terms of uptime percentages and recovery targets.

None of that counts the quieter cost: a corporate treasurer who starts routing volume through a second provider after one missed cut-off, or a retail customer who never fully trusts a mobile transfer again after watching one disappear into a spinner.

Those consequences are harder to put into a downtime calculation, but they are also harder to reverse.

What this means for how payment providers compete

The processors and vendors that win the next decade of instant payments won't be the ones with the busiest feature roadmap. They'll be the ones a bank's operations team never has to think about, because the architecture is active-active by design rather than relying on failover after the fact. Critical dependencies are spread across providers and geographies so that one bad update can't take the whole book down at once. Recovery time objectives are tested against real scenarios often enough that the two-hour clock actually means something. The system degrades gracefully under strain instead of falling over.

The distinction is important. Active-active architecture, geographic diversity and documented recovery objectives are useful design choices, but none of them proves reliability on its own. Their value comes from whether they remove single points of failure, contain disruption and continue working when the assumptions behind the architecture are put under pressure.

None of that shows up as a headline feature in a sales deck. It shows up as the thing that never happened: the outage that didn't make the news, the cut-off that got hit anyway, the customer who never had a reason to wonder if their money was safe.

Increasingly, that is the entire pitch.

References

  1. Bank of England. (2015). Independent Review of RTGS Outage on 20 October 2014.

  2. Committee on Payment and Settlement Systems (CPSS) & International Organization of Securities Commissions (IOSCO). (2012). Principles for Financial Market Infrastructures.

  3. Committee on Payments and Market Infrastructures (CPMI) & IOSCO. (2022). Implementation monitoring of the Principles for Financial Market Infrastructures.

  4. European Payments Council. (2025). SEPA Instant Credit Transfer (SCT Inst) Rulebook.

  5. European Union. (2022). Regulation (EU) 2022/2554 on Digital Operational Resilience for the Financial Sector (DORA).

  6. Federal Reserve. FedNow Service documentation.

  7. Information Technology Intelligence Consulting (ITIC). (2024). 2024 Global Server Hardware, Server OS Reliability Survey.

  8. Splunk & Oxford Economics. (2024). The Hidden Costs of Downtime.

  9. The Clearing House. RTP Network overview and operating model.

  10. Visa. (2018). European service disruption incident statement.