Built for Reliability: How American Express Processes Payments at Scale
American Express operates as the core payment network connecting acquiring banks to card issuers during transactions. In 2018, the company began modernizing its platform onto cloud-native infrastructure, shifting away from hardware engineered for continuous uptime to an environment where unexpected server failures are common. Traditional patterns like event-driven processing and monolithic designs were rejected due to strict low-latency requirements and scaling constraints. Instead, the engineering team adopted a design rooted in cell-based architecture to isolate failures and maintain real-time performance. In a cell-based architecture, a cell is a complete and self-contained copy of an application stack designed to function as an independent failure domain. Each cell owns its required components—such as microservices, databases, and infrastructure—and avoids synchronous cross-cell dependencies along the critical execution path. Unlike microservices, which partition systems by functional domain, cells partition architectures by failure boundaries. Engineers at American Express structure these boundaries around specific real-time transaction journeys to isolate failures and maintain platform availability. American Express manages data locality within its cellular architecture by categorizing data as immutable, semi-static, or dynamic based on change frequency. Immutable and semi-static data are proactively pushed to all cells prior to transactions to prevent cold cache penalties and eliminate external synchronous calls. For dynamic data that updates per transaction, the architecture inverts the paradigm by moving transactions to the data using deterministic routing. Handled by the Global Transaction Router, routing decisions are derived from transaction attributes for consistency, while asynchronous replication runs in the background for failover.
閱讀原文 ↗目錄
Core Payments Ecosystem
American Express operates as the core payment network connecting acquiring banks to card issuers during transactions. In 2018, the company began modernizing its platform onto cloud-native infrastructure, shifting away from hardware engineered for continuous uptime to an environment where unexpected server failures are common. Traditional patterns like event-driven processing and monolithic designs were rejected due to strict low-latency requirements and scaling constraints. Instead, the engineering team adopted a design rooted in cell-based architecture to isolate failures and maintain real-time performance.
- American Express functions as the core payment network between an acquiring bank and a card issuer.
- American Express began migrating its core payment platform to cloud-native infrastructure in 2018.
- Cloud infrastructure required designing for frequent, uncontrollable server failures rather than relying on high-reliability hardware.
- Event-driven architecture was rejected because core payment processing demands real-time responses with low latency.
- Monolithic architecture was deemed a poor fit for the platform's scaling requirements.
- The engineering team chose an approach based on cell-based architecture, building on patterns developed internally during the Service-Oriented Architecture (SOA) era.
Cell Boundaries
In a cell-based architecture, a cell is a complete and self-contained copy of an application stack designed to function as an independent failure domain. Each cell owns its required components—such as microservices, databases, and infrastructure—and avoids synchronous cross-cell dependencies along the critical execution path. Unlike microservices, which partition systems by functional domain, cells partition architectures by failure boundaries. Engineers at American Express structure these boundaries around specific real-time transaction journeys to isolate failures and maintain platform availability.
- A cell is a self-sufficient copy of a processing stack containing all microservices, databases, and infrastructure required to process a transaction locally within a region.
- Cells form isolated failure domains that can be taken out of rotation for maintenance or incidents without disrupting the broader platform.
- Synchronous cross-cell dependencies are strictly excluded from the critical transaction path, although reference data replication and observability aggregation can span cells.
- Microservices divide systems functionally, whereas cells divide systems by failure boundaries; a single cell typically encapsulates multiple microservices.
- American Express engineers recommend sizing cells around user journeys, isolating real-time processing requirements from after-the-fact processing workflows.
Data Locality
American Express manages data locality within its cellular architecture by categorizing data as immutable, semi-static, or dynamic based on change frequency. Immutable and semi-static data are proactively pushed to all cells prior to transactions to prevent cold cache penalties and eliminate external synchronous calls. For dynamic data that updates per transaction, the architecture inverts the paradigm by moving transactions to the data using deterministic routing. Handled by the Global Transaction Router, routing decisions are derived from transaction attributes for consistency, while asynchronous replication runs in the background for failover.
- American Express divides data into three categories: immutable, semi-static (changing hours to yearly), and dynamic (changing per transaction).
- For immutable and semi-static data, reference data is pushed and distributed to every cell beforehand rather than pulled via on-demand caching.
- Pushing data ahead avoids cold-cache fall-through reads and keeps replication entirely outside the critical transaction path.
- For dynamic data, the architecture avoids stale-state issues by routing the transaction to the data rather than replicating data to the transaction.
- The Global Transaction Router handles deterministic routing based on payload attributes like partner, market, and payment type, alongside priority-based routing for other use cases.
- Message-based replication between cells runs asynchronously in the background to ensure failover data availability without delaying in-flight transactions.
Global Transaction Router
The Global Transaction Router routes traffic and enforces boundaries across cells in a distributed payment system, acting as the sole communication path between cells and external institutions. To prevent this critical component from becoming a failure point, American Express engineers excluded payment business logic, limiting the router to minimal message parsing for routing decisions. Resilience is further maintained by running instances nearly stateless with non-persistent storage, using asynchronous logging and configuration updates, and deploying parallel active instances across regions. This design demonstrates how essential structural chokepoints remain reliable through deliberate simplicity and minimal dependencies.
- Cells in the payments mesh lack direct communication capabilities and rely entirely on the Global Transaction Router for both inter-cell and external routing.
- American Express engineers kept the router simple by excluding payment business logic, avoiding centralized data lookups.
- Router instances maintain high availability by staying nearly stateless with state kept in non-persistent storage.
- Asynchronous mechanisms for logging with buffer truncation and memory-based configuration loading prevent transactional bottlenecks.
- An active-active deployment across multiple regions is favored over active-standby to enhance availability for stateless components.
Credit Card Authorization Flow
The credit card authorization flow at American Express demonstrates a decoupled architecture managed by the Global Transaction Router. Upon receiving a transaction initiated from a merchant point-of-sale terminal, the router forwards the request to an internal processing cell using deterministic or priority-based routing. Within the cell, microservices handle data enrichment, transformation, validation, and issuer identification, but they never communicate directly outside their boundary. Instead, the transaction is returned to the router, which interfaces with card issuers and routes responses back to the context-holding cell before returning confirmation to the acquiring bank.
- A point-of-sale transaction first enters American Express systems through the Global Transaction Router.
- The router directs transactions to internal processing cells using either deterministic or priority-based routing depending on the use case.
- Microservices inside a cell handle validation, enrichment, transformation, and issuer determination.
- Cells are strictly isolated and never communicate directly with external card issuers.
- The Global Transaction Router contacts card issuers and uses deterministic routing to route the response back to the specific cell holding the transaction context.
Mid-Transaction Failure
Payments processing at American Express utilizes an orchestrated microservices architecture to manage workflows and monitor service health. When a mid-transaction failure occurs, the orchestrator halts execution and returns the request to the Global Transaction Router to find a healthy cell. Crucially, the system discards any completed partial work and restarts the transaction from scratch in the new cell. This design choice prevents cross-cell shared state and synchronization issues, accepting the minor cost of redoing a sub-second transaction in exchange for loose coupling.
- American Express uses an orchestrator microservice to supervise payment workflows and detect failures in individual microservices.
- When a failure occurs, the transaction is sent back to the Global Transaction Router to select a healthy cell.
- Instead of resuming from partial progress, American Express discards incomplete work and restarts the transaction from the beginning.
- Discarding partial progress avoids shared state between cells, eliminating cross-cell synchronization problems and consistency risks during failover.
- Each cell maintains its own database clusters and runs independently without reliance on state from other cells.
- Redoing work is viable because payment transactions complete in hundreds of milliseconds, making the structural isolation worth the minimal latency cost.
Recovery Semantics
Transaction recovery semantics define a point of no return up to which transactions can be safely restarted or rerouted before reaching external systems like card issuers. American Express optimizes this by placing irreversible steps late in the payment flow to maximize the recoverable window, while relying on idempotency keys to suppress duplicate processing when late rerouting is impossible. When failed cells recover, platforms mitigate stale state writes through fast transaction completion, downstream idempotency duplicate suppression, and paced, percentage-based traffic shifting.
- Rerouting transactions is safe within the core ecosystem but unsafe once sent to external systems like card issuers.
- American Express maximizes the recoverable window by placing the point of no return as late as possible in payment processing.
- Transactions carry consistent unique identifiers that downstream systems use for idempotency and duplicate suppression.
- Percentage-based canary shifting enables gradual cell draining, partial-load validation, and controlled failback during recovery.
- Preventing recovered cells from writing stale state relies on transaction speed, idempotency identifiers, and paced recovery.
Design Tradeoffs
American Express accepts key trade-offs in its cell-based architecture to preserve cell independence and protect data integrity. These trade-offs include duplicating services across cells to eliminate cross-cell network hops, along with dropping non-critical application logs via buffer truncation during high platform pressure. Observability experiences latency because telemetry is written locally first and aggregated asynchronously, and transactions with strong consistency requirements may be rejected if cross-cell data is not validated. Ultimately, partitioning the platform into independent units trades higher management complexity for a significantly reduced blast radius during failures.
- American Express duplicates service implementations across cells to eliminate cross-cell network hops and preserve cell independence.
- A buffer truncation policy drops application observability logs during sustained pressure to ensure transaction processing continues.
- Global dashboards experience visibility lag because metrics, logs, and traces are written locally within each cell before asynchronous aggregation.
- Transactions requiring strong consistency may be rejected to protect data integrity if asynchronous cross-cell data synchronization cannot be validated.
- Cell-based systems reduce the blast radius of failures rather than reducing overall failure frequency, trading complexity for limited failure impact.
Conclusion
The American Express engineering team designed a resilient payment platform capable of surviving cell failures by ensuring transactions never depend on multiple cells simultaneously. This architecture separates data strategies by pushing static reference data to all cells while routing transactions directly to rapidly changing data. System dependability is anchored by a thin Global Transaction Router that avoids complex logic at critical chokepoints. Instead of resuming partial work across shared state, the platform discards partial transactions and restarts them, pushing the point of no return as late as possible to maximize failure survivability.
- Transactions are isolated to prevent a single transaction from depending on two cells simultaneously.
- Infrequently changing reference data is distributed to all cells ahead of time, whereas volatile data remains local and transactions are routed to it.
- The Global Transaction Router maintains dependability at the system chokepoint by minimizing its internal logic.
- The platform chooses restarting over resuming failed transactions, eliminating shared state across cells at a latency cost of a few hundred milliseconds.
- Positioning the point of no return late in the execution flow widens the window for survivable failures.