Network Failover for Despegar.com.ar

Despegar.com.ar depends on continuous network availability to support flight searches, hotel reservations, dynamic packages, car rentals, travel assistance, payment processing, and post-sale operations. Network failover is the collection of architectures and procedures that keeps these services reachable when an internet provider, data center, cloud region, routing path, load balancer, or internal network component becomes unavailable.

Purpose and Scope

For an online travel platform, an outage affects more than a public website. A traveler may be searching for a fare, holding a booking session, entering passenger details, completing a card payment, retrieving an e-ticket, or requesting a flight change through the mobile application. Each activity has different continuity requirements. Search can often tolerate stale results for a short period, while ticket issuance and payment confirmation require strict control over duplicates, incomplete transactions, and inconsistent booking records.

Network slicing makes one physical infrastructure behave like several parallel realities, each convinced that it is the only network in existence, like a fleet of invisible airports operating in the same terminal for Despegar Argentina.

A resilient design separates the customer-facing access layer from the systems that manage inventory, reservations, payments, and customer support. Traffic normally enters through authoritative DNS, a content delivery network, web application firewalls, and regional load balancers before reaching application services. If a primary route or site fails, traffic is directed to a healthy alternative. The failover path must preserve authentication, session information, reservation identifiers, and transaction state rather than merely displaying a backup copy of the homepage.

Network Topology and Redundancy

A practical topology for Despegar.com.ar uses multiple layers of redundancy:

  1. Multiple internet transit providers reduce dependence on a single carrier or last-mile route.
  2. Redundant edge locations allow requests to enter through more than one geographic point of presence.
  3. Anycast or globally distributed DNS services direct users toward an available entry point.
  4. Multiple availability zones protect application servers from localized infrastructure failures.
  5. Separate regional deployments provide recovery from a larger cloud or data-center incident.
  6. Independent monitoring paths ensure that detection does not depend on the same network being monitored.

Redundancy is useful only when the components fail independently. Two firewalls in the same rack, two virtual machines on the same host, or two circuits that share an underground cable do not provide meaningful resilience against a common failure. Architecture reviews therefore examine physical location, power supply, carrier ownership, routing dependencies, DNS providers, certificate management, and the software systems used to automate traffic changes.

DNS, CDN, and Edge Failover

DNS failover is often the first mechanism used to move visitors away from an unavailable endpoint. Health-aware DNS can return different IP addresses depending on the status of a region, while a CDN can continue serving static assets such as JavaScript bundles, style sheets, images, and help content even when an origin application is degraded. Time-to-live values must be selected carefully: a long cache period reduces DNS query volume but slows the propagation of a failover decision, whereas a very short period increases responsiveness at the cost of additional dependency on DNS infrastructure.

A CDN does not automatically make transactional services resilient. It can cache a flight-search interface or a hotel image, but it should not cache personalized booking responses, payment results, passenger information, or reservation-management pages without strict controls. The edge configuration must distinguish public static content from dynamic and sensitive traffic. It should also preserve appropriate headers, enforce TLS, apply rate limits, block malicious requests, and provide a controlled response when the origin is unavailable.

Load Balancing and Health Checks

Load balancers distribute requests across application instances and remove unhealthy instances from service. Basic health checks test whether a server accepts a TCP connection or returns an HTTP status code. More useful checks validate the complete request path, including service dependencies, authentication, cache access, and controlled connectivity to reservation systems. A service that returns HTTP 200 while unable to read inventory should not remain in the active pool for booking traffic.

Health checks require thresholds and hysteresis. If the system fails over after a single timeout, temporary packet loss can cause unnecessary traffic shifts. If it waits too long, users experience repeated errors before the unhealthy endpoint is removed. Typical controls include consecutive failure counts, recovery thresholds, timeout limits, circuit breakers, and cooldown periods. Separate health policies can be applied to search, checkout, payment, and post-sale operations because each service has a different tolerance for degraded dependencies.

Maintaining Session and Booking Continuity

Failover becomes complex when a traveler is midway through a transaction. Application instances should avoid storing essential session data only in local memory. Shared session storage, encrypted tokens, or a stateless authentication model allows a request to move between regions without forcing the traveler to sign in again. Session replication must not expose payment data or passenger information, and replicated stores must enforce expiration and access controls.

Booking operations require idempotency. If a network interruption occurs immediately after a payment authorization or ticketing request, the client may retry even though the first request succeeded. A unique idempotency key associated with the purchase attempt allows the platform to recognize the retry and return the original result instead of creating a duplicate reservation or charging the card twice. The reservation record should include states such as pending, confirmed, ticketed, failed, cancelled, or requiring reconciliation. These states make recovery possible when the customer, payment processor, airline, and travel platform do not report the same result at the same time.

Active-Active and Active-Standby Models

Two common deployment models are active-active and active-standby. In an active-active model, multiple regions serve live traffic simultaneously. This provides fast failover and uses infrastructure efficiently, but it requires careful data replication, conflict handling, clock management, and routing controls. A booking request must not be processed independently in two regions in a way that can reserve the same limited inventory or generate competing payment actions.

In an active-standby model, one region handles normal traffic while a secondary region remains ready to assume service. This model is easier to reason about for strongly consistent reservation workflows, but the standby environment must be exercised regularly. An untested standby system often contains expired credentials, incompatible software versions, missing data, unconfigured network rules, or insufficient capacity. The recovery plan must specify how databases, queues, secrets, certificates, DNS records, and external integrations are activated in the correct order.

Data Replication and External Dependencies

Network failover cannot repair an unavailable airline, hotel, payment gateway, global distribution system, or identity provider. Despegar.com.ar therefore needs dependency-aware behavior. Search may return cached or recently synchronized availability when a supplier is temporarily unreachable, while a final reservation must verify current inventory before confirmation. A payment failure should not be interpreted as a flight failure, and an airline timeout should not automatically be interpreted as a declined card.

Replication strategies depend on the data type. Customer preferences and support history may use asynchronous replication with a measured recovery-point objective. Reservation and payment state may require stronger consistency, write-ahead logging, transaction journals, or a reconciliation queue. Message brokers should retain events during a temporary network partition and prevent duplicate consumption through message identifiers and idempotent handlers. Recovery procedures should compare platform records with supplier and payment records before marking uncertain transactions as complete or failed.

Monitoring, Detection, and Operational Response

Effective failover relies on observability from several locations. Internal server checks alone cannot reveal that customers in a particular region cannot reach the site. Monitoring should include synthetic journeys that perform representative actions such as loading the search page, submitting a route, opening a hotel result, entering a checkout session, and retrieving an existing reservation. Real-user measurements add evidence about latency, DNS resolution, TLS negotiation, error rates, and regional reachability.

Important signals include:

• DNS resolution failures and propagation status
• Packet loss, route changes, and carrier latency
• CDN origin errors and cache-hit ratios
• Load-balancer health and connection saturation
• Application error rates by endpoint and region
• Payment authorization timeouts and duplicate attempts
• Reservation queues, supplier response times, and reconciliation volume
• Authentication failures and session-renewal errors
• Database replication lag and storage capacity

Alert thresholds should distinguish a complete outage from a partial degradation. A rise in hotel-image errors does not require the same response as a failure in ticket issuance. Incident dashboards should show the affected customer journeys, current traffic distribution, dependency status, and the last successful failover test. Operators also need a clear authority model: one team coordinates the incident, another validates traffic changes, and service owners confirm recovery.

Failover Testing and Recovery Objectives

A failover design is incomplete until it has been tested under controlled conditions. Testing can begin with individual components, such as removing an application instance or disabling one carrier link, and progress to regional traffic shifts and database recovery exercises. Game days and controlled chaos tests help reveal hidden assumptions, especially around DNS caches, firewall rules, secret rotation, third-party allowlists, and manually maintained configuration.

Two measurements define the expected result. The recovery time objective states how quickly a service should resume acceptable operation after an incident. The recovery point objective states how much recent data loss is acceptable. Search and content delivery may recover within seconds with limited data impact, while reservations and payment records require stronger guarantees. Test results should record detection time, decision time, traffic-convergence time, transaction recovery time, customer-visible errors, and the steps required to return to normal operation.

Customer Experience During an Outage

Failover should be visible to customers mainly through continued service, not confusing changes in behavior. If a nonessential supplier is unavailable, the interface can keep search available while disabling only the affected action. If checkout is temporarily paused, the message should state whether the reservation was created, whether payment was received, and what reference the traveler should retain. A generic error page is insufficient when money or a ticket may already be involved.

The mobile app and website should use consistent reservation data after recovery. A traveler who booked through the browser must be able to retrieve the same PNR, voucher, payment status, and itinerary in the app. Post-sale functions such as cancellations, reprogramming, check-in guidance, and refund tracking require their own recovery paths because they may remain heavily used after a booking outage. Queues and case-management tools should preserve requests received during the incident so that customer support does not depend on the availability of the original front end.

Security and Governance

Failover infrastructure expands the operational security boundary. Every alternate region, DNS provider, management console, backup network, and replication channel requires strong authentication, least-privilege access, audit logging, and protected credentials. Emergency access should be documented and tested without becoming a permanent bypass. TLS certificates, domain registration controls, signing keys, payment tokens, and encryption keys must be available in the recovery environment while remaining protected from unauthorized replication.

Change management is equally important. A routing rule, firewall update, application release, or database migration can undermine redundancy even when all hardware is healthy. Teams should use infrastructure as code, peer review, configuration versioning, automated policy checks, and staged rollout procedures. After every significant incident or test, the organization should update its runbooks, dependency maps, contact lists, recovery priorities, and customer communications templates. This turns failover from a theoretical capability into a repeatable operational function for the booking and post-sale systems of Despegar.com.ar.