Руководства

Multi-region failover with a WireGuard mesh

Продвинутый40 мин чтенияОбновлено 8 июля 2026 г.
Краткий ответ

Три инстанса в трёх странах, объединённые через WireGuard, с приложением, привязанным только к mesh-сети, и DNS-записью с проверкой работоспособности перед ней, дают настоящую отказоустойчивость менее чем за $110 в месяц. Ограничение дизайна — репликация базы данных: синхронная в пределах региона, асинхронная между ними.

01 Choose three independent failure domains

Different countries, ideally different transit mixes. Amsterdam, Ashburn and Singapore is the classic triangle. Three is the minimum for quorum — two gives you a split-brain problem rather than redundancy.

02 Build the mesh

Each node gets a stable private address. Every node peers with every other node; with three nodes that is three tunnels.

# node A — /etc/wireguard/mesh.conf
[Interface]
Address = 10.10.0.1/24
ListenPort = 51821
PrivateKey = <A private>

[Peer]                       # node B
PublicKey = <B public>
Endpoint = b.example.net:51821
AllowedIPs = 10.10.0.2/32
PersistentKeepalive = 25

[Peer]                       # node C
PublicKey = <C public>
Endpoint = c.example.net:51821
AllowedIPs = 10.10.0.3/32
PersistentKeepalive = 25

03 Bind services to the mesh only

The database, the cache and the internal API should listen on 10.10.0.x, never on the public address. This removes an entire class of exposure without a single firewall rule.

04 Реплицируйте базу данных правильно

Synchronous replication across regions is impractical — a 160 ms round trip becomes the floor for every write. Use asynchronous replication across regions and know your recovery point objective.

05 Проверка здоровья на уровне DNS

Short TTLs plus health-checked failover records, or anycast if your regions support it. Then run a failure drill: kill a region deliberately, on a weekday, while you are watching.

Часто задаваемые вопросы

Why not just buy a bigger server?

У более крупного сервера количество точек отказа такое же, как у маленького: одна. Доступность обеспечивается независимостью, а не мощностью.

How far apart can etcd or Postgres synchronous replicas be?

Держите синхронные реплики в пределах примерно 100 мс друг от друга. За этим порогом задержка записи становится доминирующей стоимостью каждой транзакции.