指南

Multi-region failover with a WireGuard mesh

高级阅读时间 40 分钟更新时间 2026年7月8日
简短回答

三地三实例,通过WireGuard组网,应用仅绑定到网格,并在前面放置健康检查的DNS记录,就能以每月不到110美元实现真正的容错。设计约束在于数据库复制:区域内同步,区域间异步。

01 Choose three independent failure domains

Different countries, ideally different transit mixes. Amsterdam, Ashburn and Singapore is the classic triangle. Three is the minimum for quorum — two gives you a split-brain problem rather than redundancy.

02 Build the mesh

Each node gets a stable private address. Every node peers with every other node; with three nodes that is three tunnels.

# node A — /etc/wireguard/mesh.conf
[Interface]
Address = 10.10.0.1/24
ListenPort = 51821
PrivateKey = <A private>

[Peer]                       # node B
PublicKey = <B public>
Endpoint = b.example.net:51821
AllowedIPs = 10.10.0.2/32
PersistentKeepalive = 25

[Peer]                       # node C
PublicKey = <C public>
Endpoint = c.example.net:51821
AllowedIPs = 10.10.0.3/32
PersistentKeepalive = 25

03 Bind services to the mesh only

The database, the cache and the internal API should listen on 10.10.0.x, never on the public address. This removes an entire class of exposure without a single firewall rule.

04 适当复制数据库

Synchronous replication across regions is impractical — a 160 ms round trip becomes the floor for every write. Use asynchronous replication across regions and know your recovery point objective.

05 在DNS层进行健康检查

Short TTLs plus health-checked failover records, or anycast if your regions support it. Then run a failure drill: kill a region deliberately, on a weekday, while you are watching.

常见问题

Why not just buy a bigger server?

较大的服务器与小型服务器具有相同的故障域数量:一个。可用性来自独立性,而非容量。

How far apart can etcd or Postgres synchronous replicas be?

保持同步副本之间的距离在约100毫秒以内。超过这一距离,写入延迟将成为每笔交易的主要成本。