100.64.0.0/10 was a bad default: it is exactly Tailscale's range, and carrier-grade NAT's. There is no IPv4 range that is free on every host, so there is now no default at all — IPv4 is off until --ipv4-range names one. IPv6 is unaffected and still works out of the box, because a ULA derived from the network id collides with essentially nothing. The more serious problem this exposed: the range is an input to the address derivation, and each agent derives every peer's address itself. Two members configured with different ranges would therefore derive different addresses for each other and IPv4 would silently misroute. So the range now travels in the announcement — not as a request and never trusted, only so the mismatch is seen. A peer whose range disagrees gets no IPv4 address here, keeps working over IPv6, and the reason is reported with both ranges named. The announcement format goes to version 2. postcard is not self-describing, so an older peer cannot read it; the version check already catches that and now says which side needs updating. The (Ipv4Addr, u8) tuple that had spread across six modules is now an Ipv4Range with validation, Display and FromStr, so a bad --ipv4-range is refused with a reason instead of being accepted and misbehaving later. It is also rejected when passed without --wireguard rather than ignored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
271 lines
12 KiB
Markdown
271 lines
12 KiB
Markdown
# The WireGuard data plane
|
|
|
|
WireGuard is the first IP plugin. It carries user traffic between
|
|
participants while the control plane keeps doing its own job: deciding who is
|
|
in the network and carrying each participant's opaque announcement.
|
|
|
|
Module boundaries are in [architecture.md](architecture.md), the control
|
|
protocol in [protocol.md](protocol.md), the security consequences in
|
|
[threat-model.md](threat-model.md).
|
|
|
|
## Userspace, not the kernel
|
|
|
|
WireGuard here is [boringtun]'s protocol state machine running in this
|
|
process. There is **no kernel WireGuard module** and **no `wg` tool**: the same
|
|
code runs everywhere, and the protocol can be exercised in tests without any
|
|
privileges at all.
|
|
|
|
The only privileged step left is creating a packet interface so the operating
|
|
system can hand us IP packets, and even that is behind a trait
|
|
([`TunFactory`]) with an in-memory implementation.
|
|
|
|
| | needs privileges | what it proves |
|
|
|---|---|---|
|
|
| `MemoryTunFactory` | no | handshake, encryption, routing, address ownership |
|
|
| `SystemTunFactory`, attaching | none, if the interface was prepared | traffic actually reaches the OS |
|
|
| `SystemTunFactory`, creating | `CAP_NET_ADMIN` | the same, at the cost of a capability |
|
|
|
|
`SystemTunFactory` attaches to an interface that already exists and only
|
|
creates one when it does not. A persistent interface created by root and owned
|
|
by the user lets the agent run with no privileges at all; see *Running
|
|
unprivileged* in [../README.md](../README.md#running-unprivileged).
|
|
|
|
[boringtun]: https://docs.rs/boringtun
|
|
[`TunFactory`]: https://docs.rs/tsunagi
|
|
|
|
## Where the packets go
|
|
|
|
The plugin does not know and does not care. It is handed a `PacketLink` per
|
|
peer by the agent and runs a WireGuard tunnel over it:
|
|
|
|
```text
|
|
TUN device (IP packets) PacketLink per peer
|
|
| |
|
|
v v
|
|
destination address -> peer --Tunn.encapsulate--> ciphertext -> transport
|
|
source address checked <--Tunn.decapsulate-- ciphertext <- transport
|
|
```
|
|
|
|
Reachability — hole punching, relay fallback — belongs to the transport, which
|
|
today is iroh. That is the whole reason the plugin's announcement says *who* it
|
|
is and never *where* it is: there is no address for a peer to advertise, get
|
|
wrong, or lie about.
|
|
|
|
**Two peers behind NAT work exactly as well as iroh does.** iroh hole punches a
|
|
direct path when it can and falls back to a relay when it cannot; the tunnel
|
|
rides on whichever it got. There is no separate STUN, no separate hole punching
|
|
and no second set of NAT problems to solve for WireGuard.
|
|
|
|
## Checking it from outside
|
|
|
|
`tsunagi status` asks a running agent over its local control socket and prints
|
|
what it sees, including whether each tunnel has actually handshaken. See
|
|
[../README.md](../README.md#checking-that-it-works).
|
|
|
|
## Deterministic overlay addressing
|
|
|
|
A mesh with no coordinator cannot hand out addresses, so everyone derives their
|
|
own. The result is an IPv6 unique local address (RFC 4193):
|
|
|
|
```text
|
|
prefix (/64) = 0xfd || SHA-256( LP(domain) || LP("prefix") || LP(network_id) )[0..7]
|
|
iid (64b) = SHA-256( LP(domain) || LP("interface") || LP(network_id) || LP(wg_public_key) )[0..8]
|
|
address = prefix || iid
|
|
```
|
|
|
|
with `domain = "tsunagi-wireguard-overlay-v1"` and `LP(x) = u32_be(len(x)) || x`,
|
|
the same unambiguous encoding the rest of the project uses.
|
|
|
|
Two consequences matter:
|
|
|
|
* every member of a network derives the **same `/64`**, so the overlay is one
|
|
subnet that nobody had to allocate;
|
|
* a member's address is bound to its WireGuard public key, so address
|
|
ownership can be checked locally rather than believed.
|
|
|
|
## IPv4 alongside IPv6
|
|
|
|
The overlay carries IPv4 as well, but **it is off unless you name a range**,
|
|
and every member must name the same one:
|
|
|
|
```bash
|
|
tsunagi up --network lab --secret "$SECRET" --wireguard --ipv4-range 10.77.0.0/16
|
|
```
|
|
|
|
Two reasons it has no default.
|
|
|
|
**There is no IPv4 range that is free everywhere.** `100.64.0.0/10` is
|
|
Tailscale's and carrier-grade NAT's, `10.0.0.0/8` and `192.168.0.0/16` are on
|
|
half the networks in the world, `172.17.0.0/16` is Docker. Picking one
|
|
requires knowing what is already in use on every machine that will join, which
|
|
is the operator's knowledge, not ours. IPv6 needs none of this: a ULA derived
|
|
from the network id collides with essentially nothing.
|
|
|
|
**The range is an input to the derivation.** Each agent computes every peer's
|
|
address itself, so two members configured with different ranges would derive
|
|
different addresses for each other and IPv4 would silently misroute. The range
|
|
therefore travels in the announcement — not as a request, and never trusted,
|
|
but so that a mismatch is *detected*. When it happens, the offending peer gets
|
|
no IPv4 address here, keeps working over IPv6, and the reason is reported:
|
|
|
|
```text
|
|
! wireguard: peer SDsEb/WF is configured with the IPv4 overlay range
|
|
10.81.0.0/16 but this agent uses 10.80.0.0/16; every member must use the
|
|
same one. That peer has no IPv4 address here and is reachable over IPv6 only.
|
|
```
|
|
|
|
**IPv4 addresses can also collide with each other.** A 64 bit interface
|
|
identifier makes an IPv6 collision impossible in practice; IPv4 has nothing
|
|
like that room. In a `/16` with 50 members the chance that two members derive
|
|
the same address is roughly 2%. A mesh with no coordinator cannot allocate
|
|
around it, so the collision is resolved instead: the member whose WireGuard
|
|
public key sorts lower keeps the address, a rule every member computes
|
|
identically and therefore agrees on without exchanging anything. The other
|
|
member has no IPv4 address and remains reachable over IPv6. Pick a roomy
|
|
range — a `/16` for a handful of machines, larger for more — and the odds stay
|
|
small.
|
|
|
|
The honest summary: **IPv6 always works. IPv4 is opt-in, needs agreement, and
|
|
degrades predictably when it does not get it.** Allocating IPv4 properly needs
|
|
the agreed state described in [sync-model.md](sync-model.md).
|
|
|
|
## Address ownership is enforced, not announced
|
|
|
|
Kernel WireGuard enforces `AllowedIPs`. In userspace that is our job, and
|
|
[`device`] does it on both sides:
|
|
|
|
* **outbound**, a packet is routed to the peer that *owns* its destination
|
|
address; a destination nobody owns is counted as unroutable and dropped;
|
|
* **inbound**, a decrypted packet is dropped unless its *source* is exactly the
|
|
address derived for the peer whose tunnel decrypted it.
|
|
|
|
Both apply to IPv4 and IPv6 alike.
|
|
|
|
So a participant cannot receive traffic addressed to somebody else and cannot
|
|
forge traffic that appears to come from somebody else. A participant who knows
|
|
the network secret can mint many keys and therefore occupy many addresses, but
|
|
it cannot choose to collide with an existing member without finding a hash
|
|
preimage.
|
|
|
|
The announcement also carries the address the peer believes it has. It is never
|
|
used — only cross-checked — so a version skew produces a clear rejection rather
|
|
than silent non-connectivity.
|
|
|
|
[`device`]: https://docs.rs/tsunagi
|
|
|
|
## MTU
|
|
|
|
Two constraints pull against each other.
|
|
|
|
**IPv6 sets a floor of 1280 bytes** (RFC 8200), and Linux enforces it
|
|
brutally: an interface whose MTU drops below 1280 loses IPv6 entirely — its
|
|
`/proc/sys/net/ipv6/conf/<dev>` directory disappears and `ip -6 address add`
|
|
answers `Invalid argument`. So the overlay MTU cannot go below 1280, and the
|
|
plugin refuses a smaller one at startup instead of letting it fail obscurely.
|
|
|
|
**The transport sets a ceiling.** Every packet rides in one datagram and
|
|
WireGuard adds 32 bytes, so a link must carry `mtu + 32` = 1312 bytes. A direct
|
|
QUIC path typically offers around 1380, which fits. A relayed path can offer
|
|
less, and then full-size packets do not fit: they are dropped and counted as
|
|
`dropped_oversize`, never truncated, and the plugin reports the exact numbers
|
|
when the tunnel is set up.
|
|
|
|
There is no room left to trade, so the default MTU is exactly 1280.
|
|
Fragmenting a packet across several datagrams would lift the ceiling and is
|
|
not implemented.
|
|
|
|
## Lifecycle
|
|
|
|
* A network is activated → the plugin loads or creates its key for that
|
|
network, derives the interface name, and creates the packet interface. If
|
|
that fails — no privileges, for instance — the key and the announcement still
|
|
work and the interface is retried on the next reconcile.
|
|
* A peer announces its key → recorded.
|
|
* A data link to that peer arrives → recorded.
|
|
* Reconciliation starts a tunnel for every peer that has **both**, and removes
|
|
tunnels for peers that lost either.
|
|
* A network is deactivated, or the agent shuts down → the interface and every
|
|
tunnel go away. The key stays, so coming back keeps the same overlay address.
|
|
|
|
There is no external configuration file and no command line tool, so unlike a
|
|
kernel-WireGuard setup there is nothing outside this process for anybody to
|
|
edit. Reconciliation is purely "do the running tunnels match what is known".
|
|
|
|
## Using it
|
|
|
|
```bash
|
|
# On both machines
|
|
tsunagi up --network lab --secret "$SECRET" --wireguard
|
|
```
|
|
|
|
See the two-machine walkthrough in [../README.md](../README.md#trying-it-on-two-machines).
|
|
|
|
From the library:
|
|
|
|
```rust,no_run
|
|
use std::sync::Arc;
|
|
use tsunagi::config::{AgentConfig, StoragePaths, TransportPolicy};
|
|
use tsunagi::dataplane::IpPlugin;
|
|
use tsunagi::dataplane::wireguard::{MemoryTunFactory, WireguardConfig, WireguardPlugin};
|
|
use tsunagi::identity::{NetworkName, NetworkSecret};
|
|
use tsunagi::{Agent, Result};
|
|
|
|
#[tokio::main]
|
|
async fn main() -> Result<()> {
|
|
let paths = StoragePaths::user_default()?;
|
|
|
|
// MemoryTunFactory needs no privileges; swap in SystemTunFactory for a
|
|
// real interface.
|
|
let plugin = WireguardPlugin::open(
|
|
WireguardConfig::new(paths.state_dir.join("wireguard")),
|
|
Arc::new(MemoryTunFactory::new()),
|
|
)
|
|
.await
|
|
.expect("wireguard plugin");
|
|
|
|
let agent = Agent::spawn(
|
|
AgentConfig::new(paths)
|
|
.with_transport(TransportPolicy::N0Defaults)
|
|
.with_plugin(plugin.clone() as Arc<dyn IpPlugin>),
|
|
)
|
|
.await?;
|
|
|
|
let network = agent
|
|
.join_network(&NetworkName::new("lab")?, &NetworkSecret::generate())
|
|
.await?;
|
|
|
|
if let Some(view) = plugin.overview(network) {
|
|
println!("{} on {}", view.interface, view.overlay_address);
|
|
}
|
|
agent.shutdown().await;
|
|
Ok(())
|
|
}
|
|
```
|
|
|
|
## Limits and future work
|
|
|
|
* **Full mesh only.** Every member runs a tunnel to every other member.
|
|
Routing through an intermediate participant is not implemented.
|
|
* **IPv4 is opt-in, must be agreed, and can collide.** See above. A proper
|
|
allocator needs agreed state.
|
|
* **No routes, DNS or firewall rules.** The plugin creates its interface and
|
|
nothing else. Anything beyond the overlay `/64` is the operator's business.
|
|
* **Membership is session-scoped.** A peer leaves the overlay when its control
|
|
session ends; surviving a long absence is the same future work.
|
|
* **Userspace costs CPU.** Kernel WireGuard is faster. A kernel backend could
|
|
return behind the same boundary, but it would give up transport-provided NAT
|
|
traversal unless paired with a local proxy.
|
|
* **A persistent TUN interface needs `keep_addr_on_down`.** Without a process
|
|
attached it has no carrier, and Linux then flushes its IPv6 addresses. The
|
|
setup printed by `tsunagi tun-setup` sets it; the agent checks the address is
|
|
present *and usable* — not tentative, not DAD-failed — before attaching, and
|
|
reports what it actually found.
|
|
* **The agent cannot assign the overlay address itself.** The `tun` crate sets
|
|
addresses through an IPv4-only ioctl, so the IPv6 overlay address must come
|
|
from `ip -6 address add` or an equivalent. The agent verifies the address is
|
|
present, via `/proc/net/if_inet6`, and refuses with the exact command rather
|
|
than running an interface that could never receive anything. Doing it
|
|
in-process would mean speaking netlink, which is not implemented.
|
|
* **The system interface path is not exercised by the default suite**, because
|
|
it needs privileges. Everything else about the data plane is.
|