IPv6 works end to end between two machines; IPv4 silently did not, and the
agent said nothing useful about why.
Allocation moved the address from something derivable before startup to
something agreed at run time, so an interface configured by an earlier
`tun-setup` carries a different address than the one allocated. The kernel
then sends packets with that stale source and every peer drops them as not
belonging to us — correct behaviour, invisible cause. Meanwhile pings to
our own allocated address fall into the tunnel and land in the "nobody
owns this" counter.
The agent now checks whether its allocated address is assigned anywhere on
the host — by binding a UDP socket to it, which needs no privileges and no
platform code — and reports the exact `ip address add` command until it
is, mentioning that another address of the range has to go.
`tun-setup` no longer prints a derived IPv4 address, because that number
is now wrong by construction. It says the agent will print the real one.
The unroutable counter keeps one destination as a sample, in status output
too. A bare count says something is wrong; the address says what.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Derived IPv4 addresses could not survive anything: they changed with the
range, and there was no way for a member to come back to the one it had.
Addresses are now allocated and recorded as signed facts, which is the
first slice of the model in docs/sync-model.md.
src/state/ holds one record per author per network, carrying that author's
complete current statement, signed with its persistent device key over a
length-prefixed canonical encoding. Merging follows the model's rules: a
higher version wins, an older one never rolls back a newer, duplicates are
idempotent, absence from a snapshot is not deletion, and a same-version
conflict is resolved identically on every replica and reported rather than
letting replicas diverge. Records are persisted in state.sqlite, with the
record and the author's version counter committed in one transaction
before anything is announced, and distributed as a State control message
that is merged into what the receiver already holds.
No vote, deliberately, despite the request. A majority is not a trust root
here — anyone with the secret can mint identities — and a quorum would
stall with one peer online and diverge across a partition. Signatures plus
a deterministic merge converge without either failure mode: two members
claiming one address at once are resolved by the lower endpoint id, and
the loser allocates again with a higher version.
The range moved from the plugin to the agent, defaults to 10.13.37.0/24,
and is now agreed rather than configured per member: a joining agent
adopts what the network already uses, so --ipv4-range only matters for
whoever starts it. The announcement went back to identity only (version 3)
since the range travels in signed records now.
A release tombstone exists and merges correctly, but nothing emits one
yet.
116 tests. The headline ones: an address survives restarting both agents,
three members get three distinct addresses, and a member started with a
different range adopts the one in use. Confirmed by hand with two CLI
agents restarted end to end.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
100.64.0.0/10 was a bad default: it is exactly Tailscale's range, and
carrier-grade NAT's. There is no IPv4 range that is free on every host, so
there is now no default at all — IPv4 is off until --ipv4-range names one.
IPv6 is unaffected and still works out of the box, because a ULA derived
from the network id collides with essentially nothing.
The more serious problem this exposed: the range is an input to the address
derivation, and each agent derives every peer's address itself. Two members
configured with different ranges would therefore derive different addresses
for each other and IPv4 would silently misroute. So the range now travels
in the announcement — not as a request and never trusted, only so the
mismatch is seen. A peer whose range disagrees gets no IPv4 address here,
keeps working over IPv6, and the reason is reported with both ranges named.
The announcement format goes to version 2. postcard is not
self-describing, so an older peer cannot read it; the version check already
catches that and now says which side needs updating.
The (Ipv4Addr, u8) tuple that had spread across six modules is now an
Ipv4Range with validation, Display and FromStr, so a bad --ipv4-range is
refused with a reason instead of being accepted and misbehaving later. It
is also rejected when passed without --wireguard rather than ignored.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every member now also derives an IPv4 address, from the same inputs as its
IPv6 one, into 100.64.0.0/10 by default. The range is configurable and IPv4
can be turned off with --no-ipv4.
IPv4 is honestly weaker than IPv6 here and the code says so. A 64 bit
interface identifier makes an IPv6 collision impossible in practice; IPv4
has nothing like that room, and in a /10 with 50 members two will derive the
same address about 0.03% of the time. A mesh with no coordinator cannot
allocate around that, so a collision is detected and resolved instead: the
member whose public key sorts lower keeps the address, a rule every member
computes identically and therefore agrees on. The other keeps IPv6 and is
flagged in the status. IPv6 always works; IPv4 almost always works and
degrades predictably.
Routing and address-ownership enforcement now cover both families: a packet
goes to the peer that owns its destination, and a decrypted packet is
dropped unless its source is an address derived for the peer that sent it,
IPv4 included.
Six new tests, among them a real IPv4 packet crossing a tunnel next to an
IPv6 one, a spoofed IPv4 source being dropped, and an IPv6-only overlay.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The setup recipe failed with a missing sysctl directory and "RTNETLINK
answers: Invalid argument". The cause was the default MTU of 1100.
IPv6 requires a minimum MTU of 1280 (RFC 8200) and Linux enforces it by
tearing IPv6 down on any interface below it: the per-device
/proc/sys/net/ipv6/conf entries disappear and an address can no longer be
assigned. Evidence on the test host: every interface at 1280 or above has
an IPv6 conf directory, every interface below it (1230, 1100) has none.
So the overlay MTU is now 1280, which is also the floor. A smaller value is
refused when the plugin opens, naming the reason, rather than surfacing as
an obscure netlink error after the user has already run four commands.
That leaves no slack against the other constraint: a packet needs mtu + 32
bytes of transport datagram, so 1312. A direct QUIC path offers roughly
1380 and fits; a relayed path may not, so the plugin now reports the exact
numbers when a link cannot carry a full-size packet, instead of only
counting silent drops.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The setup this tool printed did not work, and the agent then correctly
refused to start. A persistent TUN interface has no carrier until a process
attaches to it, and Linux flushes IPv6 addresses from an interface that
loses carrier unless net.ipv6.conf.<dev>.keep_addr_on_down is set, which it
is not by default. So `ip -6 address add` on a freshly created interface
silently lost the address before the agent ever ran.
The recipe now brings the link up first, sets keep_addr_on_down, and adds
the address with `nodad` — without which duplicate address detection can
never finish on an interface with no carrier and the address stays
tentative and unusable.
The agent's own retry loop made this worse: it attached, failed the address
check, dropped the device and toggled the carrier, which flushed the
address again. The check now runs before attaching to an existing
interface, so looking is not destructive.
Failures are self-diagnosing now: the check parses the IFA_F_* flags, tells
tentative and DAD-failed apart from missing, and lists the addresses the
interface actually has.
Four new tests, including one that reads this host's real /proc/net/if_inet6
and one that pins the ordering of the setup commands.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Creating a network interface needs CAP_NET_ADMIN, but that is a one-time
setup step rather than something the agent must hold for its whole life.
SystemTunFactory now attaches to an interface that already exists and only
creates one when it does not. A persistent interface created by root and
owned by the user therefore lets the agent run with no privileges and no
capabilities at all. When attaching, nothing is reconfigured, since doing so
would need exactly the privileges we are avoiding.
New `tsunagi tun-setup` prints the three commands to run once as root,
resolving the derived interface name and overlay address for the network.
This also fixes a real gap: the overlay address was passed to the factory
and thrown away, so an interface the agent created had no address and could
never have received anything. The `tun` crate sets addresses through an
IPv4-only ioctl and cannot assign an IPv6 one at all, so the agent now
verifies the address is present via /proc/net/if_inet6 and refuses with the
exact command to run instead of coming up broken. Doing it in-process would
mean speaking netlink, which is not implemented and is recorded as such.
Not verified on this machine: no sudo is available here, so the privileged
setup and the attach path were not executed end to end.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"n0" is Number 0, the company behind iroh, and the name leaked from iroh's
own preset into this project's user interface, where it explains nothing.
The value is now --transport relay, which says what it does; n0 stays as an
accepted alias.
Also spells out, in the CLI help, the README, the threat model and the
TransportPolicy docs, whose infrastructure is involved: address records are
published to and resolved from dns.iroh.link, and the fallback relays are
Number 0's, in the US, EU and Asia-Pacific.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Running `tsunagi up` twice with the same arguments failed with "network ...
is already active", and then dropped the iroh endpoint without closing it.
A configured network is activated automatically at startup, so the second
run found it already up. `join_network` is declarative — "be a member of
this network" — so joining one that is already active now succeeds and
changes nothing. `activate_network` stays strict for callers that
specifically want to know whether an inactive network was started.
The CLI now closes the agent on the error path too, and handles SIGTERM as
well as Ctrl-C, so a service manager stopping the agent gets the same clean
shutdown an interactive user does.
Also documents the two lookups people conflate: resolving one endpoint's
address is iroh's public pkarr/DNS service and works today, which is why
`--peer <endpoint-id>` needs no address; finding who is in a network is this
project's `NetworkDiscovery` and is still static bootstrap only. Notes in
the README and the threat model that `n0` and `direct` publish this
endpoint's addresses to a public third-party service.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Corrects the architecture on two points raised in review, while the project
is still small enough to change cheaply.
1. Control and data are separated *logically*, not physically.
The old reading — "nothing but control may ride on iroh" — threw away iroh's
whole value and would have forced the data plane to reimplement STUN, ICE and
a relay. Now both planes ride on iroh with different ALPNs and different
connections, so the data plane inherits hole punching and relay fallback,
while proto/ still knows nothing about packets and dataplane/ knows nothing
about the control protocol.
New boundary: PacketTransport / PacketLink, an authenticated unreliable
datagram channel per (network, peer, protocol). tsunagi/data/1 runs the same
membership handshake, then DataOpen/DataOpenAck, then QUIC datagrams. Only
the smaller endpoint id dials, so exactly one link exists per pair.
A plugin is handed links and never learns reachability, so the WireGuard
announcement shrank to a public key: there is no address left to lie about.
2. WireGuard now runs in userspace, on boringtun's protocol state machine.
No kernel module, no wg tool, no ip shell-out, no loopback proxy: the wgtool,
backend and bridge modules are gone. Only creating a TUN device needs
privileges, and that sits behind TunFactory, so the entire data plane —
handshake, encryption, routing, address ownership — is tested with none.
Address ownership is enforced rather than believed: outbound packets go to
the owner of the destination address, inbound packets are dropped unless
their source is the address derived for the peer that sent them.
3. A `tsunagi` binary: secret, doctor, id, up. It owns the runtime, the
logging subscriber and Ctrl-C, which the library still refuses to.
Also fixes a reference cycle where IrohTransport held Arc<Inner>, which kept
the databases open and the directory lock held after shutdown; two storage
tests caught it once the cycle existed.
81 tests pass offline with no privileges, including real IPv6 packets
crossing a real WireGuard tunnel over real iroh connections. Verified by
hand: two CLI processes forming a mesh both on loopback and via n0 discovery
using only an endpoint id.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The first IP plugin, built on the data plane boundary the core already had.
Plugin:
- one X25519 key per network in the plugin's own wireguard.sqlite, separate
from the iroh identity and from the network secret; a damaged store is an
error, never a silently regenerated identity
- deterministic IPv6 ULA overlay: every member derives the same /64 from the
network id and its own /128 from its WireGuard public key, so no
coordinator allocates addresses
- AllowedIPs are derived locally, never taken from a peer's announcement, so
a member cannot claim another member's overlay address; a mismatched claim
is rejected
- bounded, versioned, validated announcement carried as the existing opaque
capability payload, which the core still never parses
- each agent builds its own full-mesh configuration (N-1 peers) and
reconciles on every change and on a timer, repairing drift
- WireguardBackend abstraction: RecordingBackend in memory, and WgToolBackend
driving real wg/ip on Linux, split into a pure planner plus parsers and a
thin executor so everything interesting is testable without root
Core, three generic additions the plugin needed:
- IpPlugin::on_network_activated, so per-network state is ready before peers
- PluginContext for re-announcements and error reports from plugin tasks,
with errors counted by the owning network runtime
- IpPlugin::shutdown, awaited with a grace period, so system objects go away
94 tests pass offline with no privileges: 35 new WireGuard unit tests and 12
integration tests over real iroh connections. The real wg/ip backend needs
root and is behind --ignored in tests/wireguard_system.rs; it was not run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Working library with real iroh connections, not an interface sketch:
- persistent device identity in state.sqlite, stable across restarts
- deterministic network space derived from name + secret via HKDF-SHA256,
with frozen labels and unambiguous length-prefixed encoding
- replaceable discovery returning unverified candidates only; static
bootstrap, in-memory test backend and a composite
- real iroh connections plus an explicit mutual membership proof:
HMAC-SHA256 over a role-separated transcript bound to the TLS exporter,
the network id and both endpoint identities
- small versioned control protocol: handshake, announcement, ping/pong
- multiple networks per agent with enforced isolation
- automatic reconnect with bounded backoff and jitter
- mandatory state vs disposable cache, with a real directory ownership lock
- status snapshots, event stream and honest diagnostics
47 integration and unit tests cover the required scenarios offline on
loopback. Snapshots, revocations and WireGuard are designed for and
documented, not implemented.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>