Nine orphaned test processes were found still running from the day before,
three of them spinning on a core each for twenty hours. The code they ran
is several changes old and the mesh test passes twenty times over now, so
the wedge itself is gone — but nothing in the way it was waited on was
bounded, which is why a wedge lasted a day instead of failing a run.
The harness enforced its deadline only between probes. A probe that never
returned — one call into a wedged runtime, which is exactly what a status
request is — waited for ever inside the deadline it was supposed to obey.
The probe is now bounded too, so the same wedge fails the test in thirty
seconds.
Shutdown claimed to be bounded and was not. The plugins had a grace
period; the network runtimes, the accept loop, the plugin request loop and
the endpoint close did not, and a peer that stops reading is enough to
hold any of them open. Each now gets a grace period and is aborted after
it. The overlay packet loop was not stopped at all: it ends when the
device reports end of stream, which a live interface never does, so it
outlived the interface it was reading. And a plugin's grace period
abandoned the future without stopping the task behind it, so the helper
is public and `wg-quic` uses it on its own runtime.
The local control socket was unbounded in both directions. A wedged agent
left `tsunagi status` hanging with nothing on screen and no way out but
Ctrl-C; it now says the agent did not answer, after five seconds, and
falls back to the state store as it already did for a socket that refuses
a connection. On the serving side, a connection that sends no request no
longer holds a task open.
Tests cover the mechanism — a task that stops on its own is not aborted,
one that ignores the grace is cut off and drops what it held — and both
sides of the change in behaviour: a probe that never answers fails its
deadline, and a silent agent is reported rather than waited out.
Also: the binary opts out of rustdoc, since it shares a name with the
library and `cargo doc` cannot put both in one directory.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The module table still had the plugin inside the core and no mention of
the overlay or the DNS view, and the stale path in the testing notes
pointed at a directory that had moved.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
First step of separating the layers. The library and the binary are now
crates/tsunagi and crates/tsunagi-cli, which means the plugin crate to
come can be told apart from the core by the compiler rather than by
discipline.
Falls out of it immediately: the CLI's dependencies stop being features
of the library. clap, anstream and tracing-subscriber were optional
dependencies behind a `cli` feature that every library user had to
remember to turn off; now they belong to the crate that uses them, and
the library defaults to no features at all.
The one test that drives the binary moved beside it — a library cannot
depend on a binary built from a crate that depends on the library — and
was rewritten against the public API instead of the test harness.
AGENTS.md said to prefer one crate. It now says the system level and its
plugins are separate crates, for the reason above, and that everything
else stays one crate.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Derived IPv4 addresses could not survive anything: they changed with the
range, and there was no way for a member to come back to the one it had.
Addresses are now allocated and recorded as signed facts, which is the
first slice of the model in docs/sync-model.md.
src/state/ holds one record per author per network, carrying that author's
complete current statement, signed with its persistent device key over a
length-prefixed canonical encoding. Merging follows the model's rules: a
higher version wins, an older one never rolls back a newer, duplicates are
idempotent, absence from a snapshot is not deletion, and a same-version
conflict is resolved identically on every replica and reported rather than
letting replicas diverge. Records are persisted in state.sqlite, with the
record and the author's version counter committed in one transaction
before anything is announced, and distributed as a State control message
that is merged into what the receiver already holds.
No vote, deliberately, despite the request. A majority is not a trust root
here — anyone with the secret can mint identities — and a quorum would
stall with one peer online and diverge across a partition. Signatures plus
a deterministic merge converge without either failure mode: two members
claiming one address at once are resolved by the lower endpoint id, and
the loser allocates again with a higher version.
The range moved from the plugin to the agent, defaults to 10.13.37.0/24,
and is now agreed rather than configured per member: a joining agent
adopts what the network already uses, so --ipv4-range only matters for
whoever starts it. The announcement went back to identity only (version 3)
since the range travels in signed records now.
A release tombstone exists and merges correctly, but nothing emits one
yet.
116 tests. The headline ones: an address survives restarting both agents,
three members get three distinct addresses, and a member started with a
different range adopts the one in use. Confirmed by hand with two CLI
agents restarted end to end.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every member now also derives an IPv4 address, from the same inputs as its
IPv6 one, into 100.64.0.0/10 by default. The range is configurable and IPv4
can be turned off with --no-ipv4.
IPv4 is honestly weaker than IPv6 here and the code says so. A 64 bit
interface identifier makes an IPv6 collision impossible in practice; IPv4
has nothing like that room, and in a /10 with 50 members two will derive the
same address about 0.03% of the time. A mesh with no coordinator cannot
allocate around that, so a collision is detected and resolved instead: the
member whose public key sorts lower keeps the address, a rule every member
computes identically and therefore agrees on. The other keeps IPv6 and is
flagged in the status. IPv6 always works; IPv4 almost always works and
degrades predictably.
Routing and address-ownership enforcement now cover both families: a packet
goes to the peer that owns its destination, and a decrypted packet is
dropped unless its source is an address derived for the peer that sent it,
IPv4 included.
Six new tests, among them a real IPv4 packet crossing a tunnel next to an
IPv6 one, a spoofed IPv4 source being dropped, and an IPv6-only overlay.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Corrects the architecture on two points raised in review, while the project
is still small enough to change cheaply.
1. Control and data are separated *logically*, not physically.
The old reading — "nothing but control may ride on iroh" — threw away iroh's
whole value and would have forced the data plane to reimplement STUN, ICE and
a relay. Now both planes ride on iroh with different ALPNs and different
connections, so the data plane inherits hole punching and relay fallback,
while proto/ still knows nothing about packets and dataplane/ knows nothing
about the control protocol.
New boundary: PacketTransport / PacketLink, an authenticated unreliable
datagram channel per (network, peer, protocol). tsunagi/data/1 runs the same
membership handshake, then DataOpen/DataOpenAck, then QUIC datagrams. Only
the smaller endpoint id dials, so exactly one link exists per pair.
A plugin is handed links and never learns reachability, so the WireGuard
announcement shrank to a public key: there is no address left to lie about.
2. WireGuard now runs in userspace, on boringtun's protocol state machine.
No kernel module, no wg tool, no ip shell-out, no loopback proxy: the wgtool,
backend and bridge modules are gone. Only creating a TUN device needs
privileges, and that sits behind TunFactory, so the entire data plane —
handshake, encryption, routing, address ownership — is tested with none.
Address ownership is enforced rather than believed: outbound packets go to
the owner of the destination address, inbound packets are dropped unless
their source is the address derived for the peer that sent them.
3. A `tsunagi` binary: secret, doctor, id, up. It owns the runtime, the
logging subscriber and Ctrl-C, which the library still refuses to.
Also fixes a reference cycle where IrohTransport held Arc<Inner>, which kept
the databases open and the directory lock held after shutdown; two storage
tests caught it once the cycle existed.
81 tests pass offline with no privileges, including real IPv6 packets
crossing a real WireGuard tunnel over real iroh connections. Verified by
hand: two CLI processes forming a mesh both on loopback and via n0 discovery
using only an endpoint id.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The first IP plugin, built on the data plane boundary the core already had.
Plugin:
- one X25519 key per network in the plugin's own wireguard.sqlite, separate
from the iroh identity and from the network secret; a damaged store is an
error, never a silently regenerated identity
- deterministic IPv6 ULA overlay: every member derives the same /64 from the
network id and its own /128 from its WireGuard public key, so no
coordinator allocates addresses
- AllowedIPs are derived locally, never taken from a peer's announcement, so
a member cannot claim another member's overlay address; a mismatched claim
is rejected
- bounded, versioned, validated announcement carried as the existing opaque
capability payload, which the core still never parses
- each agent builds its own full-mesh configuration (N-1 peers) and
reconciles on every change and on a timer, repairing drift
- WireguardBackend abstraction: RecordingBackend in memory, and WgToolBackend
driving real wg/ip on Linux, split into a pure planner plus parsers and a
thin executor so everything interesting is testable without root
Core, three generic additions the plugin needed:
- IpPlugin::on_network_activated, so per-network state is ready before peers
- PluginContext for re-announcements and error reports from plugin tasks,
with errors counted by the owning network runtime
- IpPlugin::shutdown, awaited with a grace period, so system objects go away
94 tests pass offline with no privileges: 35 new WireGuard unit tests and 12
integration tests over real iroh connections. The real wg/ip backend needs
root and is behind --ignored in tests/wireguard_system.rs; it was not run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Working library with real iroh connections, not an interface sketch:
- persistent device identity in state.sqlite, stable across restarts
- deterministic network space derived from name + secret via HKDF-SHA256,
with frozen labels and unambiguous length-prefixed encoding
- replaceable discovery returning unverified candidates only; static
bootstrap, in-memory test backend and a composite
- real iroh connections plus an explicit mutual membership proof:
HMAC-SHA256 over a role-separated transcript bound to the TLS exporter,
the network id and both endpoint identities
- small versioned control protocol: handshake, announcement, ping/pong
- multiple networks per agent with enforced isolation
- automatic reconnect with bounded backoff and jitter
- mandatory state vs disposable cache, with a real directory ownership lock
- status snapshots, event stream and honest diagnostics
47 integration and unit tests cover the required scenarios offline on
loopback. Snapshots, revocations and WireGuard are designed for and
documented, not implemented.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>