Twice now a report has read "dns not serving" and been taken for a broken resolver. It was accurate both times: the agent had been started without `--dns`. That is the flag's fault, not the reader's — a resolver that disappears because one word was not retyped is worse than none, since the names simply stop working. So the setting belongs to the device now: `--dns` turns it on and it stays on, `tsunagi dns off` turns it off, and `tsunagi dns on` turns it on for an agent that is already running, without restarting it. `tsunagi dns` says what it is doing, or what it will do at the next start when nothing is running. One agent has one identity and as many networks as it likes, so one DNS service serves them all: each network is a zone named after it, and joining or leaving one changes what resolves with no restart. A question carries a name and not the network it belongs to, so the suffix decides and nothing is shared between zones — a member of one network is not a name in another. `--dns-zone` is gone with that: there is no single zone to name any more, and a network name may contain dots, so `--network lab.internal` is how you get `music.lab.internal`. It listens on loopback only, where it always could have. Binding the overlay address put the zones in front of the whole mesh, and with several networks on one agent that would have answered one network's questions about another's names. That made a gap plain: a second network on an agent had no addresses at all, because the configured range belongs to whichever network took it first, so its members had nothing to allocate from and no names to answer with. A second network now uses the range **derived from its own network id** — every member derives the same one from something they all already have, so it is an agreement rather than a local invention. It is held back for a moment first, because a network that already exists has a range of its own and a joiner should adopt it rather than argue; that wait is what keeps "the first member settles it" true. And a network needs no ceremony to start. `tsunagi up --network lab` with no secret resolves the obvious way: the one network of that name this device already has, or — when there is none — a fresh random secret, printed in full with the single line to send the others. That is the ad-hoc case, one person makes a network and passes the command round, and it was previously two steps with a flag people could not find. The secret is printed only when the agent invented it, because then there is nowhere else to read it from; one that was supplied is not echoed. Two networks of one name and no secret is the one case with no answer, and it says so rather than choosing. Releasing now also stops this agent claiming again. The periodic check would otherwise publish a fresh claim in the moment between the goodbye and the teardown, turning a release into a hello nobody asked for. Exercised with the real binary: a zone per network as a second one is joined into a running agent, the resolver switched on and off while it runs, and an ad-hoc network printing its secret and the line to share. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
7.4 KiB
Testing
How to run everything is in ../README.md. Rules for writing tests are in ../AGENTS.md.
Ground rules
Tests use real iroh endpoints on loopback, real handshakes, real SQLite in per-test temporary directories, and independent agent instances. Discovery is substitutable; iroh, authentication, message passing and persistent storage are not.
The suite runs with no internet, no DHT, no public relay, no administrator
rights and no changes to OS network settings: endpoints bind 127.0.0.1:0 and
[::1]:0, relays are disabled, address lookup is cleared, port mapping is
disabled, and net-report probing is reduced to its minimum.
Synchronisation is always "wait for a specific event or condition under one
overall deadline" (wait_event, wait_until, 30 s). The deadline covers the
probe as well as the gaps between probes: a call into a wedged agent that
never answers fails the test rather than hanging the process, which is the
difference between a red run and a test binary still burning a core the next
day. settle() exists only for asserting that something did not happen.
Ports are dynamic and directories are isolated, so tests run in parallel.
Several library instances in one process is exactly that. It is not a test of several system processes, and is not presented as one.
What is covered
| # | scenario | file |
|---|---|---|
| 1 | deterministic identity: same name + secret ⇒ same space on different devices; a changed name or secret changes it; hostname, device key and restart do not | tests/identity.rs |
| 2 | four agents find each other, authenticate for real and exchange distinguishable messages; a late joiner is picked up; opaque plugin capabilities cross the control plane | tests/multi_peer.rs |
| 3 | an attacker who knows the address and the correct public NetworkId but not the secret is rejected at the handshake |
tests/authentication.rs |
| 4 | one agent in two networks: statuses and messages do not mix; a session authenticated for one network cannot speak for the other; deactivating one leaves the other running | tests/network_isolation.rs |
| 5 | full stop and recreation from the same database: identity and settings survive, sessions come back automatically, a new local UDP port does not break recovery | tests/restart.rs |
| 6 | changing the secret through the library API: device identity survives, old sessions and messages get no access to the new space, and the retired space stays retired across a restart | tests/restart.rs |
| 7 | missing, corrupt and stale cache do not prevent connecting; a corrupt mandatory store is a clear error and never a fresh identity; a newer schema is refused; secrets stay out of status and Debug |
tests/cache_and_state.rs |
| 8 | a dead candidate and a vanished peer do not block the others; retries are bounded and stop when the network is deactivated | tests/resilience.rs |
| 9 | wrong version, a message before authentication, a proof replayed on another connection, an oversized frame and a Hello for an inactive network are all rejected without taking the agent down |
tests/authentication.rs |
| 10 | a second agent on the same state directory gets a clear error; after a clean stop the directory reopens; shutdown ends background tasks and refuses further work; independent agents coexist in one process | tests/resilience.rs |
| 11 | leaving a network frees the address for the others, says plainly when there was nobody to tell, and rejoining afterwards is not mistaken for a stale record; a wipe empties both directories and the next start is a stranger, while a directory that is not ours is refused | tests/leaving.rs, tests/cache_and_state.rs |
| 12 | a network can be joined into a running agent over the control socket and is live at once; a name with no secret resumes the one network of that name, invents one when there is none, and refuses to choose between two; two networks on one agent each get a range of their own | tests/local_control.rs, tests/network_isolation.rs, CLI unit tests |
tests/wireguard.rs drives the WireGuard data plane over real iroh
connections. Everything is real except the packet interface: real agents, real
control plane, real data links, real WireGuard handshakes and encryption from
boringtun, with an in-memory TUN device so none of it needs privileges. It
covers real IPv6 packets travelling both ways through a tunnel, a three-agent
mesh, a peer that sends from an address it does not own being dropped, packets
for unowned addresses being counted rather than broadcast, a departing peer
losing its tunnel, two networks keeping separate interfaces and keys, restart
keeping the WireGuard identity, shutdown removing every interface, a forged
overlay claim being rejected, and the core carrying the payload without
interpreting it.
Unit tests in crates/tsunagi/src/state/ cover the signed record model directly: tampering
with any field breaks verification, a newer version wins while an older one
never rolls back, two authors claiming one address resolve the same way no
matter the merge order, one key used in two places is reported rather than
silently merged, a release survives a late-arriving old claim, a bad record in
a batch does not stop the rest, and allocation is deterministic, spread out,
walks past everything taken and reports a full range instead of handing out a
duplicate.
tests/local_control.rs covers the local control socket end to end: a client
asking a running agent for status over a real Unix socket, joining and
leaving a network through it, a leftover socket file being replaced while a
live one is not, and the derived socket path staying short enough to bind.
crates/tsunagi-cli/tests/dns_service.rs runs the real binary: the resolver
comes up with no interface to attach it to, the listener is not rebuilt on
the way past, a name outside every zone is refused, each network gets a zone
of its own as it is joined, and the resolver can be switched on and off
while the agent runs.
tests/discovery.rs covers the discovery contract itself: a static bootstrap
candidate is enough to join, several backends compose, entries are withdrawn
when a network stops, and a forgotten network stays forgotten across a restart.
Unit tests in crates/tsunagi/src/proto/handshake.rs cover the transcript construction
itself: role separation, channel binding, identity and network binding,
unambiguous encoding, and rejection under the wrong key.
Unit tests in crates/tsunagi-wg-quic/src/ cover key clamping against the RFC
7748 vector, overlay derivation, announcement validation including the
address-hijack attempt, interface naming, and IP header parsing against
truncated and nonsense input.
tests/end_to_end.rs is the vertical slice: persistent identity → network
space → discovery → iroh → authentication → message exchange.
What the default suite does not cover is the real TUN interface, because
that needs CAP_NET_ADMIN. Everything above it does run.
Not covered, and not claimed to be
Listed in sync-model.md: snapshots, revocations, long partitions, hostname renames, recovery of a returning participant, NAT traversal, relay fallback, and multi-process or multi-host deployment. None of these are implemented, and none are marked as passing.
Debugging a test
TSUNAGI_TEST_LOG=tsunagi=debug cargo test --test multi_peer -- --nocapture
The library never installs a global subscriber; the harness opts in only when that variable is set.