Nine orphaned test processes were found still running from the day before, three of them spinning on a core each for twenty hours. The code they ran is several changes old and the mesh test passes twenty times over now, so the wedge itself is gone — but nothing in the way it was waited on was bounded, which is why a wedge lasted a day instead of failing a run. The harness enforced its deadline only between probes. A probe that never returned — one call into a wedged runtime, which is exactly what a status request is — waited for ever inside the deadline it was supposed to obey. The probe is now bounded too, so the same wedge fails the test in thirty seconds. Shutdown claimed to be bounded and was not. The plugins had a grace period; the network runtimes, the accept loop, the plugin request loop and the endpoint close did not, and a peer that stops reading is enough to hold any of them open. Each now gets a grace period and is aborted after it. The overlay packet loop was not stopped at all: it ends when the device reports end of stream, which a live interface never does, so it outlived the interface it was reading. And a plugin's grace period abandoned the future without stopping the task behind it, so the helper is public and `wg-quic` uses it on its own runtime. The local control socket was unbounded in both directions. A wedged agent left `tsunagi status` hanging with nothing on screen and no way out but Ctrl-C; it now says the agent did not answer, after five seconds, and falls back to the state store as it already did for a socket that refuses a connection. On the serving side, a connection that sends no request no longer holds a task open. Tests cover the mechanism — a task that stops on its own is not aborted, one that ignores the grace is cut off and drops what it held — and both sides of the change in behaviour: a probe that never answers fails its deadline, and a silent agent is reported rather than waited out. Also: the binary opts out of rustdoc, since it shares a name with the library and `cargo doc` cannot put both in one directory. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
104 lines
6.4 KiB
Markdown
104 lines
6.4 KiB
Markdown
# Testing
|
|
|
|
How to run everything is in [../README.md](../README.md). Rules for writing
|
|
tests are in [../AGENTS.md](../AGENTS.md).
|
|
|
|
## Ground rules
|
|
|
|
Tests use real iroh endpoints on loopback, real handshakes, real SQLite in
|
|
per-test temporary directories, and independent agent instances. Discovery is
|
|
substitutable; **iroh, authentication, message passing and persistent storage
|
|
are not**.
|
|
|
|
The suite runs with no internet, no DHT, no public relay, no administrator
|
|
rights and no changes to OS network settings: endpoints bind `127.0.0.1:0` and
|
|
`[::1]:0`, relays are disabled, address lookup is cleared, port mapping is
|
|
disabled, and net-report probing is reduced to its minimum.
|
|
|
|
Synchronisation is always "wait for a specific event or condition under one
|
|
overall deadline" (`wait_event`, `wait_until`, 30 s). The deadline covers the
|
|
probe as well as the gaps between probes: a call into a wedged agent that
|
|
never answers fails the test rather than hanging the process, which is the
|
|
difference between a red run and a test binary still burning a core the next
|
|
day. `settle()` exists only for asserting that something did *not* happen.
|
|
Ports are dynamic and directories are isolated, so tests run in parallel.
|
|
|
|
Several library instances in one process is exactly that. It is **not** a test
|
|
of several system processes, and is not presented as one.
|
|
|
|
## What is covered
|
|
|
|
| # | scenario | file |
|
|
|---|---|---|
|
|
| 1 | deterministic identity: same name + secret ⇒ same space on different devices; a changed name or secret changes it; hostname, device key and restart do not | `tests/identity.rs` |
|
|
| 2 | four agents find each other, authenticate for real and exchange distinguishable messages; a late joiner is picked up; opaque plugin capabilities cross the control plane | `tests/multi_peer.rs` |
|
|
| 3 | an attacker who knows the address and the correct public `NetworkId` but not the secret is rejected at the handshake | `tests/authentication.rs` |
|
|
| 4 | one agent in two networks: statuses and messages do not mix; a session authenticated for one network cannot speak for the other; deactivating one leaves the other running | `tests/network_isolation.rs` |
|
|
| 5 | full stop and recreation from the same database: identity and settings survive, sessions come back automatically, a new local UDP port does not break recovery | `tests/restart.rs` |
|
|
| 6 | changing the secret through the library API: device identity survives, old sessions and messages get no access to the new space, and the retired space stays retired across a restart | `tests/restart.rs` |
|
|
| 7 | missing, corrupt and stale cache do not prevent connecting; a corrupt mandatory store is a clear error and never a fresh identity; a newer schema is refused; secrets stay out of status and `Debug` | `tests/cache_and_state.rs` |
|
|
| 8 | a dead candidate and a vanished peer do not block the others; retries are bounded and stop when the network is deactivated | `tests/resilience.rs` |
|
|
| 9 | wrong version, a message before authentication, a proof replayed on another connection, an oversized frame and a `Hello` for an inactive network are all rejected without taking the agent down | `tests/authentication.rs` |
|
|
| 10 | a second agent on the same state directory gets a clear error; after a clean stop the directory reopens; shutdown ends background tasks and refuses further work; independent agents coexist in one process | `tests/resilience.rs` |
|
|
|
|
`tests/wireguard.rs` drives the WireGuard data plane over real iroh
|
|
connections. Everything is real except the packet interface: real agents, real
|
|
control plane, real data links, real WireGuard handshakes and encryption from
|
|
boringtun, with an in-memory TUN device so none of it needs privileges. It
|
|
covers real IPv6 packets travelling both ways through a tunnel, a three-agent
|
|
mesh, a peer that sends from an address it does not own being dropped, packets
|
|
for unowned addresses being counted rather than broadcast, a departing peer
|
|
losing its tunnel, two networks keeping separate interfaces and keys, restart
|
|
keeping the WireGuard identity, shutdown removing every interface, a forged
|
|
overlay claim being rejected, and the core carrying the payload without
|
|
interpreting it.
|
|
|
|
Unit tests in `crates/tsunagi/src/state/` cover the signed record model directly: tampering
|
|
with any field breaks verification, a newer version wins while an older one
|
|
never rolls back, two authors claiming one address resolve the same way no
|
|
matter the merge order, one key used in two places is reported rather than
|
|
silently merged, a release survives a late-arriving old claim, a bad record in
|
|
a batch does not stop the rest, and allocation is deterministic, spread out,
|
|
walks past everything taken and reports a full range instead of handing out a
|
|
duplicate.
|
|
|
|
`tests/local_control.rs` covers the local control socket end to end: a client
|
|
asking a running agent for status over a real Unix socket, a leftover socket
|
|
file being replaced while a live one is not, and the derived socket path
|
|
staying short enough to bind.
|
|
|
|
`tests/discovery.rs` covers the discovery contract itself: a static bootstrap
|
|
candidate is enough to join, several backends compose, entries are withdrawn
|
|
when a network stops, and a forgotten network stays forgotten across a restart.
|
|
|
|
Unit tests in `crates/tsunagi/src/proto/handshake.rs` cover the transcript construction
|
|
itself: role separation, channel binding, identity and network binding,
|
|
unambiguous encoding, and rejection under the wrong key.
|
|
|
|
Unit tests in `crates/tsunagi-wg-quic/src/` cover key clamping against the RFC
|
|
7748 vector, overlay derivation, announcement validation including the
|
|
address-hijack attempt, interface naming, and IP header parsing against
|
|
truncated and nonsense input.
|
|
|
|
`tests/end_to_end.rs` is the vertical slice: persistent identity → network
|
|
space → discovery → iroh → authentication → message exchange.
|
|
|
|
What the default suite does **not** cover is the real TUN interface, because
|
|
that needs `CAP_NET_ADMIN`. Everything above it does run.
|
|
|
|
## Not covered, and not claimed to be
|
|
|
|
Listed in [sync-model.md](sync-model.md#future-tests): snapshots, revocations,
|
|
long partitions, hostname renames, recovery of a returning participant, NAT
|
|
traversal, relay fallback, and multi-process or multi-host deployment. None of
|
|
these are implemented, and none are marked as passing.
|
|
|
|
## Debugging a test
|
|
|
|
```bash
|
|
TSUNAGI_TEST_LOG=tsunagi=debug cargo test --test multi_peer -- --nocapture
|
|
```
|
|
|
|
The library never installs a global subscriber; the harness opts in only when
|
|
that variable is set.
|