9698f21d558daad62ac23417584114e887533c8e
10
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9698f21d55 | Added DHT peer resolver, fixed MTU | ||
|
|
83d3445b7e |
Meet everybody, not just the one you were told about
A device pointed at one member talked to that member and nobody else. The only candidates an agent had were the ones on its command line and whatever the cache remembered, so the network was a star around whichever peer happened to be typed — while the signed state sitting in front of it listed every other member by name. Two sources fix that, and both produce candidates rather than facts. Every author of a signed record is somebody to try. Those records reach us through anybody, so a member is known to exist, and by id, long before it is ever spoken to; an id with no address is still dialable where the endpoint's own discovery can resolve one. And members tell each other where they have seen the others. A `ControlMessage::Peers` carries each member with the addresses the sender observes for it — including the sender's own, which is the one thing nobody else can pass on — in the same `ip:`/`relay:` spelling the cache already uses, so one decoder serves both and neither can drift. An agent that only ever accepts has no candidates of its own and is exactly the one everybody was pointed at, so the addresses come from the live sessions as well as the candidate list. None of it authenticates anything. An introduction is not a vouching: the handshake decides membership as before, and a candidate from a member is tried exactly like one from a bootstrap entry or the cache. It is also deliberately the shape a distributed hash table lookup would return, so that becomes another source beside these rather than a redesign. With the mesh pairwise, the relay is what it was meant to be: the way to the one peer that cannot be reached directly, not the way the network is held together. Covered by three agents where two are told only about the first: each ends up with both of the others, and the one nobody mentioned arrives as an introduction. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0f97d60854 |
Serve a zone per network, and make a network in one command
Twice now a report has read "dns not serving" and been taken for a broken resolver. It was accurate both times: the agent had been started without `--dns`. That is the flag's fault, not the reader's — a resolver that disappears because one word was not retyped is worse than none, since the names simply stop working. So the setting belongs to the device now: `--dns` turns it on and it stays on, `tsunagi dns off` turns it off, and `tsunagi dns on` turns it on for an agent that is already running, without restarting it. `tsunagi dns` says what it is doing, or what it will do at the next start when nothing is running. One agent has one identity and as many networks as it likes, so one DNS service serves them all: each network is a zone named after it, and joining or leaving one changes what resolves with no restart. A question carries a name and not the network it belongs to, so the suffix decides and nothing is shared between zones — a member of one network is not a name in another. `--dns-zone` is gone with that: there is no single zone to name any more, and a network name may contain dots, so `--network lab.internal` is how you get `music.lab.internal`. It listens on loopback only, where it always could have. Binding the overlay address put the zones in front of the whole mesh, and with several networks on one agent that would have answered one network's questions about another's names. That made a gap plain: a second network on an agent had no addresses at all, because the configured range belongs to whichever network took it first, so its members had nothing to allocate from and no names to answer with. A second network now uses the range **derived from its own network id** — every member derives the same one from something they all already have, so it is an agreement rather than a local invention. It is held back for a moment first, because a network that already exists has a range of its own and a joiner should adopt it rather than argue; that wait is what keeps "the first member settles it" true. And a network needs no ceremony to start. `tsunagi up --network lab` with no secret resolves the obvious way: the one network of that name this device already has, or — when there is none — a fresh random secret, printed in full with the single line to send the others. That is the ad-hoc case, one person makes a network and passes the command round, and it was previously two steps with a flag people could not find. The secret is printed only when the agent invented it, because then there is nowhere else to read it from; one that was supplied is not echoed. Two networks of one name and no secret is the one case with no answer, and it says so rather than choosing. Releasing now also stops this agent claiming again. The periodic check would otherwise publish a fresh claim in the moment between the goodbye and the teardown, turning a release into a hello nobody asked for. Exercised with the real binary: a zone per network as a second one is joined into a running agent, the resolver switched on and off while it runs, and an ad-hoc network printing its secret and the line to share. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
41604225ba |
Let a device leave a network, and start over
Joining was one command and leaving was nothing at all: a network went into `state.sqlite` on the first `up` and stayed there, so a mistyped secret left a second network beside the working one with no way to remove it but editing the database by hand. `tsunagi network` lists what this device belongs to. `tsunagi network leave <id>` publishes a signed release first — while the agent is running and its sessions are up — and only then deactivates the network and removes it. The order is the whole point: signed state has no expiry, so the tombstone is the only thing that ever frees the address and the name for the others, and after the network is gone there is nothing left here to sign one with. Peers pass it on, so a member that was away hears it from them rather than from an agent that has already left. With no agent running nothing can sign or send, and the command says so instead of quietly succeeding: `--offline` drops the network locally and says plainly that the others keep the old claim. The outcome always distinguishes "published to nobody" from "not published at all", because they leave the network in different states. A network is named by its id, and a unique prefix will do. The name is refused on purpose: two networks can share one — that is exactly the situation this command exists for — and picking between them for the user is how the wrong one gets left. The author's version counter deliberately survives. Rejoining the same network with the same key must continue above the release, or every replica that holds the release would treat the new claim as stale and the returning member would be invisible for good. The protocol key does not survive: rejoining is joining, not resuming, and coming back with a key the network was told to let go claims an identity nobody holds any more. Plugins learn about it through a new `on_network_forgotten`, which is about what outlives a session rather than what a deactivation tears down. A released member also drops out of the roster `status` prints. The tombstone stays in the record set — a replica that never heard of it would otherwise reinstate the old claim — but listing an author that gave everything up as a member made leaving look like a peer that had broken. `tsunagi wipe` is the other half: it empties both directories, so the device identity, every network, every signed record and everything a protocol kept beside them go at once and the next start is a stranger. It refuses while an agent holds the directory, and refuses a directory with no `state.sqlite` in it, so a mistyped `--state-dir` cannot take somebody's documents with it. Without `--yes` it only prints what it would remove and what membership would be lost. It is not a goodbye and says so: leaving the networks first is what frees their addresses. The local control protocol is 8 — the socket carries a `Leave` request now, since only the running agent can publish the release. Exercised end to end against real agents: leaving by prefix released the address to a connected peer, leaving by name was refused, `--offline` was refused until asked for explicitly, wipe was refused while the agent ran, and the directory afterwards had no identity in it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
142fdf995c |
Make the protocol a crate of its own
tsunagi-wg-quic. The line between a protocol and the system level is now drawn by the compiler: nothing in it can reach into tsunagi beyond what tsunagi makes public, and it carries its own version — which is not the version peers compare. Two things the compiler found the moment the boundary was real. The key store was reaching into the core's `pub(crate)` file-permission helpers; those are a legitimate service of the system level, because a protocol keeping keys on disk has the same obligation the agent does, so they are public now with that said. And the test harness was about to be copied into a second crate, which is how two copies start to drift; it is a `testing` feature of the core instead, which is also what anybody writing a protocol would need. The bridges put up while things were moving are gone: the error conversion between the two levels, and the re-exports of the system level's types from the protocol crate. Imports now say which level they come from, which is the point. One deliberate deviation, stated rather than hidden. The authenticated transport stayed in the core. Moving it would have meant handing a protocol the network's keys so it could prove membership itself, and a plugin that can authenticate on the control plane is a worse trade than a module boundary is worth. So the core proves who is at the other end and the protocol owns what is said over it — the same separation, without the secret crossing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ff7e235414 |
Split the command line by level
`up` now says which level each setting belongs to, and `--help` shows the two sections. System: how the agent reaches peers, the one interface it owns, the address range, the resolver. Transport: which protocols carry packets and what they take. `--wireguard` is gone. `--protocol` takes a list and defaults to `wg-quic`, which is what the protocol is now called — WireGuard's cryptography in QUIC datagrams, so the name says what is on the wire rather than what the implementation borrows. `--protocol none` runs the control plane alone. Protocol settings moved to `-o key=value`, or `-o protocol:key=value` when several are selected. Each protocol declares its own settings and their help, so `tsunagi protocols` can list them without the agent knowing anything about any protocol, and a setting nobody takes is refused rather than dropped — a dropped setting looks exactly like one that did not work. What the user asked for is checked before anything that could fail on its own, so a misspelled protocol is not buried under a privilege error. `--wg-prefix` and `--wg-mtu` became `--interface` and `--mtu`: they were never the protocol's, and the interface they describe belongs to the agent. `--transport` became `--reach`, because "transport" now means the protocol level and using the word for iroh's path policy as well would be a collision of meaning rather than a shortage of words. The plugin gave up the last things that were not its own: the interface name it carried in its own state, and the check that this agent's address is really on an interface. Both are the agent's, and the check is now the agent's too, still said once per address rather than every round. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7c1be332e3 |
Agree a protocol by name and wire version
A peer's announcement already carried a protocol and a version; only the name was being checked. Now both are, and a peer offering a protocol at a version this build does not speak simply has no data plane — the control plane keeps working, messages and signed state still flow, and the difference is reported once rather than on every announcement. The version compared is the *wire* version, not the software version, and the trait says so: a plugin crate has its own version and it is nobody else's business. Two peers on different releases work together for as long as the bytes between them have not changed, and nothing in the negotiation may be derived from anything that moves with a release. A test pins that agreement turns on the name and version alone, with a peer whose announcement carries a payload this build has never seen. `PeerStatus` gained the protocols agreed with each peer, so an empty list is visible as what it is: a session that is up, carrying control traffic, with no protocol in common. The forging test plugin was announcing version 1 while claiming to speak the current one, so its payload was being set aside for the wrong reason. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9415866193 |
Take the interface off the protocol and give it to the agent
One agent, one interface. The plugin no longer creates one, no longer holds a TUN factory, no longer keeps a routing table and no longer decides who owns an address. What is left of it is the protocol: a WireGuard key per network, a tunnel per peer, encryption on the way out and decryption on the way in. The packet path is now explicit about where each decision lives. Out: the agent's interface reads a packet, the routing table says whose destination it is, and each protocol is asked in turn whether it can carry it there. In: the protocol decrypts and hands the packet up with the peer it came from attached, and the agent checks that peer is entitled to the source address before writing it out. A protocol proves who; only the system level knows what they may say. Two things found by making it work. `carry` sent the packet unencrypted at first. The encryption had lived in the interface loop that moved to the core, so taking that out quietly removed it — the receiving end rejected plaintext as a bad WireGuard datagram and the counters said nothing at all. Encryption belongs with the protocol and is now there, with packets dropped for having no session yet counted apart, because a handful while a tunnel comes up is normal and a number that keeps climbing is not. An agent could impose a range it could not itself route. With one interface two networks need different ranges, and "the lowest author's range wins" would have carried one agent's colliding default to everybody. The configured range is now reserved when a network is activated — on the serialised path, so the answer does not depend on which task ran first — and an agent that cannot have it proposes nothing and adopts whatever the network settles on. Leaving a network takes its address off the interface and leaves the interface; the interface goes when the agent does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
745bbaea06 |
Drop IPv6 from the overlay
The overlay address was derived from a WireGuard key, which makes it the protocol's address — and the whole point of the interface belonging to the system level is that every protocol carries traffic for the *same* addresses. A derived-per-protocol address cannot be that. So the overlay is IPv4 only: allocated at the system level, signed by the member that holds it, and the same address whichever protocol happens to be moving the packets. The derivation, its ULA prefix and its constants are gone, along with the collision rule that existed only because a derived IPv4 address has too little room to be unique — an allocated one is unique by construction. A real loss came with it and is restored explicitly. The announcement was bound to its network only as a side effect of checking the derived address, so removing that check removed the binding. It now carries the network id and rejects a mismatch. Strictly redundant, because a capability arrives on a session that already proved membership, and kept because losing a property silently is the wrong way to lose one. An unlock falls out: the MTU floor of 1280 existed because Linux tears IPv6 down below it. Without IPv6 the floor is 576, what every IPv4 host must be able to reassemble, so a relayed path with small datagrams can be matched rather than warned about. The default stays 1280. The test that forged an overlay address now forges a network id, which is what is left to lie about. One flaky assertion fixed while passing: it waited for "the interface has some address", which is briefly true of the leftover it was meant to see replaced. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
60e6b263d1 |
Split the system level and the command line into a workspace
First step of separating the layers. The library and the binary are now crates/tsunagi and crates/tsunagi-cli, which means the plugin crate to come can be told apart from the core by the compiler rather than by discipline. Falls out of it immediately: the CLI's dependencies stop being features of the library. clap, anstream and tracing-subscriber were optional dependencies behind a `cli` feature that every library user had to remember to turn off; now they belong to the crate that uses them, and the library defaults to no features at all. The one test that drives the binary moved beside it — a library cannot depend on a binary built from a crate that depends on the library — and was rewritten against the public API instead of the test harness. AGENTS.md said to prefer one crate. It now says the system level and its plugins are separate crates, for the reason above, and that everything else stays one crate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |