Commit Graph
23 Commits
Author SHA1 Message Date
ab c724981bfd Added per-network LAN broadcast relay
LAN game discovery previously dropped IPv4 broadcasts at TUN ingress. Carry
limited and subnet-directed UDP broadcasts to authenticated, opted-in members
of the source network, including destinations reached through mesh relays.
Preserve the original IP/UDP bytes and deliver received broadcasts only to the
local TUN; never reflood them or expose another pair's plaintext at transit.

Build immutable recipient snapshots on address and participation changes. The
origin sends one ordinary end-to-end encrypted copy per recipient; the existing
fast, bounded-hop transport router remains unchanged. Validate UDP framing,
source ownership and destination admission without game-specific port rules.
Keep network domains isolated and refuse implicit gateways to physical LANs.
The separate broadcast policy/domain layer is the extension point for future
authorized subnet exports; physical capture, bridging and LAN deduplication
are deliberately not implemented yet.

Persist default-on participation independently for each local network. Add
join --no-broadcast/--broadcast and network broadcast <id> [on|off], including
live updates and authenticated announcements. Joining without a flag preserves
the saved choice. Opt-out stops local origination and delivery, while opaque
unicast transit for other members keeps working.

Migrate SQLite schema 3 to 4 without replacing identities or signed state.
Use control ALPN 3 and local IPC protocol 14 for the new announcement/request
shapes; update peers and restart running agents together. The data ALPN 4
envelope remains unchanged. No release version bump, tag or push is included.

Document agent-owned commits in AGENTS.md: short English subjects, explanatory
bodies, scoped staging, honest validation, and repository-local fallback author
AB <ab@hexor.cy> only when an effective name/email is missing. Release actions
remain the user's responsibility.

Validation on Windows: cargo fmt --all -- --check; cargo check --locked
--workspace --all-targets; cargo clippy --locked --workspace --all-targets --
-D warnings; release workspace/all-target tests: 313 passed. The two existing
SQLite wipe failures (a_wipe_removes_everything_and_the_next_start_is_a_stranger
and wiping_twice_is_as_ordinary_as_wiping_once) were explicitly skipped; the
public-DHT smoke test and forwarding benchmark remain ignored by default.
New coverage exercises real iroh/WireGuard multihop fanout, single delivery,
runtime opt-out, unicast replies, domain isolation, malformed input and schema
migration. TUNs are in-memory; actual games and OS adapter selection were not
tested.
2026-09-22 18:33:52 +03:00
ab d1a0eca723 Fixec join CLI 2026-09-22 16:32:39 +03:00
ab 686c74a3bd Enabled DNS by default 2026-09-22 16:08:06 +03:00
ab 9698f21d55 Added DHT peer resolver, fixed MTU 2026-09-22 15:44:33 +03:00
ab 72eaf9a228 Fixed win dns 2026-09-22 13:32:37 +03:00
ab e735151d62 Added windown support 2026-09-22 12:34:19 +03:00
tsunagiandClaude Opus 5 3581feb9b9 Reach a peer through one that can reach both
Two members of a mesh could both reach a third and not each other, and
that pair was simply lost to one another: a packet for a peer with no data
link was counted undeliverable and dropped. Now it goes through a member
that has both.

What travels is not routes. Each agent says only which peers *it* has a
live link with — first-hand, over the control plane, one hop, never a
claim about somebody else's reachability — and everybody computes their
own way through from that. The choice is local and deterministic (the
lowest endpoint id among the peers that have a link to the destination),
so there is nothing to agree, nothing to elect, and two agents may well
route each direction differently. It is soft state: repeated while it
holds, expired when it stops, so a relay that disappears stops being
chosen without anybody revoking anything.

The one in the middle carries bytes it cannot read. A datagram is wrapped
with the peer it is for, and unwrapped on the other side into the link for
the peer it came *from* — which matters, because a packet attributed to
the carrier would be dropped as coming from an address the carrier does
not hold. The tunnel stays end to end, and the relayed datagram goes link
in, link out: it never reaches the middle's interface, so no routing,
forwarding or firewall setting of that host is involved. One hop, so a
loop cannot form without counting anything.

A protocol is handed one link per peer that now outlives the paths under
it. A direct link that dies, a hop that changes, a direct link that comes
back: none of it tears down a tunnel any more, and the size a protocol may
use does not change with the path. Where there was never a direct link at
all, the link exists anyway as long as a hop does, so a peer reachable
only through somebody still gets a tunnel.

The data ALPN is `tsunagi/data/2`: every datagram now carries a tag saying
whether it is direct, for somebody else, or from somebody else. The local
control protocol is 13, for the relay counters — what this device carried
for others is their traffic on its uplink, and that should not be
invisible. `status` says `via <peer>` on a path through somebody.

Fairness between the peers a relay carries for is deliberately not here
yet: the queues are bounded and the counters are what a limit would be
built on.

Tested with fake links for the mechanics, and end to end with three real
agents — two that cannot reach each other directly, a real WireGuard
packet crossing through the middle. The one arrangement a single host
cannot produce by itself is a pair that cannot see each other, so that is
a `testing`-only switch on the agent config and exists in no release
build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 01:32:43 +01:00
tsunagiandClaude Opus 5 3044e592bd Say when there is nobody to contact
Two freshly wiped devices joined the same network, each got an address,
and nothing happened. The report said "none known yet; nobody else has
joined", which reads as patience — as if somebody were expected and had
not arrived. What had actually happened is that neither agent had anywhere
to look: this project publishes nothing about who is in which network, so
a first meeting needs one of them to be told the other's endpoint id, and
the hints an agent remembers from earlier sessions had gone with the wipe.

Candidates are now in the report, and having none of them with nobody
connected is said plainly, with both ways out of it and this device's own
id ready to hand to the other side. Having some and no sessions yet is the
ordinary state of a network coming up, and stays ungraded.

The local control protocol is 12.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 00:51:39 +01:00
tsunagiandClaude Opus 5 9b240672e6 Put every network's address on the interface, and split up from join
Two networks on one agent, and only one of them worked. The interface
plan is exhaustive by contract — it is what the interface should carry and
nothing else — but the request that built it held a single address, so the
provisioner was told about the first network and took the second one's
address off, or never put it on. On the host that is a network whose
address the operating system has never heard of: the tunnel is up, the
status says all is well, and nothing routes. The request now carries every
address, which is also what takes one off when a network is left or
stopped.

The other half is the command line. `up --network X --secret Y` and
`network join` were two ways to do the same thing, and the one on `up`
could only be undone by restarting — which is how a network somebody left
came back, and how an invite line told the other side to start their agent
with a network baked into it. So they are one thing now, split the way the
system is: **`up` runs the agent** — the device's one process, serving
whatever it has joined, answering `status`, taking instructions — and
**`join` decides what it belongs to**, at any time, while it runs. `join`
is at the top level because it is what gets typed; `network join` is the
same command for anyone who likes the long form.

Every line that told somebody to type the old form is gone with it: the
invite after making a network, the lock error from a second `up`, the
empty-network hint in `id`, the README walkthrough and the WireGuard
document. The invite now prints the `join` line for the other machine and,
separately, the `up --peer` line for an agent that is not running yet —
two commands, because they really are two, and no amount of wording makes
starting an agent the same thing as joining a network.

Joining says how the network stood before: new, already here, or stopped
and now running again. That last one matters — joining is an instruction
to run it, so it undoes a stop, and a pause that ends without a word is a
pause nobody can rely on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 00:30:03 +01:00
tsunagiandClaude Opus 5 1b2f050c05 Start the agent without deciding anything yet
`up` demanded a network, so there was no way to run the agent in one
terminal and decide what it belongs to in another — which is the shape of
the thing now that networks are joined, stopped and left while it runs.

`--network` is optional. Without it the agent starts with whatever it is
already configured for and waits; the banner counts the networks instead of
naming one, and points at `tsunagi network join`. A secret with no network
is refused rather than ignored, because it says nothing on its own.

The periodic summary, which had one network by construction, now falls back
to the first running one — `tsunagi status` is the whole picture and this
line was never more than a glance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 23:36:36 +01:00
tsunagiandClaude Opus 5 637e2f74e4 Tell being away from giving up
Leaving was the only way out of a network, and it is the irreversible one:
it publishes a release and then removes the configuration, the secret, the
network's signed records, its cached hints and the protocol key it used.
What somebody usually wants before a reboot, a trip or an experiment is
the other thing — stop serving it and keep everything.

`tsunagi network stop <id>` closes that network's sessions, takes its
address off the interface and keeps it from starting again. Nothing is
announced, deliberately: to the others this device is away, which is an
ordinary condition they already handle, and the address and name it holds
stay reserved for it. `tsunagi network start <id>` resumes it where it left
off. Both are remembered, so a restart does what the last instruction said
rather than what the last command line happened to say.

Except when the command line says otherwise: `up --network X` starts X
whatever its stored state, because a command naming a network is an
instruction to run it. The banner now says which of the three happened —
`new`, `already here`, or `was stopped; this command starts it` — since
silently, that is a stop that comes back from the dead with nothing to
explain it.

The listing tells the three states apart too: running with its address,
stopped and kept, or configured and waiting for an agent to start. Each row
says what to type to move it, because "stop" and "leave" are a pair that
has to be easy to tell apart before the irreversible one is typed.

The local control protocol is 11.

Covered end to end against a running agent: stopping leaves it configured
and says so, stopping twice is the state asked for rather than an error,
starting brings it back, and the secret afterwards is the one from before —
so it is the same network and not a lookalike.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 23:32:23 +01:00
tsunagiandClaude Opus 5 415b6a6667 Make a network on the spot, and say which one you just started
Two reports of the same shape: a network was left and came back after a
restart, and `network join` asked for a secret it could have invented.

The first was not a bug in leaving. The secret on the start command line
derives the network id, so a command line carrying the secret of a network
you have just left recreates it on the next start — which is right, it says
to join that network, but nothing on screen said so. `up` now marks the
network `· new` or `· already here`, and warns in full when another
configured network answers to the same name. A name is a label; the id is
the identity, and the secret is what decides which of them this is. Said at
the moment it happens it is obvious; discovered later in a status report it
is a mystery, which is exactly how it went.

The second was an omission: `up` had learned to invent a secret and
`network join` had not, so the quickest possible thing — a network with
somebody for as long as it is needed, then gone — still needed a secret
generated first. Both now resolve a bare name the same way: the one network
of that name this device already has, or a fresh random secret when there
is none. It is printed in full, with the single line the other person can
paste as it stands, endpoint id included, because a secret nobody can read
is a network nobody can join.

The id is printed in full by both answers now. The shortened form belongs in
a report, where it is read; this one gets copied into the next command.

Covered end to end against a running agent: joining with no secret prints a
secret and a pasteable command with a peer in it, and joining a name this
device already has resumes that network instead of making another that
merely looks the same.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 23:19:58 +01:00
tsunagiandClaude Opus 5 0f97d60854 Serve a zone per network, and make a network in one command
Twice now a report has read "dns not serving" and been taken for a broken
resolver. It was accurate both times: the agent had been started without
`--dns`. That is the flag's fault, not the reader's — a resolver that
disappears because one word was not retyped is worse than none, since the
names simply stop working. So the setting belongs to the device now: `--dns`
turns it on and it stays on, `tsunagi dns off` turns it off, and
`tsunagi dns on` turns it on for an agent that is already running, without
restarting it. `tsunagi dns` says what it is doing, or what it will do at
the next start when nothing is running.

One agent has one identity and as many networks as it likes, so one DNS
service serves them all: each network is a zone named after it, and joining
or leaving one changes what resolves with no restart. A question carries a
name and not the network it belongs to, so the suffix decides and nothing is
shared between zones — a member of one network is not a name in another.
`--dns-zone` is gone with that: there is no single zone to name any more, and
a network name may contain dots, so `--network lab.internal` is how you get
`music.lab.internal`.

It listens on loopback only, where it always could have. Binding the overlay
address put the zones in front of the whole mesh, and with several networks
on one agent that would have answered one network's questions about
another's names.

That made a gap plain: a second network on an agent had no addresses at all,
because the configured range belongs to whichever network took it first, so
its members had nothing to allocate from and no names to answer with. A
second network now uses the range **derived from its own network id** —
every member derives the same one from something they all already have, so
it is an agreement rather than a local invention. It is held back for a
moment first, because a network that already exists has a range of its own
and a joiner should adopt it rather than argue; that wait is what keeps
"the first member settles it" true.

And a network needs no ceremony to start. `tsunagi up --network lab` with no
secret resolves the obvious way: the one network of that name this device
already has, or — when there is none — a fresh random secret, printed in
full with the single line to send the others. That is the ad-hoc case, one
person makes a network and passes the command round, and it was previously
two steps with a flag people could not find. The secret is printed only when
the agent invented it, because then there is nowhere else to read it from;
one that was supplied is not echoed. Two networks of one name and no secret
is the one case with no answer, and it says so rather than choosing.

Releasing now also stops this agent claiming again. The periodic check would
otherwise publish a fresh claim in the moment between the goodbye and the
teardown, turning a release into a hello nobody asked for.

Exercised with the real binary: a zone per network as a second one is joined
into a running agent, the resolver switched on and off while it runs, and an
ad-hoc network printing its secret and the line to share.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 23:03:42 +01:00
tsunagiandClaude Opus 5 44a799faee Split the command line along the line the system draws
`id` had grown into the place where everything was shown and changed,
including the secret of every network this device had joined — and it
printed them all in its ordinary overview, which is a poor default for
output that gets pasted into chats and issue trackers. Now that networks
have a command of their own, the boundary is the one the system already
has: `id` is this **device**, `network` is what it **belongs to**. A device
outlives every network it is in and a network outlives any device in it, so
a command that mixed them had to be read twice.

`id` keeps the key, the name and the directories, and says how many
networks there are without naming their secrets. `network secret` prints
one, or all of them, and only when asked. Nothing else ever does.

`network join` is the part that was missing entirely. One state directory
is one identity and one live agent, so a second `tsunagi up` on it is
refused — and until now that refusal was the end of the road: a network
could be left while the agent ran but never added. It goes over the control
socket, takes effect at once, and is idempotent, saying which of "joined"
and "already there" happened. With no agent running it is written to the
configuration and starts with the next `up`, and says so rather than
implying it is live. The secret travels over an owner-only socket to the
agent that stores it anyway, and `Request` has a hand-written `Debug` that
redacts it, because a derived one would put it in any log line that printed
a request.

The lock error from a second `up` now answers the question behind it: add
the network to the running agent with one command, or run a genuinely
separate agent — a second identity, with its own directories, interface and
range — with the other. That is the shape of the thing: one agent per
identity, many networks on it, one interface; a second agent is isolated,
not a second view of the first. AGENTS.md carries that as a boundary now,
since it is the kind of thing a change could quietly break.

The local control protocol is 9.

Exercised against a running agent: a second `up` refused with both routes
named, a network joined into the live agent and answering for status at
once, the same one again reported as already there, a same-name network
with a different secret joined with the warning, and `id` showing three
networks and no secrets.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 22:16:57 +01:00
tsunagiandClaude Opus 5 41604225ba Let a device leave a network, and start over
Joining was one command and leaving was nothing at all: a network went into
`state.sqlite` on the first `up` and stayed there, so a mistyped secret
left a second network beside the working one with no way to remove it but
editing the database by hand.

`tsunagi network` lists what this device belongs to. `tsunagi network leave
<id>` publishes a signed release first — while the agent is running and its
sessions are up — and only then deactivates the network and removes it. The
order is the whole point: signed state has no expiry, so the tombstone is
the only thing that ever frees the address and the name for the others, and
after the network is gone there is nothing left here to sign one with.
Peers pass it on, so a member that was away hears it from them rather than
from an agent that has already left.

With no agent running nothing can sign or send, and the command says so
instead of quietly succeeding: `--offline` drops the network locally and
says plainly that the others keep the old claim. The outcome always
distinguishes "published to nobody" from "not published at all", because
they leave the network in different states.

A network is named by its id, and a unique prefix will do. The name is
refused on purpose: two networks can share one — that is exactly the
situation this command exists for — and picking between them for the user
is how the wrong one gets left.

The author's version counter deliberately survives. Rejoining the same
network with the same key must continue above the release, or every replica
that holds the release would treat the new claim as stale and the returning
member would be invisible for good. The protocol key does not survive:
rejoining is joining, not resuming, and coming back with a key the network
was told to let go claims an identity nobody holds any more. Plugins learn
about it through a new `on_network_forgotten`, which is about what outlives
a session rather than what a deactivation tears down.

A released member also drops out of the roster `status` prints. The
tombstone stays in the record set — a replica that never heard of it would
otherwise reinstate the old claim — but listing an author that gave
everything up as a member made leaving look like a peer that had broken.

`tsunagi wipe` is the other half: it empties both directories, so the
device identity, every network, every signed record and everything a
protocol kept beside them go at once and the next start is a stranger. It
refuses while an agent holds the directory, and refuses a directory with no
`state.sqlite` in it, so a mistyped `--state-dir` cannot take somebody's
documents with it. Without `--yes` it only prints what it would remove and
what membership would be lost. It is not a goodbye and says so: leaving the
networks first is what frees their addresses.

The local control protocol is 8 — the socket carries a `Leave` request now,
since only the running agent can publish the release.

Exercised end to end against real agents: leaving by prefix released the
address to a connected peer, leaving by name was refused, `--offline` was
refused until asked for explicitly, wipe was refused while the agent ran,
and the directory afterwards had no identity in it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 21:54:05 +01:00
tsunagiandClaude Opus 5 fae3816bbf Answer DNS over both families, and only for our own interface
Questions now arrive over IPv4 or IPv6, whichever the resolver uses. The
server opens one listener per family — the overlay address or `127.0.0.1`,
and `[::1]` — and all of them are published to systemd-resolved in one
call, which is what that call requires: sending them one at a time leaves
only the last. A family that cannot be bound, IPv6 switched off in the
kernel for instance, no longer stops the other from answering.

The answers stay IPv4, because that is what the overlay is. A listening
address is disposable, unlike an address a member holds in signed state,
so serving one family over the overlay and the other over loopback costs
nothing and loses nothing.

`--no-tun` was also configuring the host. An in-memory interface has a
name and an MTU and nothing else, but everything downstream read that name
as a host interface: the resolver setting was pushed onto whatever else on
the host happened to be called `tsun0` — which, with two agents on one
machine, is another agent's live interface. A factory now says whether
what it creates is on the host, and the resolver setting goes only to an
interface the agent created. For the same reason the complaint that "the
allocated address is not on any interface" no longer fires under
`--no-tun`, where there was never going to be one; it had people looking
for something that had removed their address. The status line says
`tsun0 (in memory, --no-tun)` rather than printing an address beside a name
the operating system does not have.

That distinction also corrected a test that used the in-memory interface
as a stand-in for an unconfigured host interface. They are not the same
case, so the test now uses a factory that claims the host and puts nothing
there — a provisioner that reported a success it did not achieve — and a
second test covers `--no-tun` being an arrangement rather than a fault.

An in-memory device now reports end of stream when it is destroyed. It
never did, so the packet loop reading it could not end, and since shutdown
became bounded that cost every `--no-tun` agent the full five-second grace
before the loop was aborted instead of finishing.

And `status` says when there is no local resolver at all. Its absence is
the answer to "why does this name not resolve?", and leaving the section
out made a report with DNS switched off look exactly like one where it was
running.

The local control protocol is 7: `DnsReport` carries a list of listening
addresses and `OverlayReport` says whether the interface is on the host.

Verified against the running systemd-resolved: it takes
`127.0.0.1:5354 [::1]:5354` on one link in a single call, and forward and
reverse questions for a peer's name are answered identically over both
transports.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 21:21:55 +01:00
tsunagiandClaude Opus 5 0b3915d52b Bound every wait that could last for ever
Nine orphaned test processes were found still running from the day before,
three of them spinning on a core each for twenty hours. The code they ran
is several changes old and the mesh test passes twenty times over now, so
the wedge itself is gone — but nothing in the way it was waited on was
bounded, which is why a wedge lasted a day instead of failing a run.

The harness enforced its deadline only between probes. A probe that never
returned — one call into a wedged runtime, which is exactly what a status
request is — waited for ever inside the deadline it was supposed to obey.
The probe is now bounded too, so the same wedge fails the test in thirty
seconds.

Shutdown claimed to be bounded and was not. The plugins had a grace
period; the network runtimes, the accept loop, the plugin request loop and
the endpoint close did not, and a peer that stops reading is enough to
hold any of them open. Each now gets a grace period and is aborted after
it. The overlay packet loop was not stopped at all: it ends when the
device reports end of stream, which a live interface never does, so it
outlived the interface it was reading. And a plugin's grace period
abandoned the future without stopping the task behind it, so the helper
is public and `wg-quic` uses it on its own runtime.

The local control socket was unbounded in both directions. A wedged agent
left `tsunagi status` hanging with nothing on screen and no way out but
Ctrl-C; it now says the agent did not answer, after five seconds, and
falls back to the state store as it already did for a socket that refuses
a connection. On the serving side, a connection that sends no request no
longer holds a task open.

Tests cover the mechanism — a task that stops on its own is not aborted,
one that ignores the grace is cut off and drops what it held — and both
sides of the change in behaviour: a probe that never answers fails its
deadline, and a silent agent is reported rather than waited out.

Also: the binary opts out of rustdoc, since it shares a name with the
library and `cargo doc` cannot put both in one directory.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 20:41:19 +01:00
tsunagiandClaude Opus 5 4c84cc9e4b Tell two networks of the same name apart
A status report with two sections both headed "network LAB" reads as one
network that is somehow working and empty at once. It was two networks:
the same name with different secrets, which is two different networks
that share nothing, because a network's identity is its name *and* its
secret.

Three fixes for the one confusion.

The heading now carries the network id, so the sections are plainly
different things. A name is a label the user chose; the id is the
identity.

Joining a name that is already configured with another secret says so, at
the moment it happens, because that is almost always a mistyped secret
and until now it silently produced an empty network sitting beside a
working one. `status` flags it too, for the case where it already
happened.

And the second network's emptiness now says why. It had no address
because the only configured range was already taken by the first — one
agent has one interface, so an address belongs to one network — and
"nobody else has joined" pointed at the wrong thing entirely. It now
names the range it cannot have, the reason, and the flag that gives it
one of its own.

Nothing was wrong with the connectivity: the working network's tunnel was
up and its ping was answering throughout.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 20:24:11 +01:00
tsunagiandClaude Opus 5 142fdf995c Make the protocol a crate of its own
tsunagi-wg-quic. The line between a protocol and the system level is now
drawn by the compiler: nothing in it can reach into tsunagi beyond what
tsunagi makes public, and it carries its own version — which is not the
version peers compare.

Two things the compiler found the moment the boundary was real. The key
store was reaching into the core's `pub(crate)` file-permission helpers;
those are a legitimate service of the system level, because a protocol
keeping keys on disk has the same obligation the agent does, so they are
public now with that said. And the test harness was about to be copied
into a second crate, which is how two copies start to drift; it is a
`testing` feature of the core instead, which is also what anybody writing
a protocol would need.

The bridges put up while things were moving are gone: the error
conversion between the two levels, and the re-exports of the system
level's types from the protocol crate. Imports now say which level they
come from, which is the point.

One deliberate deviation, stated rather than hidden. The authenticated
transport stayed in the core. Moving it would have meant handing a
protocol the network's keys so it could prove membership itself, and a
plugin that can authenticate on the control plane is a worse trade than
a module boundary is worth. So the core proves who is at the other end
and the protocol owns what is said over it — the same separation, without
the secret crossing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 19:55:30 +01:00
tsunagiandClaude Opus 5 ff7e235414 Split the command line by level
`up` now says which level each setting belongs to, and `--help` shows the
two sections. System: how the agent reaches peers, the one interface it
owns, the address range, the resolver. Transport: which protocols carry
packets and what they take.

`--wireguard` is gone. `--protocol` takes a list and defaults to
`wg-quic`, which is what the protocol is now called — WireGuard's
cryptography in QUIC datagrams, so the name says what is on the wire
rather than what the implementation borrows. `--protocol none` runs the
control plane alone.

Protocol settings moved to `-o key=value`, or `-o protocol:key=value`
when several are selected. Each protocol declares its own settings and
their help, so `tsunagi protocols` can list them without the agent
knowing anything about any protocol, and a setting nobody takes is
refused rather than dropped — a dropped setting looks exactly like one
that did not work. What the user asked for is checked before anything
that could fail on its own, so a misspelled protocol is not buried under
a privilege error.

`--wg-prefix` and `--wg-mtu` became `--interface` and `--mtu`: they were
never the protocol's, and the interface they describe belongs to the
agent. `--transport` became `--reach`, because "transport" now means the
protocol level and using the word for iroh's path policy as well would
be a collision of meaning rather than a shortage of words.

The plugin gave up the last things that were not its own: the interface
name it carried in its own state, and the check that this agent's
address is really on an interface. Both are the agent's, and the check
is now the agent's too, still said once per address rather than every
round.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 19:44:21 +01:00
tsunagiandClaude Opus 5 9415866193 Take the interface off the protocol and give it to the agent
One agent, one interface. The plugin no longer creates one, no longer
holds a TUN factory, no longer keeps a routing table and no longer
decides who owns an address. What is left of it is the protocol: a
WireGuard key per network, a tunnel per peer, encryption on the way out
and decryption on the way in.

The packet path is now explicit about where each decision lives. Out: the
agent's interface reads a packet, the routing table says whose
destination it is, and each protocol is asked in turn whether it can
carry it there. In: the protocol decrypts and hands the packet up with
the peer it came from attached, and the agent checks that peer is
entitled to the source address before writing it out. A protocol proves
who; only the system level knows what they may say.

Two things found by making it work.

`carry` sent the packet unencrypted at first. The encryption had lived in
the interface loop that moved to the core, so taking that out quietly
removed it — the receiving end rejected plaintext as a bad WireGuard
datagram and the counters said nothing at all. Encryption belongs with
the protocol and is now there, with packets dropped for having no session
yet counted apart, because a handful while a tunnel comes up is normal
and a number that keeps climbing is not.

An agent could impose a range it could not itself route. With one
interface two networks need different ranges, and "the lowest author's
range wins" would have carried one agent's colliding default to
everybody. The configured range is now reserved when a network is
activated — on the serialised path, so the answer does not depend on
which task ran first — and an agent that cannot have it proposes nothing
and adopts whatever the network settles on.

Leaving a network takes its address off the interface and leaves the
interface; the interface goes when the agent does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 19:20:30 +01:00
tsunagiandClaude Opus 5 745bbaea06 Drop IPv6 from the overlay
The overlay address was derived from a WireGuard key, which makes it the
protocol's address — and the whole point of the interface belonging to
the system level is that every protocol carries traffic for the *same*
addresses. A derived-per-protocol address cannot be that.

So the overlay is IPv4 only: allocated at the system level, signed by the
member that holds it, and the same address whichever protocol happens to
be moving the packets. The derivation, its ULA prefix and its constants
are gone, along with the collision rule that existed only because a
derived IPv4 address has too little room to be unique — an allocated one
is unique by construction.

A real loss came with it and is restored explicitly. The announcement was
bound to its network only as a side effect of checking the derived
address, so removing that check removed the binding. It now carries the
network id and rejects a mismatch. Strictly redundant, because a
capability arrives on a session that already proved membership, and kept
because losing a property silently is the wrong way to lose one.

An unlock falls out: the MTU floor of 1280 existed because Linux tears
IPv6 down below it. Without IPv6 the floor is 576, what every IPv4 host
must be able to reassemble, so a relayed path with small datagrams can be
matched rather than warned about. The default stays 1280.

The test that forged an overlay address now forges a network id, which is
what is left to lie about. One flaky assertion fixed while passing: it
waited for "the interface has some address", which is briefly true of the
leftover it was meant to see replaced.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 18:40:49 +01:00
tsunagiandClaude Opus 5 60e6b263d1 Split the system level and the command line into a workspace
First step of separating the layers. The library and the binary are now
crates/tsunagi and crates/tsunagi-cli, which means the plugin crate to
come can be told apart from the core by the compiler rather than by
discipline.

Falls out of it immediately: the CLI's dependencies stop being features
of the library. clap, anstream and tracing-subscriber were optional
dependencies behind a `cli` feature that every library user had to
remember to turn off; now they belong to the crate that uses them, and
the library defaults to no features at all.

The one test that drives the binary moved beside it — a library cannot
depend on a binary built from a crate that depends on the library — and
was rewritten against the public API instead of the test harness.

AGENTS.md said to prefer one crate. It now says the system level and its
plugins are separate crates, for the reason above, and that everything
else stays one crate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 18:05:42 +01:00