Commit Graph
3 Commits
Author SHA1 Message Date
tsunagiandClaude Opus 5 fae3816bbf Answer DNS over both families, and only for our own interface
Questions now arrive over IPv4 or IPv6, whichever the resolver uses. The
server opens one listener per family — the overlay address or `127.0.0.1`,
and `[::1]` — and all of them are published to systemd-resolved in one
call, which is what that call requires: sending them one at a time leaves
only the last. A family that cannot be bound, IPv6 switched off in the
kernel for instance, no longer stops the other from answering.

The answers stay IPv4, because that is what the overlay is. A listening
address is disposable, unlike an address a member holds in signed state,
so serving one family over the overlay and the other over loopback costs
nothing and loses nothing.

`--no-tun` was also configuring the host. An in-memory interface has a
name and an MTU and nothing else, but everything downstream read that name
as a host interface: the resolver setting was pushed onto whatever else on
the host happened to be called `tsun0` — which, with two agents on one
machine, is another agent's live interface. A factory now says whether
what it creates is on the host, and the resolver setting goes only to an
interface the agent created. For the same reason the complaint that "the
allocated address is not on any interface" no longer fires under
`--no-tun`, where there was never going to be one; it had people looking
for something that had removed their address. The status line says
`tsun0 (in memory, --no-tun)` rather than printing an address beside a name
the operating system does not have.

That distinction also corrected a test that used the in-memory interface
as a stand-in for an unconfigured host interface. They are not the same
case, so the test now uses a factory that claims the host and puts nothing
there — a provisioner that reported a success it did not achieve — and a
second test covers `--no-tun` being an arrangement rather than a fault.

An in-memory device now reports end of stream when it is destroyed. It
never did, so the packet loop reading it could not end, and since shutdown
became bounded that cost every `--no-tun` agent the full five-second grace
before the loop was aborted instead of finishing.

And `status` says when there is no local resolver at all. Its absence is
the answer to "why does this name not resolve?", and leaving the section
out made a report with DNS switched off look exactly like one where it was
running.

The local control protocol is 7: `DnsReport` carries a list of listening
addresses and `OverlayReport` says whether the interface is on the host.

Verified against the running systemd-resolved: it takes
`127.0.0.1:5354 [::1]:5354` on one link in a single call, and forward and
reverse questions for a peer's name are answered identically over both
transports.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 21:21:55 +01:00
tsunagiandClaude Opus 5 0b3915d52b Bound every wait that could last for ever
Nine orphaned test processes were found still running from the day before,
three of them spinning on a core each for twenty hours. The code they ran
is several changes old and the mesh test passes twenty times over now, so
the wedge itself is gone — but nothing in the way it was waited on was
bounded, which is why a wedge lasted a day instead of failing a run.

The harness enforced its deadline only between probes. A probe that never
returned — one call into a wedged runtime, which is exactly what a status
request is — waited for ever inside the deadline it was supposed to obey.
The probe is now bounded too, so the same wedge fails the test in thirty
seconds.

Shutdown claimed to be bounded and was not. The plugins had a grace
period; the network runtimes, the accept loop, the plugin request loop and
the endpoint close did not, and a peer that stops reading is enough to
hold any of them open. Each now gets a grace period and is aborted after
it. The overlay packet loop was not stopped at all: it ends when the
device reports end of stream, which a live interface never does, so it
outlived the interface it was reading. And a plugin's grace period
abandoned the future without stopping the task behind it, so the helper
is public and `wg-quic` uses it on its own runtime.

The local control socket was unbounded in both directions. A wedged agent
left `tsunagi status` hanging with nothing on screen and no way out but
Ctrl-C; it now says the agent did not answer, after five seconds, and
falls back to the state store as it already did for a socket that refuses
a connection. On the serving side, a connection that sends no request no
longer holds a task open.

Tests cover the mechanism — a task that stops on its own is not aborted,
one that ignores the grace is cut off and drops what it held — and both
sides of the change in behaviour: a probe that never answers fails its
deadline, and a silent agent is reported rather than waited out.

Also: the binary opts out of rustdoc, since it shares a name with the
library and `cargo doc` cannot put both in one directory.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 20:41:19 +01:00
tsunagiandClaude Opus 5 142fdf995c Make the protocol a crate of its own
tsunagi-wg-quic. The line between a protocol and the system level is now
drawn by the compiler: nothing in it can reach into tsunagi beyond what
tsunagi makes public, and it carries its own version — which is not the
version peers compare.

Two things the compiler found the moment the boundary was real. The key
store was reaching into the core's `pub(crate)` file-permission helpers;
those are a legitimate service of the system level, because a protocol
keeping keys on disk has the same obligation the agent does, so they are
public now with that said. And the test harness was about to be copied
into a second crate, which is how two copies start to drift; it is a
`testing` feature of the core instead, which is also what anybody writing
a protocol would need.

The bridges put up while things were moving are gone: the error
conversion between the two levels, and the re-exports of the system
level's types from the protocol crate. Imports now say which level they
come from, which is the point.

One deliberate deviation, stated rather than hidden. The authenticated
transport stayed in the core. Moving it would have meant handing a
protocol the network's keys so it could prove membership itself, and a
plugin that can authenticate on the control plane is a worse trade than
a module boundary is worth. So the core proves who is at the other end
and the protocol owns what is said over it — the same separation, without
the secret crossing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 19:55:30 +01:00