Manage the overlay interface instead of asking for it

The agent printed a list of `ip` commands and asked a human to run them.
That is fragile in the way hand-held setup always is: a persistent TUN
does not survive a reboot, a changed address allocation needs another
manual round, and a run that died leaves a half-configured interface the
next run trips over.

On Linux the agent now creates the interface, sets the MTU, brings it up
and assigns both overlay addresses itself, over netlink in process. No
`ip` is invoked, so nothing this path does can be influenced by PATH, a
shell, or anything a remote peer said.

Cleanup stops being an action. The interface is tied to an open file
descriptor and is deliberately not persistent, so the kernel removes it
when the agent goes — cleanly, by panic, by SIGKILL or by power loss
alike. That also retires `keep_addr_on_down` and `nodad`, which existed
only because an interface nobody held open lost carrier.

Anything still left behind is repaired rather than tripped over: an
abandoned TUN is replaced along with its stale addresses. Two cases
refuse instead of guessing — a link that is not a TUN, because a name
collision is no reason to destroy somebody's bridge, and a TUN another
process holds open, because that is a working overlay belonging to
someone else.

CAP_NET_ADMIN is kept out of the effective set except around the calls
that use it. Two facts shape how: capabilities are per thread, and
netlink checks the credentials of whichever thread calls sendmsg, which
with an async client is the connection task rather than the caller. So
netlink runs on one dedicated thread with a current-thread runtime where
nothing is polled outside a block_on, and opening the TUN descriptor is
synchronous with no await between the guard and its release.

The decision of what to change is a pure function, tested on every
platform; only the execution is behind the provisioner trait. macOS and
Windows get an implementation that refuses with an explanation and falls
back to attaching to a prepared interface, plus a mock host the tests
drive the whole plugin lifecycle against.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
tsunagi
2026-09-21 14:33:19 +01:00
co-authored by Claude Opus 5
parent 944d98389f
commit b23e832a73
15 changed files with 2598 additions and 150 deletions
+9 -1
View File
@@ -17,7 +17,7 @@ cli = ["dep:clap", "dep:tracing-subscriber", "tokio/signal", "tun-device"]
# A real TUN device, so the WireGuard plugin can carry actual IP traffic.
# Needs CAP_NET_ADMIN at run time; without it the plugin still runs and its
# in-memory device can be used for tests.
tun-device = ["dep:tun"]
tun-device = ["dep:tun", "dep:rtnetlink", "dep:caps", "dep:futures-util"]
[[bin]]
name = "tsunagi"
@@ -49,6 +49,14 @@ bytes = "1.12.1"
boringtun = { version = "0.7.1", default-features = false }
tun = { version = "0.8", features = ["async"], optional = true }
# Linux-only interface provisioning. `rtnetlink` configures the interface in
# process, so no `ip` invocation is ever needed; `caps` keeps CAP_NET_ADMIN
# out of the effective set except during the moments it is used.
[target.'cfg(target_os = "linux")'.dependencies]
rtnetlink = { version = "0.23", optional = true }
caps = { version = "0.5", optional = true }
futures-util = { version = "0.3", default-features = false, optional = true }
[dev-dependencies]
tokio = { version = "1.53", features = ["rt", "rt-multi-thread", "sync", "time", "macros", "process"] }
tempfile = "3.24"