Let a device leave a network, and start over

Joining was one command and leaving was nothing at all: a network went into
`state.sqlite` on the first `up` and stayed there, so a mistyped secret
left a second network beside the working one with no way to remove it but
editing the database by hand.

`tsunagi network` lists what this device belongs to. `tsunagi network leave
<id>` publishes a signed release first — while the agent is running and its
sessions are up — and only then deactivates the network and removes it. The
order is the whole point: signed state has no expiry, so the tombstone is
the only thing that ever frees the address and the name for the others, and
after the network is gone there is nothing left here to sign one with.
Peers pass it on, so a member that was away hears it from them rather than
from an agent that has already left.

With no agent running nothing can sign or send, and the command says so
instead of quietly succeeding: `--offline` drops the network locally and
says plainly that the others keep the old claim. The outcome always
distinguishes "published to nobody" from "not published at all", because
they leave the network in different states.

A network is named by its id, and a unique prefix will do. The name is
refused on purpose: two networks can share one — that is exactly the
situation this command exists for — and picking between them for the user
is how the wrong one gets left.

The author's version counter deliberately survives. Rejoining the same
network with the same key must continue above the release, or every replica
that holds the release would treat the new claim as stale and the returning
member would be invisible for good. The protocol key does not survive:
rejoining is joining, not resuming, and coming back with a key the network
was told to let go claims an identity nobody holds any more. Plugins learn
about it through a new `on_network_forgotten`, which is about what outlives
a session rather than what a deactivation tears down.

A released member also drops out of the roster `status` prints. The
tombstone stays in the record set — a replica that never heard of it would
otherwise reinstate the old claim — but listing an author that gave
everything up as a member made leaving look like a peer that had broken.

`tsunagi wipe` is the other half: it empties both directories, so the
device identity, every network, every signed record and everything a
protocol kept beside them go at once and the next start is a stranger. It
refuses while an agent holds the directory, and refuses a directory with no
`state.sqlite` in it, so a mistyped `--state-dir` cannot take somebody's
documents with it. Without `--yes` it only prints what it would remove and
what membership would be lost. It is not a goodbye and says so: leaving the
networks first is what frees their addresses.

The local control protocol is 8 — the socket carries a `Leave` request now,
since only the running agent can publish the release.

Exercised end to end against real agents: leaving by prefix released the
address to a connected peer, leaving by name was refused, `--offline` was
refused until asked for explicitly, wipe was refused while the agent ran,
and the directory afterwards had no identity in it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
tsunagi
2026-09-21 21:54:05 +01:00
co-authored by Claude Opus 5
parent fae3816bbf
commit 41604225ba
14 changed files with 1105 additions and 4 deletions
@@ -233,3 +233,82 @@ fn the_socket_path_is_derived_and_short_enough() {
control_socket_path(&std::path::PathBuf::from("/somewhere/else"))
);
}
/// A source that can also leave, the way the binary's one does.
#[derive(Debug)]
struct Control(Agent);
impl tsunagi::ipc::unix::ReportSource for Control {
fn report(&self) -> BoxFuture<'_, StatusReport> {
Box::pin(async move { StatusReport::default() })
}
fn leave(&self, network_id: String) -> BoxFuture<'_, Result<tsunagi::ipc::LeftReport, String>> {
Box::pin(async move {
let wanted: tsunagi::NetworkId = network_id.parse().map_err(|_| "not an id")?;
let name = self
.0
.list_networks()
.await
.map_err(|err| err.to_string())?
.into_iter()
.find(|network| network.network_id == wanted)
.map(|network| network.name.as_str().to_string())
.ok_or("not a network this agent is in")?;
let outcome = self
.0
.leave_network(wanted)
.await
.map_err(|err| err.to_string())?;
Ok(tsunagi::ipc::LeftReport {
name,
announced: outcome.announced,
peers_told: outcome.peers_told as u32,
})
})
}
}
#[tokio::test]
async fn a_client_can_leave_a_network_through_the_running_agent() {
// The release can only be published by the agent that is running, and
// only while its sessions are up, so leaving goes over this socket
// rather than being done behind its back in the state store.
let discovery = SharedMemoryDiscovery::new();
let (name, secret) = network("control-leave");
let dir = TempDir::new().unwrap();
let agent = Agent::spawn(config_with(dir.path(), &discovery))
.await
.unwrap();
let other = TempDir::new().unwrap();
let peer = Agent::spawn(config_with(other.path(), &discovery))
.await
.unwrap();
let network_id = agent.join_network(&name, &secret).await.unwrap();
peer.join_network(&name, &secret).await.unwrap();
wait_for_peers(&agent, network_id, 1).await;
let socket_path = dir.path().join("control.sock");
let control = ControlSocket::bind(&socket_path, Arc::new(Control(agent.clone())))
.await
.unwrap();
let report = tsunagi::ipc::unix::leave_network(&socket_path, &network_id.to_string())
.await
.unwrap();
assert_eq!(report.name, name.as_str());
assert!(report.announced);
assert_eq!(report.peers_told, 1);
assert!(agent.list_networks().await.unwrap().is_empty());
// Asking again names the state it is in rather than failing obscurely.
let err = tsunagi::ipc::unix::leave_network(&socket_path, &network_id.to_string())
.await
.unwrap_err();
assert!(err.to_string().contains("not a network"), "{err}");
control.shutdown().await;
agent.shutdown().await;
peer.shutdown().await;
}