dsh-plugin-fix-workspace-multiple-writers #
Host-only dsh plugin. While two dsh instances share one $DSH_HOME, each keeps
the entire workspace store in memory — the JSON backend reads the document once,
the domain loads its tables once, and the registry holds its own record
snapshots. A whole-file publish from one instance therefore silently drops
additions the other made, and no process ever notices.
This plugin watches $DSH_HOME/storages/workspace.json and replays the
additions it describes through ctx.workspaceRegistry's public API. That is
the same write path a normal session create uses, so each attach emits
domain/changed and every connected browser shows the session without a reload.
What it does #
- creates a locally-unknown workspace the store describes,
- attaches a session id the store lists that this process does not account for,
- archives a session id the store archived.
What it deliberately does not do #
Never removes, never reorders, never renames, never deletes. updatedAt
cannot arbitrate a conflict: if A writes {S_A} and B writes {S_B} from stale
memory, the file holds {S_B} and S_A is simply absent. A "newer file wins"
rule makes A detach S_A — reintroducing exactly the loss this plugin fixes.
Nothing in a record says "the writer knew about S_A", so a deliberate removal
and a stale-write loss are indistinguishable, and additive-only is the lossless
choice.
Consequences worth knowing:
-
A
detachSessionperformed by another instance is undone by this union. No GUI flow detaches; only the webhook path calls it in this build. -
Session order inside a workspace is not reconciled (
attachSessionprepends), so two instances can display different orders. Cosmetic. -
Workspace order and titles are not reconciled. Cosmetic.
-
Archive reconciliation is append-only, which matches the product: there is no unarchive.
-
The plugin never asks for a removal, but a write it induces can still prune. The registry re-derives a record's whole session account from its local canonical-cwd index on every write (
WorkspaceEntity.mutate), and writes whenever that filtering changes the array. So attaching one valid session to a workspace whose record also holds an id that no longer validates — a moved or deleted session directory, which the registry itself warns about at startup — drops that already-filtered id from the durable record. The plugin has no API that could request this, and any other writer would trigger the same prune, but the trigger is often this plugin's attach. Pinned bytest/integration.test.js("attaching a valid session prunes a locally-unvalidatable id"). -
If the durable accounts cannot be read at all (the storage domain is unreachable), the planner cannot tell "already accounted" from "accounted but locally unvalidatable", so it declines to attach to any workspace it already knows. Only newly created workspaces are attached to, since a fresh record is empty and cannot be pruned.
-
Opening a workspace can land on a session another instance owns. Clicking a workspace — including its "New Session" — routes through
dsh-client-ui-workspace'sconnectWorkspace, which reuses any session withsummary.blank, a matchingsummary.cwd, membership inworkspace.sessionIds, and no archive entry; it callssessions.create({ workspaceId })only when that scan finds nothing. Theblankit consults is the client's local mirror, not the session's state: that mirror is documented as "unknown bare sessions begin conservatively blank" and is cleared only by evidence this client has engaged the session. So a session another instance created is a reuse candidate here — this plugin is what puts its id inworkspace.sessionIds— even though it is not blank on disk (its projection recordssessionListMetadata.blank: falseand its log holds aturn/start). The reuse then cannot be resumed at all: a session log has exactly one live writer (an exclusiveflockonsession.lock), so the first prompt fails withSessionAlreadyOwnedError.The predicate asks whether the workspace accounts for the session, never whether this process can open it — safe while a workspace could only account for sessions this process had attached, and no longer true once a peer's sessions are visible. It is reachable without this plugin, since the workspace account is loaded from the shared store at boot; the plugin makes it immediate rather than boot-order dependent.
Workaround: archive the peer's session, because the scan skips archived ids. Making the session non-blank does not help — the observed one was already non-blank, which is what identifies the client mirror rather than the session state as the lever. The fix belongs in
connectWorkspace: require that the session is actually openable here, or fall back tosessions.createwhen the resume is refused.
It is not atomic, and that cuts both ways. Every registry write republishes
the whole store from this process's memory, which was loaded once at open. So
between this plugin reading the store and committing an attach, a peer can
publish an addition that this process never saw; the plugin's write then reverts
the file to its own view and the peer's addition is gone from disk. The peer does
not repair it either: its plan is store-minus-local, and its own addition is
still in its memory, so it sees no divergence. The addition survives in the
peer's memory (and the file regains it whenever the peer next writes), but if
the peer exits first it is lost, because initialized: true means a restart
never re-bootstraps from the session logs.
The window is small — the read-to-commit span of one pass — but it is real, it is the same class as the bug this plugin addresses, and it cannot be closed from here. Closing it needs cross-process serialization of the read-modify-write, which is the separate "locked KV backend" route the handoff describes, not something a watcher can do.
Behaviour #
- The store path is derived, not assumed: the open domain's own unit is asked
for its file (
$DSH_HOME/storages/workspace.jsonis only the composed default, and a relocated backend root leaves a stale file behind at the old path). A directory-backed (per-record) unit has no single store document, so the plugin declines to run rather than read a directory as a file. - The store's containing directory is watched, not the file: the backend
commits by temp-write +
rename(), which replaces the inode. A watch that cannot be created (the directory does not exist yet) degrades tofs.watchFile, which polls the path itself. - Events are debounced and passes run on a single-flight chain, so a pass never interleaves with itself.
ENOENTmeans "nothing to add". An unparseable or foreign document is logged and left byte-identical — it is never "repaired" by a write.- A
pendingMutationmarker means another instance is mid create/delete; the pass is deferred. The retry is bounded (8 attempts), after which the plugin stops polling and waits for the store to change — a writer that died between its two writes would otherwise leave an indefinite poll. A pass that throws is retried on the same budget. - The pass is idempotent. Once both sides agree it plans nothing, so a converged
system performs zero durable writes and emits no
domain/changed. - The watcher is bound once, at mount, to the store's directory. Changing
$DSH_HOMEafter boot makes later passes read the new file while the watcher still listens to the old directory, so that (exotic) case needs a restart.
Configuration #
- id: dsh-plugin-fix-workspace-multiple-writers
name: ./dsh-plugin-fix-workspace-multiple-writers/lib/index.js
config:
debounceMs: 100 # event coalescing window
retryOnPendingMs: 250 # retry delay while a pendingMutation marker is up
pollIntervalMs: 1000 # fs.watchFile interval, used only by the fallback
Delete the row to remove the capability; dsh boots and behaves as before.
Row changes take effect on the next dsh start. HMR watches the modules under
the checkout's watch root — editing lib/index.js reloads a mounted plugin
live — but the loader has no watcher for a nested cordis:include YAML, so the
row list itself is read once at boot. (Only the profile's own patch layer is
HMR-registered.)
The include resolves to the checkout the launcher ran from (dsh-main
derives it from git rev-parse --show-toplevel), so this directory must exist in
that checkout's dsh-plugins/. A plugin authored in a different working copy
or clone of the repository is invisible to a dsh started elsewhere: no row is
ever created, so it appears nowhere in Settings → Plugins → Plugin list.
Verifying a running instance #
1. Did the row load? A mounted plugin logs one line at startup:
fix-workspace-multiple-writers: watching /Users/you/.dsh/storages/workspace.json (domain unit)
The parenthetical is where the path came from — domain unit, or assumed ($DSH_HOME/storages) when the storage domain could not be asked. It is silent
while converged, so that line is the only proof of life. If it is missing after a
restart, the include is not mounted at all — start dsh through
experimental/dsh-everything-edition's dsh-main / dsh-main-any-port, which
is what injects the dsh-everything-checkout-plugins include; plain dsh web
mounts none of the checkout plugins.
2. Does it reconcile? Create a divergence and watch it heal. Any stored
session whose cwd is the workspace path but which the record does not account
for will do; inject it the way a stale whole-file write would, then read back:
cd /path/to/the/workspace
SID=session-00000000-0000-7000-8000-000000000000 # a real, stored session id
node -e '
const fs=require("node:fs"),os=require("node:os"),p=require("node:path"),c=require("node:crypto");
const file=p.join(process.env.DSH_HOME||p.join(os.homedir(),".dsh"),"storages","workspace.json");
const target=fs.realpathSync(process.argv[2]||process.cwd());
const doc=JSON.parse(fs.readFileSync(file,"utf8"));
const ws=Object.values(doc.tables.workspaces).find((r)=>r.path===target);
console.log("before:",ws.updatedAt,ws.sessionIds.length);
ws.sessionIds=[...ws.sessionIds.filter((id)=>id!==process.argv[1]),process.argv[1]];
const tmp=file+"."+c.randomUUID()+".tmp";
fs.writeFileSync(tmp,JSON.stringify(doc,null,2)+"\n");fs.renameSync(tmp,file);
' "$SID"
Within ~1s the store's updatedAt for that workspace advances and $SID moves
to sessionIds[0] (attachSession prepends) — the plugin attached it through
the registry, and the log prints reconciled 0 workspace(s), 1 session(s), 0 archive(s). Without the plugin updatedAt never changes: nothing else reads
that file. An id that cannot be validated (unknown session, or a cwd that no
longer matches) is warned about and skipped instead, which is equally conclusive.
3. The real acceptance test is two live instances: run a second one with
nix run ./experimental/dsh-everything-edition#main-any-port, create a session
in the shared workspace there, and watch it appear in the first instance's
sidebar — same workspace, no page reload. Comment the row out and repeat to see
the original loss.
Zero runtime dependencies #
This is a hard constraint, not a style preference. A plugin mounted through the
loader's cordis:include resolves its own imports from its own directory, which
never reaches the profile's node_modules. Any package import fails with
ERR_MODULE_NOT_FOUND. The plugin therefore imports only node: built-ins.
Tests #
npm test # unit + fake-registry tests; 3 real-stack tests skip
DSH_PACKAGES_ROOT=<dsh>/lib/node_modules/@deepseek-ai/dsh/node_modules/@deepseek-ai npm test
test/integration.test.js boots the real storage hub, JSON backend, domain
facility and WorkspaceRegistry from an installed dsh and asserts that a store
published by a second instance reaches the registry, emits domain/changed, and
then converges to zero further writes.
test/composition.test.js mounts dsh-plugins/plugins.cordis.yml through the
installed dsh's real Cordis loader — the same cordis:include create that
dsh-app-boot performs — and asserts this row activates and runs a pass, so a
mistyped row can never leave the plugin silently inert.
DSH_PACKAGES_ROOT names the @deepseek-ai directory of a dsh install; the
plugin cannot locate one itself, because it may not import anything but
node: built-ins.