entangle #
Find the websites hidden in a GitHub account and move them to Tangled.
Repo content migration is handled separately by an autosync service that mirrors GitHub → Tangled one-way. This tool only deals with the hosting layer: which repos serve a website, from what branch and directory, on what domain, and how much of that can be recreated on Tangled without a human.
Status: on hold. The inventory and planning halves work. The apply half cannot run yet — see Blockers.
node bin/entangle.mjs scan # → data/<owner>.json
node bin/entangle.mjs report # readable inventory
node bin/entangle.mjs plan # → data/<owner>.plan.json
node bin/entangle.mjs migrate # dry run
node bin/entangle.mjs migrate --apply # write site configs on Tangled
GitHub auth comes from the gh CLI (gh auth token). Tangled auth comes from
TANGLED_HANDLE and TANGLED_APP_PASSWORD (an atproto app password, not an
account password). No dependencies; Node 20+.
What the scan looks at #
GitHub reports Pages state in three places that routinely disagree, so the scan collects all of them:
- the repo's
has_pagesflag, which stays true for years after a site is removed GET /repos/{o}/{r}/pages, the authoritative branch/path/CNAME config- live DNS and an actual HTTP request, the only evidence anyone can reach the site
On top of that it reads the repo tree to work out whether servable files exist —
an index.html at the Pages source path, committed build output, or nothing at
all — and looks for deploy config belonging to other hosts (Netlify, Vercel,
Cloudflare Pages, Firebase, Render) so sites that left GitHub Pages still show up.
Why sites get classified the way they do #
Tangled serves a branch and directory verbatim. There is no build step. That single fact drives most of the verdicts:
| Verdict | Meaning |
|---|---|
automatic |
an index.html already sits where Tangled would serve from |
domain-decision |
deployable, but it currently lives on a custom domain |
needs-build |
a generator, an Actions workflow, or GitHub's implicit Jekyll build produces the homepage, and the output never lands in the repo |
elsewhere |
no Pages config, but deploy config for another host |
dead |
nothing servable and nothing answering — it stopped being a site a while ago |
The needs-build case that catches people out is implicit Jekyll: GitHub
renders a bare README.md into the homepage unless a .nojekyll file says
otherwise. Those sites look alive and have no index.html anywhere. On Tangled
they would serve nothing.
Constraints on the Tangled side #
Taken from tangled.org/core and the hosting docs:
- one claimed domain per account —
<handle>.tngl.shfor accounts on Tangled's PDS, or a claimable subdomain underTANGLED_SITES_DOMAIN(defaulttngl.io) - exactly one repo may be the index, served at the domain root; everything
else is served at
<domain>/<repo-name> - no custom domains yet, stated as planned
- no build step, no documented SPA or 404 handling
- no size or file-count limit in
appview/sitesDeploy() - redeploys on every push to the configured branch
The last constraints combine into the main migration hazard: a site that used to
sit at a domain root and links to /style.css breaks when it moves under a
sub-path. entangle plan fetches each index.html and flags this as an
absolute-paths risk rather than leaving it to be discovered after cutover.
The Tangled API this drives #
Site config is not an atproto record — it lives in the appview's database behind
XRPC procedures in the org.tangled.temp.* namespace:
| Method | Purpose |
|---|---|
org.tangled.temp.site.getDomainClaim |
current claimed domain |
org.tangled.temp.site.claimDomain |
claim <subdomain> |
org.tangled.temp.site.releaseDomain |
release it (not permitted for handle-bound tngl.sh domains) |
org.tangled.temp.repo.getSiteConfig |
read a repo's site config |
org.tangled.temp.repo.updateSiteConfig |
set branch, dir, isIndex — and deploy |
org.tangled.temp.repo.disableSite |
remove the site and its deployed files |
temp means unstable by declaration. Expect these to move.
Authentication is standard atproto inter-service auth: resolve handle → DID →
PDS, create a session with an app password, ask the PDS for a service-auth JWT
scoped to the appview's DID (did:web:tangled.org, from APPVIEW_HOST) and the
exact method being called, then send it as a bearer token. src/tangled.mjs does
all of this.
The one identifier that matters is repoDid, the DID a knot mints for a repo. It
is not the account DID and it is not derivable from the repo name — it lives in
the sh.tangled.repo record in the owner's PDS, which is where migrate reads it
from. Repo records created before repo DIDs existed do not have one and cannot
host a site until they do.
Blockers #
Work is paused on two things outside this repo.
The site endpoints are switched off. The appview only mounts /xrpc when
XRPC_ENABLED is set, it defaults to off, and https://tangled.org/xrpc/_health
returns the appview's HTML 404 page. migrate preflights this and reports it
rather than failing obscurely.
The only other way to set site config is the web UI, which issues
PUT /{owner}/{repo}/settings/sites with form fields branch, dir, and
is_index, authenticated by a session cookie from atproto OAuth login. Scripting
that means driving the OAuth flow — a much bigger commitment than an app
password. Until XRPC is reachable, the plan output doubles as a manual
checklist.
The autosync service needs fixes before any GitHub-side change should be committed, because everything written to GitHub propagates to Tangled automatically.
Redirecting old GitHub URLs #
Not built yet; the design is settled and verified against live behaviour.
GitHub Pages cannot issue a 301, so any redirect is client-side. The routing
detail that makes one file sufficient: a request to a path under
<user>.github.io that no Pages-enabled repo claims is served by the user
site's 404.html, with the original path intact. A request under a repo that
does have Pages enabled falls through to GitHub's generic 404 instead.
So the shape is one 404.html on the <user>.github.io repo, plus Pages turned
off on each project repo whose site has moved.
Because sync is one-way GitHub → Tangled, that file lands in the repo Tangled serves and must neutralize itself on arrival:
if (location.hostname.endsWith('github.io')) {
location.replace('https://<sub>.tngl.sh' + location.pathname + location.search + location.hash)
}
Without the hostname guard, a synced 404.html at the Tangled index would bounce
<sub>.tngl.sh/missing back to itself. The same reasoning rules out committing a
shim to a repo's actual Pages branch: it would sync straight to Tangled and
become the site.
Known cost: fallthrough paths return HTTP 404 with a redirecting body, so browsers
follow it but crawlers treat the URL as gone. A <link rel="canonical"> pointing
at the Tangled URL is the best available mitigation.
Layout #
bin/entangle.mjs CLI
src/github.mjs REST client, auth from gh, pagination, rate-limit backoff
src/detect.mjs what a file tree says about how a site is produced
src/scan.mjs per-repo scan: Pages config, tree, workflows, DNS, liveness
src/report.mjs inventory as markdown
src/plan.mjs map sites onto a Tangled domain, flag sub-path breakage
src/tangled.mjs atproto identity, service auth, the site XRPC calls
src/migrate.mjs apply a plan, conservatively
data/ is gitignored: scan output names private repos and describes their
contents.