This repository has no description
flarebot docs skill-packages.md
26 kB

Skill package contract #

Format 1 is a JSON envelope containing formatVersion: 1 and a files array. Each file has a relative path, content, and an optional encoding (text by default, or base64). This explicit file manifest avoids archive extraction and symlink semantics. Installation must call validateSkillPackage before writing or activating any package; validation itself performs no I/O.

Every package contains a root SKILL.md. Its YAML frontmatter and Markdown instructions use the Cloudflare Agents SDK's parseSkillFrontmatter, with a validated projection into the native SkillManifestEntry contract:

---
name: summarize-report
description: Summarize a report using the supplied reference checklist.
metadata:
  version: "1.0.0"
---

Read references/checklist.md, then summarize the supplied report.

Names are lowercase ASCII slugs of at most 64 characters. The required metadata.version is a semantic version of at most 64 characters. Descriptions must be nonempty and at most 1,024 characters. Instructions must be nonempty and at most 32 KiB of UTF-8; the complete SKILL.md is limited to 64 KiB. Duplicate YAML keys, unsupported fields and invalid metadata fail validation. Optional frontmatter fields are license (200 characters), compatibility (500), and allowed-tools (2,000). The metadata map permits up to 32 string entries with 64-character keys and 1,024-character values; the reserved __proto__ key is rejected before record decoding. allowed-tools is descriptive; it does not grant permission or enable a script runner.

Markdown files may live at the root or under any safe relative folder, preserving standard Skill references such as SKILL-MECHANICS.md and agents/guide.md. Other optional files live under resources/, references/, assets/ or scripts/, apart from the root permissions.json declaration described below. Markdown under scripts/ is inert documentation, not an executable script. Text resources support .md, .txt, .json, .csv, .yaml, .yml, .svg; base64 resources support .png, .jpg, .jpeg, .gif, .webp, .pdf. Scripts support text .js, .mjs, .py, .sh, .bash. Compile TypeScript before packaging. Markdown extensions are case-insensitive; other extensions are case-sensitive. The validator records scripts as inert native resources; execution requires a separately configured runner and capability policy.

Paths have at most 240 characters and ASCII alphanumeric, dot, underscore and hyphen segments, each starting with an alphanumeric character. Absolute paths, traversal, escapes, backslashes, hidden segments, trailing dots and Windows device names are rejected. Paths must be unique even after case folding, and no path can be both a file and a directory. Text rejects NUL and unpaired surrogates; base64 must use canonical padding without whitespace. Each decoded file is at most 256 KiB, with at most 128 files and 2 MiB of decoded content per package. Consumers must also bound incoming request bytes before parsing the envelope.

The installer supplies reserved Skill names and origin separately from package bytes. A reserved name fails with duplicate_skill; replacement needs an explicit update flow. Origins are a bundled release, uploaded filename, or HTTPS URL without embedded credentials. URL provenance strips query and fragment and checks both raw and normalized length (2,048 characters); it is a display record, not fetch authorization. Package-supplied provenance fields are rejected.

Successful validation returns the normalized files, metadata, origin, script inventory and a native SkillManifestEntry, ready for agents/skills fromManifest. It promotes metadata.version into the native entry's version. The SHA-256 fingerprint includes the format version and sorted canonical file representations, so file ordering and origin do not change it, while content or encoding changes do. It establishes content identity, not publisher trust.

Failures use SkillPackageError with a stable code, affected path and corrective message. No partial package is returned. Unknown format versions are rejected; future migrations must explicitly validate the old contract and construct a new version rather than silently reinterpret stored files. Preserve the accepted format version and fingerprint alongside installed provenance.

pnpm test:skill-package exercises this contract inside native workerd and loads accepted entries through the real Agents SDK SkillRegistry, including activation and resource reads without a script runner.

Bundled Skills #

skills/bundled.json ships two scriptless reference packages:

  • summarize-report: produce an evidence-based summary using a report checklist.
  • prepare-decision: compare supplied options using a decision-brief template.

Both contain instructions and a Markdown reference. They are available locally without a fetch or external catalog, pass the same package validator, and record the current Flarebot release as their bundled origin. Settings → Extensions lists their versions and lets the owner enable or disable each Skill. New bundled Skills default to enabled. Native source construction filters disabled entries; Think activation is connected in the dedicated activation task.

PersonalSkills stores only owner preferences in native Agent SQLite. The release continues to own bundled files, so a new release replaces bundled instructions, resources and package versions without overwriting disabled choices. Preferences for temporarily removed names are retained in case a later release restores the Skill. Names are stable release identities and must not be reassigned to unrelated Skills. Preference mutations compare both revision and current package fingerprint, so a stale browser cannot unknowingly apply an old edit to new package content.

pnpm test:bundled-skills verifies the native sources, owner isolation, restart durability and release replacement behaviour. Bundled packages remain scriptless; changing that trust boundary requires a separate permission and execution design.

Installed package storage #

Installed packages use the customer's existing ATTACHMENTS R2 binding in a separate skills/v1/<durable-object-id>/ namespace. The vendor control plane does not receive package content. Each immutable object key contains the Skill name, package version, SHA-256 fingerprint and a unique installation ID. The object is the validated format-1 JSON envelope, preserving text and base64 representations. Keeping a package in one object avoids partial file-tree publication and lets the existing native fromManifest source consume a verified entry.

Native Agent SQLite stores only the index: metadata, checksum, origin, enabled state, install time, installation ID and revision. Listing Skills and constructing the native source catalog read this index without fetching R2 objects. Installation validates before writing R2 and commits a new index pointer in one synchronous transaction after the upload succeeds. Revisions increase across deletion and reinstallation; edits based on a superseded revision fail. New and replaced installed packages start disabled, including replacements of an enabled package. The owner explicitly enables the newly reviewed content.

An Effect semaphore serializes uploads, removals and garbage collection. Preference edits remain synchronous and can invalidate a pending upload's compare-and-swap. R2 failure before publication leaves the existing index unchanged. Removal first unpublishes the index, then deletes the unreachable object. Failed cleanup does not misreport an already committed mutation as failed: startup and hourly native Agent maintenance sweep unreferenced objects within this DO's namespace. A crash or lost upload response can leave an orphan until a later sweep; even a write that settles after startup is discovered by subsequent scheduled sweeps. Cleanup retains every currently indexed object and never traverses attachment or another owner's namespace.

There are at most 64 installed Skills. Decoded content retains the format-1 2 MiB limit; serialized objects are capped at 13 MiB to account for JSON escaping. Lazy package reads revalidate the envelope and compare its checksum before native activation or resource projection. Missing or corrupt objects fail closed with a reinstall error. Sources capture the enabled index at construction and reject content after its revision changes, including disable/re-enable. A freshly built source reflects the next admitted turn. A bundled release that reserves an installed name takes precedence and makes the installed entry unavailable without deleting its data.

This storage boundary performs no URL fetching and never executes scripts.

Review and installation #

Extensions → Skills accepts a public GitHub folder URL, a direct public HTTPS JSON package URL, or a local format-1 JSON package. Upload and URL review use authenticated, same-origin HTTP routes under /api/skills/. URL acquisition uses the native Think fetch tool with no redirects, a 15-second download timeout and the same serialized-size bound as uploads. Embedded credentials, private hosts and literal IPs are rejected; application package validation runs after native fetch and before publication.

GitHub folder import accepts https://github.com/owner/repo/tree/ref/path and recursively includes regular .md files under that folder, including its root SKILL.md. Branch names containing slashes must encode those slashes as %2F. The importer resolves the ref once to a commit, reads that commit's recursive folder tree, and downloads each selected file from the same commit. The review records the pinned GitHub URL. Truncated trees, symbolic Markdown links, invalid paths, missing SKILL.md, and packages exceeding the existing file/count/size limits fail before publication. Non-Markdown files are omitted. Markdown links outside the chosen folder are not followed.

A standard Skill without metadata.version receives 0.0.0+git.<commit> during import. The native Agents SDK parses its original frontmatter; the importer adds only the missing version and serializes that metadata with the original instruction body into the reviewed package. A supplied version is preserved and validated. Downloads use native Think fetch with four-way concurrency, fixed GitHub hosts, no redirects or credentials, a bounded tree response, and the installer's existing 25-second operation deadline. Each Markdown download must match its Git blob checksum and byte count. UTF-8 byte-order marks are preserved in supporting files and normalized before root frontmatter parsing; invalid UTF-8 is rejected. Public API rate limits produce a retryable download error. Confirmation installs the reviewed snapshot without another GitHub fetch.

The review shows package metadata, including optional license, compatibility and validated custom metadata, plus provenance, instructions, resource inventory, scripts and the declared allowed-tools hint. Metadata and instructions are displayed as escaped text. Installing a package does not execute anything and leaves it disabled. Declared tool hints are not executable permission grants. A duplicate name shows the installed version and requires an explicit replacement action; its captured revision still has to match at the atomic index commit.

There is one in-memory pending review per owner, valid for five minutes. Its expiry timer prevents hibernation, while a public native Agent keepAlive() lease uses alarm heartbeats to prevent inactivity eviction. Replacement, confirmation and expiry release both. An acquisition that finishes after cancellation or replacement cannot publish its snapshot or release a newer review's lease. Preparing another package replaces the review; exceptional runtime restarts or deployments still require a fresh review. No R2 write or installed index entry is created by review. Confirmation consumes the opaque review ID once and installs exactly the validated bytes held by that review. URL content is not fetched again at confirmation, preventing a changed remote response from replacing the contents the owner saw. A failed or interrupted confirmation may require reloading Skills to determine whether its commit finished; the storage layer still cannot publish a partially enabled package.

Native tests verify an idle SDK heartbeat without browser polling, automatic timer cleanup, one-use confirmation and deliberate restart invalidation. Automatic deployed-edge hibernation is not reproduced by these local tests; the ownership arrangement follows the documented Durable Object lifecycle and the pinned Agents SDK lifetime API.

The UI uses TanStack Query mutations for preparation/confirmation and refetches the full Skills list after installation. Owner-session or dialog changes retire pending requests and their results. Package validation errors carry a bounded field path and correction; expired, replaced and stale-version reviews require reviewing the package again. Removal and the richer installed-detail view follow in the Skill management task.

Manage installed and bundled packages #

Settings lists version, description, enabled state and provenance for every Skill. Inspect opens the validated package metadata and complete instruction body, with its file inventory and declared tool hint. Individual files load only when chosen: text is escaped, scripts remain inert, and binary files download as attachments. Inspection works for disabled packages and does not change their enablement.

The authenticated POST /api/skills/manage boundary accepts inspect/file/remove commands with the source, name, fingerprint and revision. It checks the current identity before and after loading R2; changes during a read invalidate its result. File lookup uses the validated package inventory, never a filesystem or arbitrary R2 key. Bundled packages can be inspected and disabled, but only installed packages can be removed. Removal checks the current revision and unpublishes the index before existing best-effort R2 cleanup.

Replacement opens the same package review flow, requiring the selected name and current replacement revision. The owner explicitly confirms installation disabled. Inspection/file queries retire on connection loss, offline changes or package revision changes. Mutations refetch the full Skill list, without redeployment. Missing or checksum-invalid objects surface a corrective error with replacement and removal still available. Listing remains metadata-only and does not fetch all packages from R2 to guess their health.

pnpm test:skill-management covers native inspection/file bytes, owner isolation, read/write races, malformed and stale identities, restart, missing/corrupt objects, and removal/GC. The packaged Settings suite exercises the Kumo management flow.

Think activation #

Each conversation registers one stable native SkillSource. Think refreshes its metadata catalog before inference, including when the catalog was empty at startup. The catalog adds only names and descriptions to the current owner's instructions. It does not load installed package bytes or script bodies.

Think's native activate_skill tool loads the selected instructions on demand. The parent runtime checks source, name, revision, package fingerprint and current enablement before and after the load. Replacement, removal, disablement or storage failure returns an unavailable activation without breaking the turn. A later turn refreshes the catalog and can recover. Activation uses native fromManifest content projection and native tool results, so multiple Skills remain distinct in the transcript. Tool activity records the Skill name, source identity and fingerprint after parsed arguments arrive, including streamed tool calls.

The conversation explicitly enables activation and resource reading in its tool allowlist. Script execution is separate; neither operation runs scripts or grants their declared tool hints. Previously activated text remains part of conversation history after a Skill is disabled; subsequent activations require current permission.

On-demand resource reads #

Native activation enumerates resource paths without their contents. The agent uses read_skill_resource({name, path}) to fetch an individual reference, asset or script as inert content. The canonical package path validator rejects traversal and unsupported paths before storage access. Unknown valid paths may require a package read to determine that they are absent.

The parent uses native SkillSource.readResource and rechecks current source, revision, fingerprint and enablement after loading. Disabled, removed, replaced or unreadable packages return a fixed unavailable result. Resource bodies are not cached across owner changes.

The public Think turn configuration supplies a bounded reader for the native resource tool name. It returns structured provenance, Skill version, MIME type, encoding, content and a truncation flag. The complete escaped JSON result is at most 64 KiB; truncation preserves Unicode code points and base64 quartets. Resource IDs remain stable for a source/name/path across package versions; the separate fingerprint identifies the exact package. This projection leaves native source bytes unchanged for other SDK consumers.

Activity records the qualified resource name, source, version and package fingerprint. Diagnostics include the tool name and hashed capability/source identifiers and fingerprint, without resource bodies, filenames or origin URLs. Common text formats retain the package contract from above; script reads never execute code and binary reads remain explicitly base64 encoded.

Executable permissions #

An optional root permissions.json declares the maximum host access a package requires. It is strict JSON, text encoded, at most 8 KiB, and covered by the same package fingerprint as instructions and script bytes. For example:

{"network": false, "workspace": "read", "externalCapabilities": []}

Omitted fields default to no network, no workspace and no external capabilities. Workspace values are none, read, or read-write. External capabilities are at most eight distinct exact {source: {kind, id}, id, fingerprint} references from the shared capability catalog. Script execution is required only when the validated inventory contains scripts. Requirements are displayed in install review and Skill settings. They describe requested access; package content and allowed-tools never grant it. Ordinary activation and resource reading retain their independent enablement checks.

Owner settings distinguish execution, network, workspace read, workspace write, and external capabilities, each using Allow / Ask / Never. New installations, replacements and even identical-byte reinstalls start with execution Ask and all other categories Never. Choices survive disable/re-enable, but belong to the immutable installation ID and package fingerprint. Updates compare both current Skill identity and permission revision. Each change records timestamp, revision and before/after choices transactionally. Settings shows the latest 20 changes for that installation; storage retains the latest 256 changes across Skills.

SkillExecutionPermissions resolves a script from the current enabled metadata and applies the shared source policy to each required category. Any required Never blocks the run; otherwise any Ask requires approval. Its capability fingerprint binds the installation, package, Skill revision, permission revision and shared policy revision. The existing durable approval layer must bind that reference to the exact script input when execution is connected.

The server-only invocation grant validates declared access and current authority at every host boundary, suppresses read results if authority changes while I/O is pending, and expires on cancellation or completion. Workspace adapters must also check immediately before committing a write. External calls require their own current Allow policy; approving a script does not approve arbitrary nested MCP arguments. Changes to either the Skill or downstream capability revoke later host operations. These checks are for trusted host adapters; scripts never receive the grant object, raw Agent bindings or arbitrary application tool sets.

The runner uses native Cloudflare isolation. It explicitly sets network: false and supplies only guarded workspace adapters: the pinned native runner otherwise grants implicit read access when given a workspace, and its raw network: true option cannot mediate individual requests. The permission service exposes no browser-callable execution method; the native Think action below owns script admission.

Isolated script execution #

Version 0.2 executes self-contained JavaScript .js and .mjs modules. A script exports a default function accepting (input, ctx). Python and Bash packages can still be installed and inspected, but their execution returns an explicit unsupported-runtime result. Compile multi-file JavaScript before packaging; package resources are accessed through ctx.files, not module imports.

export default async function run(input, ctx) {
  const report = await ctx.workspace.readFile("report.txt");
  const summary = `${input.title}: ${report.length} characters`;
  console.log(summary);
  await ctx.workspace.writeFile("summary.txt", summary);
  return { summary };
}

For this example, permissions.json requires workspace: "read-write". The owner must permit execution, workspace reads and workspace writes. The native Think capability action includes the exact script, current authority fingerprint, JSON input and initial workspace files in its durable approval. Changing the package or permission authority invalidates an old pending action. Only the server-owned one-use invocation can supply its trusted approval grant.

The parent loads and verifies the current immutable package before passing its script to the Cloudflare Agents SDK runner. A public Worker Loader adapter places the user source in a separate module inside a Dynamic Worker. A small trusted runner module imports it and bounds logs/results before crossing RPC; user code cannot access the SDK executor's lexical host objects. No arbitrary script source executes in PersonalAgent. Each load has no application bindings and globalOutbound: null.

The action accepts at most 32 KiB of complete JSON input. Its optional workspace is an explicit list of {path, content} files included in that input. This is a private workspace for one invocation, not a mount of the live conversation or owner filesystem. The adapter allows at most eight files, 16 KiB UTF-8 per file and 64 KiB total. Paths are portable relative paths without traversal; directory listing and a bounded nonrecursive glob can only discover these files. Reads and writes enforce their separate permissions. Changed files are returned as bounded tool data and do not modify the live workspace. Failed or cancelled runs do not publish their workspace output. ctx.output.writeFile is unavailable.

Raw fetch and connect remain blocked. A package declaring network access can use ctx.tools.network_fetch({url}) only after its network policy permits the run. This uses the native Think fetch tool for bounded public HTTPS text reads, with no redirects, custom credentials or ambient browser session. Each declared external capability has a stable tool alias advertised in the action description. Those wrappers validate arguments and recheck both Skill authority and the current downstream capability before calling the existing MCP execution boundary. A script's approval never approves a downstream Ask action implicitly.

Before any native host proxy call, the trusted wrapper clones only JSON own data properties and bounds the complete encoded argument array to 32 KiB. It rejects getters, toJSON, prototype keys and non-JSON values. Since the pinned native codec serializes those arguments again, host calls also verify that its relevant codec intrinsics remain unchanged after importing the user module. Output-only scripts may mutate their realm, but cannot use a changed codec to reach the host.

A parent-owned 20-second deadline includes package loading and revokes all host operations on expiry or cancellation. The native runner also has a 20-second execution timeout. Dynamic Worker configuration caps CPU at 1,000 ms and subrequests at 32; the platform supplies its 128 MB per-isolate memory limit. These custom limits and memory limits are platform enforcement, not locally proven resource exhaustion tests. The pinned local runtime does not expose a physical termination handle. Cancellation returns promptly and makes later host calls fail; it does not claim to physically stop all JavaScript immediately. A script still cannot bypass the isolated runtime's network/binding boundary after cancellation.

The complete native result is at most 48 KiB, reserving space for provenance and workspace output in the final 64 KiB tool result. stdout and stderr each have an 8 KiB escaped-JSON budget; the returned JSON value has a 24 KiB budget. Oversized values produce an explicit output-limit failure rather than malformed partial JSON. Syntax/runtime/import failures use fixed structured errors. Tool activity retains Skill version, source and fingerprint, shows exit status and bounded inert stdout/stderr, and marks unsuccessful script results as failed without crashing the agent turn.