Skill package contract #
Format 1 is a JSON envelope containing formatVersion: 1 and a files array.
Each file has a relative path, content, and an optional encoding (text
by default, or base64). This explicit file manifest avoids archive extraction
and symlink semantics. Installation must call validateSkillPackage before
writing or activating any package; validation itself performs no I/O.
Every package contains a root SKILL.md. Its YAML frontmatter and Markdown
instructions use the Cloudflare Agents SDK's parseSkillFrontmatter, with a
validated projection into the native SkillManifestEntry contract:
---
name: summarize-report
description: Summarize a report using the supplied reference checklist.
metadata:
version: "1.0.0"
---
Read references/checklist.md, then summarize the supplied report.
Names are lowercase ASCII slugs of at most 64 characters. The required
metadata.version is a semantic version of at most 64 characters. Descriptions
must be nonempty and at most 1,024 characters. Instructions must be nonempty and
at most 32 KiB of UTF-8; the complete SKILL.md is limited to 64 KiB. Duplicate
YAML keys, unsupported fields and invalid metadata fail validation. Optional
frontmatter fields are license (200 characters), compatibility (500), and
allowed-tools (2,000). The metadata map permits up to 32 string entries with
64-character keys and 1,024-character values; the reserved __proto__ key is
rejected before record decoding. allowed-tools is descriptive;
it does not grant permission or enable a script runner.
Markdown files may live at the root or under any safe relative folder, preserving
standard Skill references such as SKILL-MECHANICS.md and agents/guide.md.
Other optional files live under resources/, references/, assets/ or scripts/,
apart from the root permissions.json declaration described below. Markdown under
scripts/ is inert documentation, not an executable script.
Text resources support .md, .txt, .json, .csv, .yaml, .yml, .svg;
base64 resources support .png, .jpg, .jpeg, .gif, .webp, .pdf.
Scripts support text .js, .mjs, .py, .sh, .bash. Compile TypeScript
before packaging. Markdown extensions are case-insensitive; other extensions are case-sensitive. The validator records scripts
as inert native resources; execution requires a separately configured runner
and capability policy.
Paths have at most 240 characters and ASCII alphanumeric, dot, underscore and hyphen segments, each starting with an alphanumeric character. Absolute paths, traversal, escapes, backslashes, hidden segments, trailing dots and Windows device names are rejected. Paths must be unique even after case folding, and no path can be both a file and a directory. Text rejects NUL and unpaired surrogates; base64 must use canonical padding without whitespace. Each decoded file is at most 256 KiB, with at most 128 files and 2 MiB of decoded content per package. Consumers must also bound incoming request bytes before parsing the envelope.
The installer supplies reserved Skill names and origin separately from package
bytes. A reserved name fails with duplicate_skill; replacement needs an explicit
update flow. Origins are a bundled release, uploaded filename, or HTTPS URL
without embedded credentials. URL provenance strips query and fragment and
checks both raw and normalized length (2,048 characters); it is a display record,
not fetch authorization. Package-supplied provenance fields are rejected.
Successful validation returns the normalized files, metadata, origin, script
inventory and a native SkillManifestEntry, ready for agents/skills
fromManifest. It promotes metadata.version into the native entry's version.
The SHA-256 fingerprint includes the format version and sorted canonical file
representations, so file ordering and origin do not change it, while content or
encoding changes do. It establishes content identity, not publisher trust.
Failures use SkillPackageError with a stable code, affected path and corrective
message. No partial package is returned. Unknown format versions are rejected;
future migrations must explicitly validate the old contract and construct a new
version rather than silently reinterpret stored files. Preserve the accepted
format version and fingerprint alongside installed provenance.
pnpm test:skill-package exercises this contract inside native workerd and loads
accepted entries through the real Agents SDK SkillRegistry, including activation
and resource reads without a script runner.
Bundled Skills #
skills/bundled.json ships two scriptless reference packages:
summarize-report: produce an evidence-based summary using a report checklist.prepare-decision: compare supplied options using a decision-brief template.
Both contain instructions and a Markdown reference. They are available locally without a fetch or external catalog, pass the same package validator, and record the current Flarebot release as their bundled origin. Settings → Extensions lists their versions and lets the owner enable or disable each Skill. New bundled Skills default to enabled. Native source construction filters disabled entries; Think activation is connected in the dedicated activation task.
PersonalSkills stores only owner preferences in native Agent SQLite. The release
continues to own bundled files, so a new release replaces bundled instructions,
resources and package versions without overwriting disabled choices. Preferences
for temporarily removed names are retained in case a later release restores the
Skill. Names are stable release identities and must not be reassigned to unrelated
Skills. Preference mutations compare both revision and current package fingerprint,
so a stale browser cannot unknowingly apply an old edit to new package content.
pnpm test:bundled-skills verifies the native sources, owner isolation, restart
durability and release replacement behaviour. Bundled packages remain scriptless;
changing that trust boundary requires a separate permission and execution design.
Installed package storage #
Installed packages use the customer's existing ATTACHMENTS R2 binding in a
separate skills/v1/<durable-object-id>/ namespace. The vendor control plane does
not receive package content. Each immutable object key contains the Skill name,
package version, SHA-256 fingerprint and a unique installation ID. The object is
the validated format-1 JSON envelope, preserving text and base64 representations.
Keeping a package in one object avoids partial file-tree publication and lets
the existing native fromManifest source consume a verified entry.
Native Agent SQLite stores only the index: metadata, checksum, origin, enabled state, install time, installation ID and revision. Listing Skills and constructing the native source catalog read this index without fetching R2 objects. Installation validates before writing R2 and commits a new index pointer in one synchronous transaction after the upload succeeds. Revisions increase across deletion and reinstallation; edits based on a superseded revision fail. New and replaced installed packages start disabled, including replacements of an enabled package. The owner explicitly enables the newly reviewed content.
An Effect semaphore serializes uploads, removals and garbage collection. Preference edits remain synchronous and can invalidate a pending upload's compare-and-swap. R2 failure before publication leaves the existing index unchanged. Removal first unpublishes the index, then deletes the unreachable object. Failed cleanup does not misreport an already committed mutation as failed: startup and hourly native Agent maintenance sweep unreferenced objects within this DO's namespace. A crash or lost upload response can leave an orphan until a later sweep; even a write that settles after startup is discovered by subsequent scheduled sweeps. Cleanup retains every currently indexed object and never traverses attachment or another owner's namespace.
There are at most 64 installed Skills. Decoded content retains the format-1 2 MiB limit; serialized objects are capped at 13 MiB to account for JSON escaping. Lazy package reads revalidate the envelope and compare its checksum before native activation or resource projection. Missing or corrupt objects fail closed with a reinstall error. Sources capture the enabled index at construction and reject content after its revision changes, including disable/re-enable. A freshly built source reflects the next admitted turn. A bundled release that reserves an installed name takes precedence and makes the installed entry unavailable without deleting its data.
This storage boundary performs no URL fetching and never executes scripts.
Review and installation #
Extensions → Skills accepts a public GitHub folder URL, a direct public HTTPS
JSON package URL, or a local format-1 JSON package. Upload and URL review use authenticated, same-origin HTTP routes
under /api/skills/. URL acquisition uses the native Think fetch tool with no
redirects, a 15-second download timeout and the same serialized-size bound as
uploads. Embedded credentials, private hosts and literal IPs are rejected;
application package validation runs after native fetch and before publication.
GitHub folder import accepts https://github.com/owner/repo/tree/ref/path and
recursively includes regular .md files under that folder, including its root
SKILL.md. Branch names containing slashes must encode those slashes as %2F.
The importer resolves the ref once to a commit, reads that commit's recursive
folder tree, and downloads each selected file from the same commit. The review
records the pinned GitHub URL. Truncated trees, symbolic Markdown links, invalid
paths, missing SKILL.md, and packages exceeding the existing file/count/size
limits fail before publication. Non-Markdown files are omitted. Markdown links
outside the chosen folder are not followed.
A standard Skill without metadata.version receives 0.0.0+git.<commit> during
import. The native Agents SDK parses its original frontmatter; the importer adds
only the missing version and serializes that metadata with the original instruction
body into the reviewed package. A supplied version is preserved and validated.
Downloads use native Think fetch with four-way concurrency, fixed GitHub hosts,
no redirects or credentials, a bounded tree response, and the installer's existing
25-second operation deadline. Each Markdown download must match its Git blob
checksum and byte count. UTF-8 byte-order marks are preserved in supporting files
and normalized before root frontmatter parsing; invalid UTF-8 is rejected.
Public API rate limits produce a retryable download
error. Confirmation installs the reviewed snapshot without another GitHub fetch.
The review shows package metadata, including optional license, compatibility and
validated custom metadata, plus provenance, instructions, resource inventory,
scripts and the declared allowed-tools hint. Metadata and instructions are displayed as
escaped text. Installing a package does not execute anything and leaves it
disabled. Declared tool hints are not executable permission grants. A duplicate
name shows the installed version and requires an explicit replacement action;
its captured revision still has to match at the atomic index commit.
There is one in-memory pending review per owner, valid for five minutes. Its expiry
timer prevents hibernation, while a public native Agent keepAlive() lease uses
alarm heartbeats to prevent inactivity eviction. Replacement, confirmation and
expiry release both. An acquisition that finishes after cancellation or replacement
cannot publish its snapshot or release a newer review's lease. Preparing another
package replaces the review; exceptional runtime restarts or deployments still
require a fresh review. No R2 write or installed index entry is created by review.
Confirmation consumes the
opaque review ID once and installs exactly the validated bytes held by that
review. URL content is not fetched again at confirmation, preventing a changed
remote response from replacing the contents the owner saw. A failed or interrupted
confirmation may require reloading Skills to determine whether its commit finished;
the storage layer still cannot publish a partially enabled package.
Native tests verify an idle SDK heartbeat without browser polling, automatic timer cleanup, one-use confirmation and deliberate restart invalidation. Automatic deployed-edge hibernation is not reproduced by these local tests; the ownership arrangement follows the documented Durable Object lifecycle and the pinned Agents SDK lifetime API.
The UI uses TanStack Query mutations for preparation/confirmation and refetches the full Skills list after installation. Owner-session or dialog changes retire pending requests and their results. Package validation errors carry a bounded field path and correction; expired, replaced and stale-version reviews require reviewing the package again. Removal and the richer installed-detail view follow in the Skill management task.
Manage installed and bundled packages #
Settings lists version, description, enabled state and provenance for every Skill. Inspect opens the validated package metadata and complete instruction body, with its file inventory and declared tool hint. Individual files load only when chosen: text is escaped, scripts remain inert, and binary files download as attachments. Inspection works for disabled packages and does not change their enablement.
The authenticated POST /api/skills/manage boundary accepts inspect/file/remove
commands with the source, name, fingerprint and revision. It checks the current
identity before and after loading R2; changes during a read invalidate its result.
File lookup uses the validated package inventory, never a filesystem or arbitrary
R2 key. Bundled packages can be inspected and disabled, but only installed packages
can be removed. Removal checks the current revision and unpublishes the index
before existing best-effort R2 cleanup.
Replacement opens the same package review flow, requiring the selected name and current replacement revision. The owner explicitly confirms installation disabled. Inspection/file queries retire on connection loss, offline changes or package revision changes. Mutations refetch the full Skill list, without redeployment. Missing or checksum-invalid objects surface a corrective error with replacement and removal still available. Listing remains metadata-only and does not fetch all packages from R2 to guess their health.
pnpm test:skill-management covers native inspection/file bytes, owner isolation,
read/write races, malformed and stale identities, restart, missing/corrupt objects,
and removal/GC. The packaged Settings suite exercises the Kumo management flow.
Think activation #
Each conversation registers one stable native SkillSource. Think refreshes its
metadata catalog before inference, including when the catalog was empty at
startup. The catalog adds only names and descriptions to the current owner's
instructions. It does not load installed package bytes or script bodies.
Think's native activate_skill tool loads the selected instructions on demand.
The parent runtime checks source, name, revision, package fingerprint and current
enablement before and after the load. Replacement, removal, disablement or storage
failure returns an unavailable activation without breaking the turn. A later
turn refreshes the catalog and can recover. Activation uses native fromManifest
content projection and native tool results, so multiple Skills remain distinct
in the transcript. Tool activity records the Skill name, source identity and
fingerprint after parsed arguments arrive, including streamed tool calls.
The conversation explicitly enables activation and resource reading in its tool allowlist. Script execution is separate; neither operation runs scripts or grants their declared tool hints. Previously activated text remains part of conversation history after a Skill is disabled; subsequent activations require current permission.
On-demand resource reads #
Native activation enumerates resource paths without their contents. The agent
uses read_skill_resource({name, path}) to fetch an individual reference, asset
or script as inert content. The canonical package path validator rejects
traversal and unsupported paths before storage access. Unknown valid paths may
require a package read to determine that they are absent.
The parent uses native SkillSource.readResource and rechecks current source,
revision, fingerprint and enablement after loading. Disabled, removed, replaced
or unreadable packages return a fixed unavailable result. Resource bodies are
not cached across owner changes.
The public Think turn configuration supplies a bounded reader for the native resource tool name. It returns structured provenance, Skill version, MIME type, encoding, content and a truncation flag. The complete escaped JSON result is at most 64 KiB; truncation preserves Unicode code points and base64 quartets. Resource IDs remain stable for a source/name/path across package versions; the separate fingerprint identifies the exact package. This projection leaves native source bytes unchanged for other SDK consumers.
Activity records the qualified resource name, source, version and package fingerprint. Diagnostics include the tool name and hashed capability/source identifiers and fingerprint, without resource bodies, filenames or origin URLs. Common text formats retain the package contract from above; script reads never execute code and binary reads remain explicitly base64 encoded.
Executable permissions #
An optional root permissions.json declares the maximum host access a package
requires. It is strict JSON, text encoded, at most 8 KiB, and covered by the same
package fingerprint as instructions and script bytes. For example:
{"network": false, "workspace": "read", "externalCapabilities": []}
Omitted fields default to no network, no workspace and no external capabilities.
Workspace values are none, read, or read-write. External capabilities are
at most eight distinct exact {source: {kind, id}, id, fingerprint} references
from the shared capability catalog. Script execution is required only when the
validated inventory contains scripts. Requirements are displayed in install
review and Skill settings. They describe requested access; package content and
allowed-tools never grant it. Ordinary activation and resource reading retain
their independent enablement checks.
Owner settings distinguish execution, network, workspace read, workspace write, and external capabilities, each using Allow / Ask / Never. New installations, replacements and even identical-byte reinstalls start with execution Ask and all other categories Never. Choices survive disable/re-enable, but belong to the immutable installation ID and package fingerprint. Updates compare both current Skill identity and permission revision. Each change records timestamp, revision and before/after choices transactionally. Settings shows the latest 20 changes for that installation; storage retains the latest 256 changes across Skills.
SkillExecutionPermissions resolves a script from the current enabled metadata
and applies the shared source policy to each required category. Any required
Never blocks the run; otherwise any Ask requires approval. Its capability
fingerprint binds the installation, package, Skill revision, permission revision
and shared policy revision. The existing durable approval layer must bind that
reference to the exact script input when execution is connected.
The server-only invocation grant validates declared access and current authority at every host boundary, suppresses read results if authority changes while I/O is pending, and expires on cancellation or completion. Workspace adapters must also check immediately before committing a write. External calls require their own current Allow policy; approving a script does not approve arbitrary nested MCP arguments. Changes to either the Skill or downstream capability revoke later host operations. These checks are for trusted host adapters; scripts never receive the grant object, raw Agent bindings or arbitrary application tool sets.
The runner uses native Cloudflare isolation. It explicitly
sets network: false and supplies only guarded workspace adapters:
the pinned native runner otherwise grants implicit read access when given a
workspace, and its raw network: true option cannot mediate individual requests.
The permission service exposes no browser-callable execution method; the native
Think action below owns script admission.
Isolated script execution #
Version 0.2 executes self-contained JavaScript .js and .mjs modules. A script
exports a default function accepting (input, ctx). Python and Bash packages can
still be installed and inspected, but their execution returns an explicit
unsupported-runtime result. Compile multi-file JavaScript before packaging;
package resources are accessed through ctx.files, not module imports.
export default async function run(input, ctx) {
const report = await ctx.workspace.readFile("report.txt");
const summary = `${input.title}: ${report.length} characters`;
console.log(summary);
await ctx.workspace.writeFile("summary.txt", summary);
return { summary };
}
For this example, permissions.json requires workspace: "read-write". The
owner must permit execution, workspace reads and workspace writes. The native
Think capability action includes the exact script, current authority fingerprint,
JSON input and initial workspace files in its durable approval. Changing the
package or permission authority invalidates an old pending action. Only the
server-owned one-use invocation can supply its trusted approval grant.
The parent loads and verifies the current immutable package before passing its
script to the Cloudflare Agents SDK runner. A public Worker Loader adapter
places the user source in a separate module inside a Dynamic Worker. A small
trusted runner module imports it and bounds logs/results before crossing RPC;
user code cannot access the SDK executor's lexical host objects. No arbitrary
script source executes in PersonalAgent. Each load has no application bindings
and globalOutbound: null.
The action accepts at most 32 KiB of complete JSON input. Its optional workspace
is an explicit list of {path, content} files included in that input. This is a
private workspace for one invocation, not a mount of the live conversation or
owner filesystem. The adapter allows at most eight files, 16 KiB UTF-8 per file
and 64 KiB total. Paths are portable relative paths without traversal; directory
listing and a bounded nonrecursive glob can only discover these files. Reads
and writes enforce their separate permissions. Changed files are returned as
bounded tool data and do not modify the live workspace. Failed or cancelled runs
do not publish their workspace output. ctx.output.writeFile is unavailable.
Raw fetch and connect remain blocked. A package declaring network access can
use ctx.tools.network_fetch({url}) only after its network policy permits the
run. This uses the native Think fetch tool for bounded public HTTPS text reads,
with no redirects, custom credentials or ambient browser session. Each declared
external capability has a stable tool alias advertised in the action description.
Those wrappers validate arguments and recheck both Skill authority and the
current downstream capability before calling the existing MCP execution boundary.
A script's approval never approves a downstream Ask action implicitly.
Before any native host proxy call, the trusted wrapper clones only JSON own data
properties and bounds the complete encoded argument array to 32 KiB. It rejects
getters, toJSON, prototype keys and non-JSON values. Since the pinned native
codec serializes those arguments again, host calls also verify that its relevant
codec intrinsics remain unchanged after importing the user module. Output-only
scripts may mutate their realm, but cannot use a changed codec to reach the host.
A parent-owned 20-second deadline includes package loading and revokes all host operations on expiry or cancellation. The native runner also has a 20-second execution timeout. Dynamic Worker configuration caps CPU at 1,000 ms and subrequests at 32; the platform supplies its 128 MB per-isolate memory limit. These custom limits and memory limits are platform enforcement, not locally proven resource exhaustion tests. The pinned local runtime does not expose a physical termination handle. Cancellation returns promptly and makes later host calls fail; it does not claim to physically stop all JavaScript immediately. A script still cannot bypass the isolated runtime's network/binding boundary after cancellation.
The complete native result is at most 48 KiB, reserving space for provenance and workspace output in the final 64 KiB tool result. stdout and stderr each have an 8 KiB escaped-JSON budget; the returned JSON value has a 24 KiB budget. Oversized values produce an explicit output-limit failure rather than malformed partial JSON. Syntax/runtime/import failures use fixed structured errors. Tool activity retains Skill version, source and fingerprint, shows exit status and bounded inert stdout/stderr, and marks unsuccessful script results as failed without crashing the agent turn.