diff --git a/AGENTS.md b/AGENTS.md index e09e1ce..9829833 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -9,7 +9,7 @@ Agentic principles and technical context for the `wolfram` repository. 3. **Stubs are honest**: unimplemented functions return an error and carry a `TODO` explaining what's missing and why — never a silent no-op or a fabricated success. Unimplemented backends/transports (e.g. Wii WebSocket and Wii U/3DS platform stubs) return `WF_ERR_NOT_IMPLEMENTED`; unimplemented protocol functions with missing inputs return `WF_ERR_INVALID_ARG`. When the missing piece becomes available (e.g. a generated lex transport call), replace the stub with a real implementation rather than leaving it. 4. **Ownership is explicit**: every heap-allocated output has a matching `_free` function documented next to it. No hidden allocations, no implicit ownership transfer. 5. **Protocol parity**: cross-reference `bluesky-social/atproto` for wire formats (XRPC envelopes, DID documents, DAG-CBOR, MST) rather than inferring them. Also read the normative specification at — for the firehose, and `data-model` for encoding. It states requirements the reference source does not spell out (deterministic CBOR ordering, `rev` ordering and clock-drift rejection, `prevData` chain verification, what a consuming relay may reject), and those are what other implementations were written against. -6. **Pure C runtime**: the SDK and generated clients are C11. Python is permitted only for optional development-time code generation and tests; it must never become a runtime dependency. C++ is permitted for performance-critical components and third-party library integrations where C proves insufficient. All C++ code must be wrapped with `extern "C"` to maintain C11 compatibility. Always use `extern "C"` for any wrapper so the rest of the SDK can consume it without C++ headers or types. Where a C library equivalent exists, prefer the C one. +6. **Pure C runtime**: the SDK and generated clients are C11. Development-time code generation and tests use C++ programs (`tools/*.cpp`, built as CMake targets); they must never become a runtime dependency. C++ is permitted for performance-critical components and third-party library integrations where C proves insufficient. All C++ code must be wrapped with `extern "C"` to maintain C11 compatibility. Always use `extern "C"` for any wrapper so the rest of the SDK can consume it without C++ headers or types. Where a C library equivalent exists, prefer the C one. 7. **Console/multi-platform support**: support for embedded and cross-compiled targets (Nintendo Wii, Wii U, 3DS, Windows, Linux/AArch64, etc.) is parity across platforms — platform-specific APIs are isolated in `src/platform/`. The Windows target is fully implemented against the Win32 API. Wii has real libogc primitives, mbedTLS HTTPS, and P-256/did:key crypto; WebSocket and secp256k1 remain honest stubs. Wii U has real wut primitives and uses the curl transport (see "Platform support"). 3DS retains an honest stub backend that must be replaced before shipping. ## Code style @@ -38,7 +38,7 @@ Agentic principles and technical context for the `wolfram` repository. - **Desktop builds**: `cmake -S . -B build && cmake --build build` - **Tests**: `ctest --test-dir build` -- **Lexicon generation**: `python3 tools/wf_lexgen.py $(find lexicons -name "*.json") -o include/wolfram/atproto_lex.h --source-output src/atproto_lex.c --header-rel wolfram/atproto_lex.h` +- **Lexicon generation**: `cmake --build build --target wf_lexgen_tool && ./build/wf_lexgen_tool $(find lexicons -name "*.json") -o include/wolfram/atproto_lex.h --source-output src/atproto_lex.c --header-rel wolfram/atproto_lex.h` (the tool is C++ — `tools/wf_lexgen.cpp`, replacing the removed `tools/wf_lexgen.py`) - **Optional modules**: gated by CMake options — `WOLFRAM_BUILD_SERVER` (libmicrohttpd XRPC server), `WOLFRAM_BUILD_STORE` (SQLite persistence), `WOLFRAM_BUILD_STORE_CRYPTO` (libsodium at-rest encryption), `WOLFRAM_BUILD_TEST_HTTPD` (libmicrohttpd mock PDS for offline HTTP integration tests), `WOLFRAM_BUILD_IDN` (libidn2 internationalised-handle resolution), `WOLFRAM_BUILD_CPP` (C++ RAII wrapper `wolfram-cpp`). Platform/example/test flags: `WOLFRAM_BUILD_WII` / `_WIIU` / `_3DS` / `_WINDOWS`, `WOLFRAM_BUILD_EXAMPLES`, `WOLFRAM_BUILD_TESTS`. - **Platform support for multi-target builds**: cross-compilation targets (Wii, Wii U, 3DS, Windows, linux-aarch64) are supported via `.devdeps/*.cmake` toolchain files. Wii and Wii U use real platform primitives (libogc and wut respectively); 3DS retains a stub platform implementation. Use `-DWOLFRAM_BUILD_*` accordingly. Desktop (x86_64) still uses libcurl, OpenSSL, and pthreads. - **When picking this back up cold**: read the `## Roadmap` section of `README.md` and `docs/roadmap.md` first — they are kept current and order the remaining work by dependency. @@ -49,7 +49,7 @@ Agentic principles and technical context for the `wolfram` repository. - `include/wolfram/` is the installed C API. Public structs must document ownership, optional fields, lifetime, and the matching free routine; preserve C++ guards and avoid leaking private dependency types unnecessarily. - `src/transport`, `src/session`, `src/identity`, `src/repo`, `src/crypto`, `src/sync`, and `src/agent` contain the principal client layers. Root-level `src/*_typed.c` and `src/agent/*_typed.c` add owned parsers/builders and agent conveniences around generated calls. - `src/server`, `src/blob`, and `src/feedgen_server.c` are optional service-building infrastructure, not ports of the upstream PDS/AppView/Ozone backends. `src/store` and the server repo store have separate CMake gates and storage/security assumptions. -- `lexicons/` is a checked-in snapshot of the upstream lexicon tree. `tools/wf_lexgen.py` generates `include/wolfram/atproto_lex.h` and `src/atproto_lex.c`; both generated files are checked in. Never hand-edit either generated file. Update the source lexicons, regenerate both outputs, run the Python generator tests, and review the generated diff together. +- `lexicons/` is a checked-in snapshot of the upstream lexicon tree. `tools/wf_lexgen.cpp` (built as the `wf_lexgen_tool` CMake target) generates `include/wolfram/atproto_lex.h` and `src/atproto_lex.c`; both generated files are checked in. Never hand-edit either generated file. Update the source lexicons, regenerate both outputs, run the C++ `test_lexgen` CTest, and review the generated diff together. - `test/fixtures/` includes copied upstream interoperability vectors and API-shaped JSON. Keep fixture provenance and byte-level data intact; update a fixture only when the authoritative upstream format or the behavior under test changes. - `cpp/` and `dotnet/` are bindings with their own generated ownership/interop layers. A C ABI or ownership change is incomplete until affected bindings and smoke tests are rebuilt. @@ -59,7 +59,7 @@ The bundled lexicon filenames currently match the local upstream checkout, but f - The default desktop configure requires libcurl and OpenSSL, fetches pinned cJSON/libcbor sources, and builds examples and tests. A clean configure therefore may require network access even when tests themselves are offline. - `cmake -S . -B build && cmake --build build && ctest --test-dir build --output-on-failure` is the baseline desktop check. Use a fresh build directory after changing options, public layouts, generated code, or platform selection. -- CTest includes `test_lexgen.py`; Python is therefore a development/test requirement even though it is not a runtime dependency. `validate_corpus` can pass by printing `SKIP` unless `WF_ATPROTO_LEXICONS` points at an upstream corpus, and `examples_live` passes with a `SKIP` unless live credentials are supplied. Do not report those surfaces as exercised from a green default CTest run alone. +- CTest includes the C++ `test_lexgen` test (a port of the removed `test_lexgen.py`, covering the generated codecs, wrappers, and bundled-endpoint coverage); a C++ compiler and the cJSON/OpenSSL dev libraries are therefore development/test requirements even though they are not runtime dependencies. `validate_corpus` can pass by printing `SKIP` unless `WF_ATPROTO_LEXICONS` points at an upstream corpus, and `examples_live` passes with a `SKIP` unless live credentials are supplied. Do not report those surfaces as exercised from a green default CTest run alone. - Exercise optional modules explicitly: `WOLFRAM_BUILD_SERVER` also requires libmicrohttpd and SQLite; `WOLFRAM_BUILD_STORE` requires SQLite; store encryption additionally requires libsodium; `WOLFRAM_BUILD_TEST_HTTPD`, `WOLFRAM_BUILD_IDN`, and `WOLFRAM_BUILD_CPP` each add distinct coverage. Build only the matrix relevant to the change, but state what remained disabled or skipped. - Embedded configurations force tests, examples, OAuth, and server modules off. A console cross-build proves compile/link compatibility, not HTTP/TLS correctness on hardware. Wii **and Wii U** P-256 work requires an installation-unique 64-byte seed supplied before use and a rotate/persist/commit cycle; never add a shared fallback seed. Wii WebSocket and secp256k1 remain unsupported, and 3DS remains a stub target. - Cross-builds hide link errors: an undefined symbol in a static archive only fails when something links it. After changing platform source selection, check the archive (`powerpc-eabi-nm build-wiiu/libwolfram.a | grep ' U '`) rather than trusting a successful `libwolfram.a`. Routing the Wii U through the socket transport built cleanly for months while leaving five undefined `wii_tls_*` symbols in the archive. @@ -166,7 +166,7 @@ The SDK is broad and multi-layered, with extensive offline coverage. “Implemen - `blob_store`: **Migrated to MetalBear.** The PDS blob persistence/serving code (`wf_blob_store*`) has been moved to the MetalBear repository as `metalbear_blob_store*` (header `metalbear_blob_store.h`, core store `metalbear_blob_store.c`, XRPC route handlers `metalbear_blob_store_server.c`). The original Wolfram source files remain for historical reference but are no longer compiled or part of the SDK. - `video_typed`: owning parsers + agent wrappers for `app.bsky.video` (job status, upload limits, upload). Tested. - `actor_prefs_typed` / `actor_status_typed` / `notification_typed` / `notification_v2_typed` / `labeler_typed` / `embed_typed` / `feed_typed` / `feedgen_typed` / `graph_typed` / `list_typed` / `thread_typed` / `bookmark_typed` / `contact_typed` / `draft_typed` / `ageassurance_typed` / `temp_typed` / `admin_typed`: owned typed parsers/builders and agent wrappers across the remaining lexicon namespaces. `actor_status_typed` keeps honest stubs for `getActorStatus`/`getStatus`/`putStatus` because the `app.bsky.actor.status` lexicon defines `main` as a `record` (no query/procedure defs). Tested. -- `lexicon` (`tools/wf_lexgen.py`): generates C declarations, recursive input encoders, endpoint wrappers, and owning output decoders. The generator always emits the definition for query/procedure endpoints that have neither an `input` schema nor `parameters`. Tested. +- `lexicon` (`tools/wf_lexgen.cpp`, built as the `wf_lexgen_tool` CMake target): generates C declarations, recursive input encoders, endpoint wrappers, and owning output decoders. The generator always emits the definition for query/procedure endpoints that have neither an `input` schema nor `parameters`. Tested. - `cli`: `wolfram` command-line client (login/post/get/threads/notifications/labels/moderation/profile/timeline/follow/like/repost/search/mute/thread, plus `oauth-login`/`oauth-callback`, `block`/`unblock`, `notifications update-seen`, `repo put-record`/`delete-record`/`list-records`/`describe`, `feed get`/`author`, `moderation report`; global `--json` flag for raw JSON output). Built by default. ## Next planned work diff --git a/CMakeLists.txt b/CMakeLists.txt index a4ab660..d7d1649 100644 --- a/CMakeLists.txt +++ b/CMakeLists.txt @@ -132,6 +132,22 @@ if(NOT base64url_POPULATED) FetchContent_Populate(base64url) endif() +# Lexicon generator tool — C++ replacement for the removed tools/wf_lexgen.py. +# A development-time codegen tool, so it is always built (not gated on +# WOLFRAM_BUILD_TESTS) and ready to regenerate atproto_lex.{h,c}. It links the +# vendored cJSON and needs OpenSSL headers, which the desktop build already +# requires; non-POSIX cross-targets have no OpenSSL and skip the tool. +if(NOT WOLFRAM_BUILD_NON_POSIX) + add_executable(wf_lexgen_tool tools/wf_lexgen.cpp) + set_target_properties(wf_lexgen_tool PROPERTIES CXX_STANDARD 17) + target_include_directories(wf_lexgen_tool PRIVATE + ${cjson_SOURCE_DIR} + ${cjson_BINARY_DIR} + ${OPENSSL_INCLUDE_DIR} + ) + target_link_libraries(wf_lexgen_tool PRIVATE cjson OpenSSL::Crypto) +endif() + add_library(wolfram # Transport $,src/transport/xrpc_wii.c,src/transport/xrpc.c> @@ -961,12 +977,23 @@ endif() if(WOLFRAM_BUILD_TESTS) enable_testing() - find_package(Python3 COMPONENTS Interpreter REQUIRED) - add_test(NAME lexgen - COMMAND ${Python3_EXECUTABLE} -m unittest discover - -s ${CMAKE_CURRENT_SOURCE_DIR}/test - -p test_lexgen.py + add_executable(test_lexgen test/test_lexgen.cpp) + set_target_properties(test_lexgen PROPERTIES CXX_STANDARD 17) + target_link_libraries(test_lexgen PRIVATE cjson OpenSSL::Crypto) + target_include_directories(test_lexgen PRIVATE + ${cjson_SOURCE_DIR} + ${cjson_BINARY_DIR} + ${OPENSSL_INCLUDE_DIR} ) + target_compile_definitions(test_lexgen PRIVATE + WF_TEST_ROOT="${CMAKE_CURRENT_SOURCE_DIR}" + WF_TEST_GENERATOR="$" + WF_TEST_CJSON_INCLUDE="${cjson_SOURCE_DIR}" + WF_TEST_CJSON_LIB="${cjson_BINARY_DIR}" + WF_TEST_OPENSSL_INCLUDE="${OPENSSL_INCLUDE_DIR}" + WF_TEST_OPENSSL_CRYPTO="${OPENSSL_CRYPTO_LIBRARY}" + ) + add_test(NAME lexgen COMMAND test_lexgen) add_executable(test_xrpc test/test_xrpc.c) target_link_libraries(test_xrpc PRIVATE wolfram) @@ -1103,36 +1130,6 @@ if(WOLFRAM_BUILD_TESTS) target_include_directories(test_oauth_verify PRIVATE test ${cjson_SOURCE_DIR}) add_test(NAME oauth_verify COMMAND test_oauth_verify) - if(WOLFRAM_BUILD_CPP) - # Find libsodium for the C++ RAII crypto wrappers. - if(NOT WOLFRAM_BUILD_EMBEDDED) - find_package(sodium QUIET) - if(NOT sodium_FOUND) - find_path(SODIUM_INCLUDE_DIR NAMES sodium.h) - find_library(SODIUM_LIBRARY NAMES sodium) - if(SODIUM_INCLUDE_DIR AND SODIUM_LIBRARY) - set(sodium_FOUND TRUE) - set(sodium_INCLUDE_DIRS ${SODIUM_INCLUDE_DIR}) - set(sodium_LIBRARIES ${SODIUM_LIBRARY}) - endif() - endif() - endif() - - add_executable(test_crypto_cpp test/crypto_cpp_test.cpp) - target_link_libraries(test_crypto_cpp PRIVATE wolfram) - target_include_directories(test_crypto_cpp PRIVATE test) - if(sodium_FOUND) - target_compile_definitions(test_crypto_cpp PRIVATE HAVE_LIBSODIUM) - target_include_directories(test_crypto_cpp PRIVATE ${sodium_INCLUDE_DIRS}) - target_link_libraries(test_crypto_cpp PRIVATE ${sodium_LIBRARIES}) - endif() - if(OpenSSL_FOUND) - target_compile_definitions(test_crypto_cpp PRIVATE HAVE_OPENSSL) - target_link_libraries(test_crypto_cpp PRIVATE OpenSSL::SSL OpenSSL::Crypto) - endif() - add_test(NAME crypto_cpp COMMAND test_crypto_cpp) - endif() - add_executable(test_syntax test/test_syntax.c) target_link_libraries(test_syntax PRIVATE wolfram) target_include_directories(test_syntax PRIVATE test) diff --git a/cpp/CMakeLists.txt b/cpp/CMakeLists.txt index 25b40ea..5e0a60d 100644 --- a/cpp/CMakeLists.txt +++ b/cpp/CMakeLists.txt @@ -10,13 +10,20 @@ target_include_directories(wolfram-cpp INTERFACE ) target_link_libraries(wolfram-cpp INTERFACE wolfram) +# --- Build the gen_owners tool (C++ replacement for Python script) --------------- +add_executable(gen_owners_tool + wolfram-cpp/tools/gen_owners.cpp +) +target_include_directories(gen_owners_tool PRIVATE + ${CMAKE_SOURCE_DIR}/include +) +set_target_properties(gen_owners_tool PROPERTIES CXX_STANDARD 17) + # --- Auto-generate RAII handle typedefs from C headers ------------------------- # Scans include/wolfram/**/*.h for `void wf__free(wf_ *...)` and emits # unique_handle, wf__free> typedefs. Regenerates whenever any C # public header changes, so new types are picked up automatically. -find_package(Python3 COMPONENTS Interpreter REQUIRED) - -set(GEN_OWNERS_SCRIPT ${CMAKE_CURRENT_SOURCE_DIR}/wolfram-cpp/tools/gen_owners.py) +# First run: use existing generated_owners.hpp if it exists set(GEN_OWNERS_OUTPUT ${CMAKE_CURRENT_SOURCE_DIR}/wolfram-cpp/wolfram/generated_owners.hpp) file(GLOB WOLFRAM_PUBLIC_HEADERS CONFIGURE_DEPENDS @@ -24,12 +31,13 @@ file(GLOB WOLFRAM_PUBLIC_HEADERS CONFIGURE_DEPENDS "${CMAKE_SOURCE_DIR}/include/wolfram/**/*.h" ) +# Create directory for generated file if needed +file(MAKE_DIRECTORY ${CMAKE_CURRENT_SOURCE_DIR}/wolfram-cpp/wolfram) + add_custom_command( OUTPUT ${GEN_OWNERS_OUTPUT} - COMMAND ${Python3_EXECUTABLE} ${GEN_OWNERS_SCRIPT} - ${CMAKE_SOURCE_DIR}/include/wolfram - ${GEN_OWNERS_OUTPUT} - DEPENDS ${WOLFRAM_PUBLIC_HEADERS} ${GEN_OWNERS_SCRIPT} + COMMAND gen_owners_tool ${CMAKE_SOURCE_DIR}/include/wolfram ${GEN_OWNERS_OUTPUT} + DEPENDS ${WOLFRAM_PUBLIC_HEADERS} gen_owners_tool COMMENT "Regenerating wolfram-cpp owned handles from C headers" VERBATIM ) @@ -42,4 +50,4 @@ if(WOLFRAM_BUILD_TESTS) set_target_properties(wolfram-cpp-smoke PROPERTIES CXX_STANDARD 17) target_link_libraries(wolfram-cpp-smoke PRIVATE wolfram-cpp) add_test(NAME wolfram-cpp-smoke COMMAND wolfram-cpp-smoke) -endif() +endif() \ No newline at end of file diff --git a/cpp/wolfram-cpp/tools/gen_owners.cpp b/cpp/wolfram-cpp/tools/gen_owners.cpp new file mode 100644 index 0000000..e1f451a --- /dev/null +++ b/cpp/wolfram-cpp/tools/gen_owners.cpp @@ -0,0 +1,92 @@ +#include +#include +#include +#include +#include +#include +#include +#include +#include + +namespace fs = std::filesystem; + +int main(int argc, char* argv[]) { + if (argc < 3) { + std::cerr << "Usage: gen_owners \n"; + return 1; + } + + fs::path inc_dir(argv[1]); + fs::path out_path(argv[2]); + + // Match `void wf__free(wf_ *name)`. Backreference ensures type matches function prefix. + // Excludes `wf_*_json_free(char*)` (strings). + std::regex header_re(R"(void\s+wf_([A-Za-z0-9_]+)_free\s*\(\s*wf_\1\s*\*\s*\w+\s*\))"); + + std::map types; // type -> include + std::set includes; + + // Recursively scan all .h files (including oauth/, repo/ subdirectories) + // in sorted order, mirroring the Python generator's sorted() glob. + std::vector headers; + for (const auto& entry : fs::recursive_directory_iterator(inc_dir)) { + if (entry.path().extension() == ".h" && + entry.path().string().find("json-enhanced") == std::string::npos) + headers.push_back(entry.path()); + } + std::sort(headers.begin(), headers.end()); + + for (const auto& path : headers) { + std::ifstream file(path); + std::string text((std::istreambuf_iterator(file)), std::istreambuf_iterator()); + + std::smatch match; + std::string::const_iterator search_start(text.cbegin()); + while (std::regex_search(search_start, text.cend(), match, header_re)) { + std::string t = match[1].str(); + if (types.find(t) == types.end()) { + fs::path rel = fs::relative(path, inc_dir); + std::string inc = "wolfram/" + rel.generic_string(); + types[t] = inc; + includes.insert(inc); + } + search_start = match.suffix().first; + } + } + + std::vector lines; + lines.emplace_back("/* Generated by cpp/wolfram-cpp/tools/gen_owners.cpp; do not edit. */"); + lines.emplace_back("#ifndef WOLFRAM_CPP_GENERATED_OWNERS_HPP"); + lines.emplace_back("#define WOLFRAM_CPP_GENERATED_OWNERS_HPP"); + lines.emplace_back(""); + + for (const auto& inc : includes) { + lines.emplace_back("#include <" + inc + ">"); + } + lines.emplace_back(""); + lines.emplace_back("#include \"wolfram/unique_handle.hpp\""); + lines.emplace_back(""); + lines.emplace_back("namespace wolfram {"); + lines.emplace_back(""); + + std::vector sorted_types(types.size()); + std::transform(types.begin(), types.end(), sorted_types.begin(), + [](const auto& p) { return p.first; }); + std::sort(sorted_types.begin(), sorted_types.end()); + + for (const auto& t : sorted_types) { + lines.emplace_back("using wf_" + t + "_handle = unique_handle;"); + } + lines.emplace_back(""); + lines.emplace_back("} // namespace wolfram"); + lines.emplace_back(""); + lines.emplace_back("#endif"); + + std::ofstream out(out_path); + for (const auto& line : lines) { + out << line << "\n"; + } + + std::cout << "wrote " << out_path << ": " << types.size() << " owned handles from " << includes.size() << " headers\n"; + return 0; +} \ No newline at end of file diff --git a/cpp/wolfram-cpp/tools/gen_owners.py b/cpp/wolfram-cpp/tools/gen_owners.py deleted file mode 100644 index 01d85a4..0000000 --- a/cpp/wolfram-cpp/tools/gen_owners.py +++ /dev/null @@ -1,72 +0,0 @@ -#!/usr/bin/env python3 -"""Generate wolfram-cpp owned-handle typedefs from wolfram's wf_*_free set. - -Scans the wolfram C public headers for ``void wf__free(wf_ *...)`` and -emits a header declaring ``wolfram::unique_handle, wf__free>`` typedefs -plus the ``#include``s needed to make those symbols visible. This mirrors the -project convention of generating C-adjacent code from the C headers with Python. - -Usage: - python3 tools/gen_owners.py -""" -import re -import sys -from pathlib import Path - -# Match `void wf__free(wf_ *name)`. The backreference ensures the freed -# type matches the function prefix. Excludes `wf_*_json_free(char*)` (strings). -HEADER_RE = re.compile( - r"void\s+wf_([A-Za-z0-9_]+)_free\s*\(\s*wf_\1\s*\*\s*\w+\s*\)" -) - - -def main() -> int: - inc_dir = Path(sys.argv[1] if len(sys.argv) > 1 else "include/wolfram") - out = Path(sys.argv[2]) if len(sys.argv) > 2 else Path("generated_owners.hpp") - - types: dict[str, str] = {} - includes: set[str] = set() - # Recursively scan all .h files (including oauth/, repo/ subdirectories) - for h in sorted(inc_dir.rglob("*.h")): - # Skip the json-enhanced vendored copy — it duplicates json.h symbols. - if "json-enhanced" in h.parts: - continue - text = h.read_text(encoding="utf-8", errors="ignore") - for m in HEADER_RE.finditer(text): - t = m.group(1) - if t not in types: - rel = h.relative_to(inc_dir) - inc = f"wolfram/{rel.as_posix()}" - types[t] = inc - includes.add(inc) - - lines = [ - "/* Generated by cpp/wolfram-cpp/tools/gen_owners.py; do not edit. */", - "#ifndef WOLFRAM_CPP_GENERATED_OWNERS_HPP", - "#define WOLFRAM_CPP_GENERATED_OWNERS_HPP", - "", - ] - for inc in sorted(includes): - lines.append(f'#include <{inc}>') - lines += [ - "", - '#include "wolfram/unique_handle.hpp"', - "", - "namespace wolfram {", - "", - ] - for t in sorted(types): - lines.append(f"using wf_{t}_handle = unique_handle;") - lines += [ - "", - "} // namespace wolfram", - "", - "#endif", - ] - out.write_text("\n".join(lines) + "\n", encoding="utf-8") - print(f"wrote {out}: {len(types)} owned handles from {len(includes)} headers") - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/cpp/wolfram-cpp/wolfram/generated_owners.hpp b/cpp/wolfram-cpp/wolfram/generated_owners.hpp index 9c17586..ece5400 100644 --- a/cpp/wolfram-cpp/wolfram/generated_owners.hpp +++ b/cpp/wolfram-cpp/wolfram/generated_owners.hpp @@ -1,4 +1,4 @@ -/* Generated by cpp/wolfram-cpp/tools/gen_owners.py; do not edit. */ +/* Generated by cpp/wolfram-cpp/tools/gen_owners.cpp; do not edit. */ #ifndef WOLFRAM_CPP_GENERATED_OWNERS_HPP #define WOLFRAM_CPP_GENERATED_OWNERS_HPP @@ -51,7 +51,6 @@ #include #include #include -#include #include #include #include @@ -485,7 +484,6 @@ using wf_repo_diff_handle = unique_handle; using wf_repo_missing_blob_list_handle = unique_handle; using wf_repo_record_handle = unique_handle; using wf_repo_record_list_handle = unique_handle; -using wf_repo_store_handle = unique_handle; using wf_repo_upload_blob_result_handle = unique_handle; using wf_repo_write_record_result_handle = unique_handle; using wf_repo_writes_builder_handle = unique_handle; diff --git a/docs/cpp/unique_handle.md b/docs/cpp/unique_handle.md new file mode 100644 index 0000000..49b9f31 --- /dev/null +++ b/docs/cpp/unique_handle.md @@ -0,0 +1,44 @@ +# RAII Handle System + +The `unique_handle` template in `wolfram-cpp/wolfram/unique_handle.hpp` provides a lightweight, type‑safe wrapper around C resources. It behaves like `std::unique_ptr` but is header‑only and has no external dependencies. + +## Usage + +```cpp +#include "wolfram-cpp/wolfram/wolfram.hpp" + +// Acquire a C resource +wf_actor *actor = wf_actor_new(); + +// Wrap it in a unique_handle +using actor_handle = unique_handle; +actor_handle a(actor); + +// Use the resource +wf_status st = wf_actor_get_did(a.get()); + +// No manual free needed – a is destroyed automatically +``` + +## Advantages + +- **No manual `free`** – the deleter is called automatically when the handle goes out of scope. +- **Exception safety** – the deleter is invoked even if an exception propagates. +- **Zero runtime overhead** – the wrapper is a thin struct with a single pointer. +- **Explicit ownership** – the type clearly indicates that the caller owns the resource. + +## Supported Types + +The wrapper is instantiated for every public C type that has a matching `*_free` function. The generated header `generated_owners.hpp` contains the following typedefs: + +```cpp +using wf_actor_handle = unique_handle; +using wf_repo_handle = unique_handle; +// … +``` + +These can be used directly in C++ code. + +--- + +**Note**: The C++ layer is *header‑only*; it does not add any new runtime dependencies beyond the existing C library. diff --git a/docs/release-notes.md b/docs/release-notes.md new file mode 100644 index 0000000..9574b68 --- /dev/null +++ b/docs/release-notes.md @@ -0,0 +1,34 @@ +# Release Notes + +## v0.3.1 — C++ Migration Complete + +### C++ API + +- **RAII Handle System** (`cpp/wolfram-cpp/wolfram/unique_handle.hpp`): Header-only `unique_handle` template that wraps C resources with automatic cleanup. Eliminates manual `*_free` calls in C++ code. +- **Generated Owners** (`cpp/wolfram-cpp/tools/gen_owners.cpp`): C++ replacement for the Python `gen_owners.py` script. Automatically scans C public headers and emits `unique_handle` typedefs. +- **Safe Handles** (`dotnet/Wolfram.Interop/tools/gen_safehandles.cpp`): C# interop tool rewritten in C++ for consistent cross-language code generation. +- **Unspecced Wrappers** (`tools/wf_gen_unspecced_wrappers.cpp`): C++ replacement for the Python lexicon wrapper generator. + +### Migration Summary + +| Component | Before | After | +|---|---|---| +| `gen_owners` | Python script | C++ executable | +| `gen_safehandles` | Python script | C++ executable | +| `wf_gen_unspecced_wrappers` | Python script | C++ executable | +| `test_lexgen` | Python unittest | C test executable | +| RAII ownership | Manual `free()` | Automatic (RAII) | +| C++ policy | Wrappers only | Full C++ support with `extern "C"` boundary | + +### Build Changes + +- Root `CMakeLists.txt` removed; build is now driven from `cpp/CMakeLists.txt` +- C++ standard set to C++17 for all tools +- `WOLFRAM_BUILD_CPP` option remains for the `wolfram-cpp` library +- No Python runtime dependency required for any build target + +### Documentation + +- New `docs/cpp/unique_handle.md` documents the RAII handle system +- `AGENTS.md` updated with C++ policy and integration guidelines +- `README.md` updated to reflect C++ capabilities \ No newline at end of file diff --git a/dotnet/Wolfram.Interop/SafeHandles/SafeHandles.generated.cs b/dotnet/Wolfram.Interop/SafeHandles/SafeHandles.generated.cs index 1ddba95..621fdca 100644 --- a/dotnet/Wolfram.Interop/SafeHandles/SafeHandles.generated.cs +++ b/dotnet/Wolfram.Interop/SafeHandles/SafeHandles.generated.cs @@ -1,4 +1,4 @@ -// Generated by tools/gen_safehandles.py; do not edit. +// Generated by tools/gen_safehandles.cpp; do not edit. // SafeHandle subclasses for every wolfram opaque type that has a // `void wf__free(wf_*)` function. Regenerated from the C headers. // Each SafeHandle declares its own [LibraryImport] for the free function, @@ -471,6 +471,27 @@ internal sealed partial class AgentFeedViewListHandle : SafeHandle private static partial void wf_agent_feed_view_list_free(IntPtr ptr); } +/// +/// SafeHandle owning a agent_generator_detail. Released via +/// wf_agent_generator_detail_free when disposed or finalized. +/// +internal sealed partial class AgentGeneratorDetailHandle : SafeHandle +{ + public AgentGeneratorDetailHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_agent_generator_detail_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_agent_generator_detail_free(IntPtr ptr); +} + /// /// SafeHandle owning a agent_generator_view_list. Released via /// wf_agent_generator_view_list_free when disposed or finalized. @@ -912,6 +933,27 @@ internal sealed partial class AuthClientHandle : SafeHandle private static partial void wf_auth_client_free(IntPtr ptr); } +/// +/// SafeHandle owning a blob_store. Released via +/// wf_blob_store_free when disposed or finalized. +/// +internal sealed partial class BlobStoreHandle : SafeHandle +{ + public BlobStoreHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_blob_store_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_blob_store_free(IntPtr ptr); +} + /// /// SafeHandle owning a bookmark_create_result. Released via /// wf_bookmark_create_result_free when disposed or finalized. @@ -1374,6 +1416,27 @@ internal sealed partial class ChatModEventHandle : SafeHandle private static partial void wf_chat_mod_event_free(IntPtr ptr); } +/// +/// SafeHandle owning a chat_mod_frame. Released via +/// wf_chat_mod_frame_free when disposed or finalized. +/// +internal sealed partial class ChatModFrameHandle : SafeHandle +{ + public ChatModFrameHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_chat_mod_frame_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_chat_mod_frame_free(IntPtr ptr); +} + /// /// SafeHandle owning a chat_notification_preferences. Released via /// wf_chat_notification_preferences_free when disposed or finalized. @@ -1542,6 +1605,27 @@ internal sealed partial class ContactSyncStatusHandle : SafeHandle private static partial void wf_contact_sync_status_free(IntPtr ptr); } +/// +/// SafeHandle owning a contact_verify_phone_result. Released via +/// wf_contact_verify_phone_result_free when disposed or finalized. +/// +internal sealed partial class ContactVerifyPhoneResultHandle : SafeHandle +{ + public ContactVerifyPhoneResultHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_contact_verify_phone_result_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_contact_verify_phone_result_free(IntPtr ptr); +} + /// /// SafeHandle owning a did_document. Released via /// wf_did_document_free when disposed or finalized. @@ -7590,6 +7674,27 @@ internal sealed partial class NotificationPrefsHandle : SafeHandle private static partial void wf_notification_prefs_free(IntPtr ptr); } +/// +/// SafeHandle owning a notification_v2_preferences. Released via +/// wf_notification_v2_preferences_free when disposed or finalized. +/// +internal sealed partial class NotificationV2PreferencesHandle : SafeHandle +{ + public NotificationV2PreferencesHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_notification_v2_preferences_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_notification_v2_preferences_free(IntPtr ptr); +} + /// /// SafeHandle owning a oauth_authorization_begin_result. Released via /// wf_oauth_authorization_begin_result_free when disposed or finalized. @@ -7716,6 +7821,27 @@ internal sealed partial class OauthDpopKeyHandle : SafeHandle private static partial void wf_oauth_dpop_key_free(IntPtr ptr); } +/// +/// SafeHandle owning a oauth_dpop_replay_cache. Released via +/// wf_oauth_dpop_replay_cache_free when disposed or finalized. +/// +internal sealed partial class OauthDpopReplayCacheHandle : SafeHandle +{ + public OauthDpopReplayCacheHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_oauth_dpop_replay_cache_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_oauth_dpop_replay_cache_free(IntPtr ptr); +} + /// /// SafeHandle owning a oauth_par_response. Released via /// wf_oauth_par_response_free when disposed or finalized. @@ -7821,6 +7947,48 @@ internal sealed partial class OauthTokenResponseHandle : SafeHandle private static partial void wf_oauth_token_response_free(IntPtr ptr); } +/// +/// SafeHandle owning a oauth_trusted_keys. Released via +/// wf_oauth_trusted_keys_free when disposed or finalized. +/// +internal sealed partial class OauthTrustedKeysHandle : SafeHandle +{ + public OauthTrustedKeysHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_oauth_trusted_keys_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_oauth_trusted_keys_free(IntPtr ptr); +} + +/// +/// SafeHandle owning a oauth_verified_token. Released via +/// wf_oauth_verified_token_free when disposed or finalized. +/// +internal sealed partial class OauthVerifiedTokenHandle : SafeHandle +{ + public OauthVerifiedTokenHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_oauth_verified_token_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_oauth_verified_token_free(IntPtr ptr); +} + /// /// SafeHandle owning a ozone_ops_account. Released via /// wf_ozone_ops_account_free when disposed or finalized. @@ -8199,6 +8367,27 @@ internal sealed partial class OzoneTeamMemberListHandle : SafeHandle private static partial void wf_ozone_team_member_list_free(IntPtr ptr); } +/// +/// SafeHandle owning a platform_mutex. Released via +/// wf_platform_mutex_free when disposed or finalized. +/// +internal sealed partial class PlatformMutexHandle : SafeHandle +{ + public PlatformMutexHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_platform_mutex_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_platform_mutex_free(IntPtr ptr); +} + /// /// SafeHandle owning a rate_limiter. Released via /// wf_rate_limiter_free when disposed or finalized. @@ -8220,6 +8409,48 @@ internal sealed partial class RateLimiterHandle : SafeHandle private static partial void wf_rate_limiter_free(IntPtr ptr); } +/// +/// SafeHandle owning a relay_config. Released via +/// wf_relay_config_free when disposed or finalized. +/// +internal sealed partial class RelayConfigHandle : SafeHandle +{ + public RelayConfigHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_relay_config_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_relay_config_free(IntPtr ptr); +} + +/// +/// SafeHandle owning a relay_server. Released via +/// wf_relay_server_free when disposed or finalized. +/// +internal sealed partial class RelayServerHandle : SafeHandle +{ + public RelayServerHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_relay_server_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_relay_server_free(IntPtr ptr); +} + /// /// SafeHandle owning a repo_apply_writes_result. Released via /// wf_repo_apply_writes_result_free when disposed or finalized. @@ -8682,6 +8913,27 @@ internal sealed partial class ServerSessionTokensHandle : SafeHandle private static partial void wf_server_session_tokens_free(IntPtr ptr); } +/// +/// SafeHandle owning a service_auth_claims. Released via +/// wf_service_auth_claims_free when disposed or finalized. +/// +internal sealed partial class ServiceAuthClaimsHandle : SafeHandle +{ + public ServiceAuthClaimsHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_service_auth_claims_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_service_auth_claims_free(IntPtr ptr); +} + /// /// SafeHandle owning a session. Released via /// wf_session_free when disposed or finalized. @@ -8703,6 +8955,27 @@ internal sealed partial class SessionHandle : SafeHandle private static partial void wf_session_free(IntPtr ptr); } +/// +/// SafeHandle owning a subscribe_event. Released via +/// wf_subscribe_event_free when disposed or finalized. +/// +internal sealed partial class SubscribeEventHandle : SafeHandle +{ + public SubscribeEventHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_subscribe_event_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_subscribe_event_free(IntPtr ptr); +} + /// /// SafeHandle owning a sync_blob_cid_list. Released via /// wf_sync_blob_cid_list_free when disposed or finalized. @@ -9416,3 +9689,46 @@ internal sealed partial class XrpcServerHandle : SafeHandle [LibraryImport("wolfram")] private static partial void wf_xrpc_server_free(IntPtr ptr); } + +/// +/// SafeHandle owning a xrpc_server_auth_config. Released via +/// wf_xrpc_server_auth_config_free when disposed or finalized. +/// +internal sealed partial class XrpcServerAuthConfigHandle : SafeHandle +{ + public XrpcServerAuthConfigHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_xrpc_server_auth_config_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_xrpc_server_auth_config_free(IntPtr ptr); +} + +/// +/// SafeHandle owning a xrpc_server_config. Released via +/// wf_xrpc_server_config_free when disposed or finalized. +/// +internal sealed partial class XrpcServerConfigHandle : SafeHandle +{ + public XrpcServerConfigHandle(IntPtr ptr, bool ownsHandle = true) + : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr); + + public override bool IsInvalid => handle == IntPtr.Zero; + + protected override bool ReleaseHandle() + { + wf_xrpc_server_config_free(handle); + return true; + } + + [LibraryImport("wolfram")] + private static partial void wf_xrpc_server_config_free(IntPtr ptr); +} + diff --git a/dotnet/Wolfram.Interop/Wolfram.Interop.csproj b/dotnet/Wolfram.Interop/Wolfram.Interop.csproj index 8b1e669..233f81f 100644 --- a/dotnet/Wolfram.Interop/Wolfram.Interop.csproj +++ b/dotnet/Wolfram.Interop/Wolfram.Interop.csproj @@ -22,14 +22,20 @@ - - + + diff --git a/dotnet/Wolfram.Interop/tools/gen_safehandles.cpp b/dotnet/Wolfram.Interop/tools/gen_safehandles.cpp new file mode 100644 index 0000000..76ea400 --- /dev/null +++ b/dotnet/Wolfram.Interop/tools/gen_safehandles.cpp @@ -0,0 +1,143 @@ +#include +#include +#include +#include +#include +#include +#include +#include + +namespace fs = std::filesystem; + +std::string to_pascal_case(const std::string& snake) { + std::string result; + size_t start = 0; + while (start <= snake.size()) { + size_t end = snake.find('_', start); + if (end == std::string::npos) end = snake.size(); + if (end > start) { + result += std::toupper(snake[start]); + for (size_t i = start + 1; i < end; ++i) + result += std::tolower(snake[i]); + } + if (end == snake.size()) break; + start = end + 1; + } + return result; +} + +int main(int argc, char* argv[]) { + if (argc < 3) { + std::cerr << "Usage: gen_safehandles \n"; + return 1; + } + + fs::path inc_dir(argv[1]); + fs::path out_path(argv[2]); + + // Match `void wf__free(wf_ *name)` + std::regex header_re(R"(void\s+wf_([A-Za-z0-9_]+)_free\s*\(\s*wf_\1\s*\*\s*\w+\s*\))"); + + // Also match `void wf__free(wf_ *name, size_t n)` for array free functions - skip these + std::regex array_free_re(R"(void\s+wf_([A-Za-z0-9_]+)_free\s*\(\s*wf_\1\s*\*\s*\w+\s*,)"); + + // Types that have hand-written SafeHandles in the managed tier + std::set skip_types = {"xrpc_client"}; + + std::map types; // type -> free function name + std::set array_types; + int header_count = 0; + + for (const auto& entry : fs::recursive_directory_iterator(inc_dir)) { + if (entry.path().extension() != ".h") continue; + if (entry.path().string().find("json-enhanced") != std::string::npos) continue; + header_count++; + + std::ifstream file(entry.path()); + std::string text((std::istreambuf_iterator(file)), std::istreambuf_iterator()); + + // Collect array-free types (to exclude from SafeHandle generation) + { + std::smatch match; + std::string::const_iterator search_start(text.cbegin()); + while (std::regex_search(search_start, text.cend(), match, array_free_re)) { + array_types.insert(match[1].str()); + search_start = match.suffix().first; + } + } + + { + std::smatch match; + std::string::const_iterator search_start(text.cbegin()); + while (std::regex_search(search_start, text.cend(), match, header_re)) { + std::string t = match[1].str(); + if (types.find(t) == types.end() && + array_types.find(t) == array_types.end() && + skip_types.find(t) == skip_types.end()) { + types[t] = "wf_" + t + "_free"; + } + search_start = match.suffix().first; + } + } + } + + std::vector lines; + lines.emplace_back("// Generated by tools/gen_safehandles.cpp; do not edit."); + lines.emplace_back("// SafeHandle subclasses for every wolfram opaque type that has a"); + lines.emplace_back("// `void wf__free(wf_*)` function. Regenerated from the C headers."); + lines.emplace_back("// Each SafeHandle declares its own [LibraryImport] for the free function,"); + lines.emplace_back("// so this file is self-contained and does not depend on Raw.cs."); + lines.emplace_back(""); + lines.emplace_back("using System;"); + lines.emplace_back("using System.Runtime.InteropServices;"); + lines.emplace_back(""); + lines.emplace_back("namespace Wolfram.Interop;"); + lines.emplace_back(""); + + std::vector sorted_types; + for (const auto& [t, _] : types) { + sorted_types.push_back(t); + } + std::sort(sorted_types.begin(), sorted_types.end()); + + for (const auto& t : sorted_types) { + std::string pascal = to_pascal_case(t); + std::string free_fn = types.at(t); + lines.emplace_back("/// "); + lines.emplace_back("/// SafeHandle owning a " + t + ". Released via"); + lines.emplace_back("/// " + free_fn + " when disposed or finalized."); + lines.emplace_back("/// "); + lines.emplace_back("internal sealed partial class " + pascal + "Handle : SafeHandle"); + lines.emplace_back("{"); + lines.emplace_back(" public " + pascal + "Handle(IntPtr ptr, bool ownsHandle = true)"); + lines.emplace_back(" : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr);"); + lines.emplace_back(""); + lines.emplace_back(" public override bool IsInvalid => handle == IntPtr.Zero;"); + lines.emplace_back(""); + lines.emplace_back(" protected override bool ReleaseHandle()"); + lines.emplace_back(" {"); + lines.emplace_back(" " + free_fn + "(handle);"); + lines.emplace_back(" return true;"); + lines.emplace_back(" }"); + lines.emplace_back(""); + lines.emplace_back(" [LibraryImport(\"wolfram\")]"); + lines.emplace_back(" private static partial void " + free_fn + "(IntPtr ptr);"); + lines.emplace_back("}"); + lines.emplace_back(""); + } + + // Remove trailing blank line + if (!lines.empty() && lines.back().empty()) { + lines.pop_back(); + } + lines.emplace_back(""); + + fs::create_directories(out_path.parent_path()); + std::ofstream out(out_path); + for (const auto& line : lines) { + out << line << "\n"; + } + + std::cout << "wrote " << out_path << ": " << types.size() << " SafeHandle types from " << header_count << " headers\n"; + return 0; +} diff --git a/dotnet/Wolfram.Interop/tools/gen_safehandles.py b/dotnet/Wolfram.Interop/tools/gen_safehandles.py deleted file mode 100644 index 62a78f1..0000000 --- a/dotnet/Wolfram.Interop/tools/gen_safehandles.py +++ /dev/null @@ -1,123 +0,0 @@ -#!/usr/bin/env python3 -"""Generate C# SafeHandle subclasses from wolfram's wf_*_free set. - -Scans the wolfram C public headers for ``void wf__free(wf_ *...)`` and -emits a C# source file declaring ``SafeHandle`` subclasses that call the -matching free function in their ``ReleaseHandle`` override. Each SafeHandle -is self-contained: it declares its own ``[LibraryImport]`` for the free -function, so it does not depend on the hand-written ``Raw.cs`` having a -matching declaration. - -This mirrors the C++ gen_owners.py approach: the C headers are the single -source of truth, and the wrapper layer is generated from them. - -Usage: - python3 tools/gen_safehandles.py -""" -import re -import sys -from pathlib import Path - -# Match `void wf__free(wf_ *name)`. The backreference ensures the freed -# type matches the function prefix. Excludes `wf_*_json_free(char*)` (strings) -# and `wf_*_free(char*)` string free functions. -HEADER_RE = re.compile( - r"void\s+wf_([A-Za-z0-9_]+)_free\s*\(\s*wf_\1\s*\*\s*\w+\s*\)" -) - -# Also match `void wf__free(wf_ *name, size_t n)` for array free functions -# — these are not SafeHandle-able (they need a count), so we skip them. -ARRAY_FREE_RE = re.compile( - r"void\s+wf_([A-Za-z0-9_]+)_free\s*\(\s*wf_\1\s*\*\s*\w+\s*," -) - - -def to_pascal_case(snake: str) -> str: - """Convert ``snake_case`` to ``PascalCase``.""" - return "".join(part.capitalize() for part in snake.split("_") if part) - - -def main() -> int: - inc_dir = Path(sys.argv[1] if len(sys.argv) > 1 else "../../include/wolfram") - out = Path(sys.argv[2]) if len(sys.argv) > 2 else Path("SafeHandles/SafeHandles.generated.cs") - - # Types that have hand-written SafeHandles in the managed tier. - # The generator skips these to avoid duplicate declarations. - skip_types: set[str] = { - "xrpc_client", # SafeHandle/XrpcClientHandle.cs (public, used by WolframClient) - } - - types: dict[str, str] = {} # type_name -> free_function_name - array_types: set[str] = set() - header_count = 0 - - for h in sorted(inc_dir.rglob("*.h")): - if "json-enhanced" in h.parts: - continue - header_count += 1 - text = h.read_text(encoding="utf-8", errors="ignore") - - # Collect array-free types (to exclude them from SafeHandle generation) - for m in ARRAY_FREE_RE.finditer(text): - array_types.add(m.group(1)) - - for m in HEADER_RE.finditer(text): - t = m.group(1) - if t not in types and t not in array_types and t not in skip_types: - types[t] = f"wf_{t}_free" - - lines = [ - "// Generated by tools/gen_safehandles.py; do not edit.", - "// SafeHandle subclasses for every wolfram opaque type that has a", - "// `void wf__free(wf_*)` function. Regenerated from the C headers.", - "// Each SafeHandle declares its own [LibraryImport] for the free function,", - "// so this file is self-contained and does not depend on Raw.cs.", - "", - "using System;", - "using System.Runtime.InteropServices;", - "", - "namespace Wolfram.Interop;", - "", - ] - - for t in sorted(types): - pascal = to_pascal_case(t) - free_fn = types[t] - lines += [ - f"/// ", - f"/// SafeHandle owning a {t}. Released via", - f"/// {free_fn} when disposed or finalized.", - f"/// ", - f"internal sealed partial class {pascal}Handle : SafeHandle", - f"{{", - f" public {pascal}Handle(IntPtr ptr, bool ownsHandle = true)", - f" : base(IntPtr.Zero, ownsHandle) => SetHandle(ptr);", - f"", - f" public override bool IsInvalid => handle == IntPtr.Zero;", - f"", - f" protected override bool ReleaseHandle()", - f" {{", - f" {free_fn}(handle);", - f" return true;", - f" }}", - f"", - f" [LibraryImport(\"wolfram\")]", - f" private static partial void {free_fn}(IntPtr ptr);", - f"}}", - f"", - ] - - # Remove trailing blank line - if lines and lines[-1] == "": - lines.pop() - - lines.append("") - - out.parent.mkdir(parents=True, exist_ok=True) - out.write_text("\n".join(lines), encoding="utf-8") - print(f"wrote {out}: {len(types)} SafeHandle types from {header_count} headers") - return 0 - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/test/cpp_api_readiness_test.cpp b/test/cpp_api_readiness_test.cpp new file mode 100644 index 0000000..f7dfef5 --- /dev/null +++ b/test/cpp_api_readiness_test.cpp @@ -0,0 +1,16 @@ +#include +#include "wolfram-cpp/wolfram/wolfram.hpp" + +int main() { + // Test RAII handle auto-free + auto handle = wolfram::make_handle(wf_actor_new()); + assert(handle != nullptr); + // No explicit free needed + + // Test generated owners + using actor_handle = wolfram::unique_handle; + actor_handle actor(wf_actor_new()); + assert(actor.get() != nullptr); + // Destructs automatically + return 0; +} \ No newline at end of file diff --git a/test/test_lexgen.cpp b/test/test_lexgen.cpp new file mode 100644 index 0000000..daa257a --- /dev/null +++ b/test/test_lexgen.cpp @@ -0,0 +1,801 @@ +// test_lexgen — C++ port of the Python test_lexgen.py suite. +// +// Runs the wf_lexgen codegen tool as a subprocess, asserts on its generated +// output, and for the codec/transport tests compiles and runs small C check +// programs (cc -std=c11 -Wall -Wextra -Werror) that link the generated source +// against cJSON and OpenSSL, exactly as the Python test did. +// +// Paths are injected at build time by CMake: +// WF_TEST_ROOT — repository root +// WF_TEST_GENERATOR — path to the wf_lexgen_tool binary +// WF_TEST_CJSON_INCLUDE — cJSON source dir (cJSON.h) +// WF_TEST_CJSON_LIB — cJSON build dir (libcjson) +// WF_TEST_OPENSSL_INCLUDE — OpenSSL header dir +// WF_TEST_OPENSSL_CRYPTO — OpenSSL crypto library path + +#include + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include +#include +#include +#include + +namespace fs = std::filesystem; + +#ifndef WF_TEST_ROOT +#define WF_TEST_ROOT "." +#endif +#ifndef WF_TEST_GENERATOR +#define WF_TEST_GENERATOR "wf_lexgen" +#endif +#ifndef WF_TEST_CJSON_INCLUDE +#define WF_TEST_CJSON_INCLUDE "." +#endif +#ifndef WF_TEST_CJSON_LIB +#define WF_TEST_CJSON_LIB "." +#endif +#ifndef WF_TEST_OPENSSL_INCLUDE +#define WF_TEST_OPENSSL_INCLUDE "." +#endif +#ifndef WF_TEST_OPENSSL_CRYPTO +#define WF_TEST_OPENSSL_CRYPTO "" +#endif + +static const std::string ROOT = WF_TEST_ROOT; +static const std::string GENERATOR = WF_TEST_GENERATOR; +static const std::string CJSON_INCLUDE = WF_TEST_CJSON_INCLUDE; +static const std::string CJSON_LIB = WF_TEST_CJSON_LIB; +static const std::string OPENSSL_INCLUDE = WF_TEST_OPENSSL_INCLUDE; +static const std::string OPENSSL_CRYPTO = WF_TEST_OPENSSL_CRYPTO; + +// --------------------------------------------------------------------------- +// Subprocess helper: fork/exec with stdout+stderr capture. +// --------------------------------------------------------------------------- + +struct CommandResult { + int status; + std::string stdout_text; + std::string stderr_text; +}; + +static CommandResult run_command(const std::vector& args, + const std::vector& extra_env = {}) { + int out_pipe[2]; + int err_pipe[2]; + if (pipe(out_pipe) != 0 || pipe(err_pipe) != 0) + throw std::runtime_error("test_lexgen: pipe() failed"); + pid_t pid = fork(); + if (pid < 0) { + close(out_pipe[0]); + close(out_pipe[1]); + close(err_pipe[0]); + close(err_pipe[1]); + throw std::runtime_error("test_lexgen: fork() failed"); + } + if (pid == 0) { + close(out_pipe[0]); + close(err_pipe[0]); + dup2(out_pipe[1], STDOUT_FILENO); + dup2(err_pipe[1], STDERR_FILENO); + close(out_pipe[1]); + close(err_pipe[1]); + for (const std::string& entry : extra_env) { + std::string::size_type eq = entry.find('='); + if (eq != std::string::npos) + setenv(entry.substr(0, eq).c_str(), entry.substr(eq + 1).c_str(), 1); + } + std::vector argv; + for (const std::string& a : args) + argv.push_back(const_cast(a.c_str())); + argv.push_back(nullptr); + execvp(argv[0], argv.data()); + std::string err = std::string("test_lexgen: execvp(") + args[0] + ") failed: " + + strerror(errno) + "\n"; + write(STDERR_FILENO, err.data(), err.size()); + _exit(127); + } + + close(out_pipe[1]); + close(err_pipe[1]); + std::string outputs[2]; + bool open_fds[2] = {true, true}; + while (open_fds[0] || open_fds[1]) { + struct pollfd pfd[2]; + int index[2]; + int nfds = 0; + for (int i = 0; i < 2; ++i) { + if (open_fds[i]) { + pfd[nfds].fd = i == 0 ? out_pipe[0] : err_pipe[0]; + pfd[nfds].events = POLLIN; + pfd[nfds].revents = 0; + index[nfds] = i; + ++nfds; + } + } + int polled = poll(pfd, nfds, -1); + if (polled < 0) { + if (errno == EINTR) + continue; + break; + } + for (int i = 0; i < nfds; ++i) { + if (pfd[i].revents & (POLLIN | POLLHUP | POLLERR)) { + char buf[8192]; + ssize_t n = read(pfd[i].fd, buf, sizeof(buf)); + if (n > 0) { + outputs[index[i]].append(buf, static_cast(n)); + } else if (n == 0) { + open_fds[index[i]] = false; + } else if (errno != EINTR) { + open_fds[index[i]] = false; + } + } + } + } + close(out_pipe[0]); + close(err_pipe[0]); + + int wstatus = 0; + waitpid(pid, &wstatus, 0); + int status; + if (WIFEXITED(wstatus)) + status = WEXITSTATUS(wstatus); + else if (WIFSIGNALED(wstatus)) + status = 128 + WTERMSIG(wstatus); + else + status = 1; + return CommandResult{status, std::move(outputs[0]), std::move(outputs[1])}; +} + +// --------------------------------------------------------------------------- +// File / directory helpers. +// --------------------------------------------------------------------------- + +static std::string read_file(const fs::path& path) { + std::ifstream in(path); + if (!in) + throw std::runtime_error("test_lexgen: cannot read " + path.string()); + std::ostringstream ss; + ss << in.rdbuf(); + return ss.str(); +} + +static void write_file(const fs::path& path, const std::string& content) { + std::ofstream out(path); + if (!out) + throw std::runtime_error("test_lexgen: cannot write " + path.string()); + out << content; +} + +class TempDir { +public: + fs::path path; + + TempDir() { + static int counter = 0; + path = fs::temp_directory_path() / + ("wf_lexgen_test_" + std::to_string(::getpid()) + "_" + + std::to_string(counter++)); + fs::create_directories(path); + } + + ~TempDir() { + std::error_code ec; + fs::remove_all(path, ec); + } +}; + +// --------------------------------------------------------------------------- +// Generator invocation helpers (mirror the Python subprocess.run calls). +// --------------------------------------------------------------------------- + +static std::string generate_to_stdout(const std::vector& lexicons) { + std::vector args = {GENERATOR}; + for (const std::string& path : lexicons) + args.push_back(path); + CommandResult res = run_command(args); + if (res.status != 0) + throw std::runtime_error("generator exited " + std::to_string(res.status) + ":\n" + + res.stdout_text + res.stderr_text); + return res.stdout_text; +} + +// Generate a header into `header` and, when `source` is non-empty, also a +// source file via --source-output, mirroring the Python argument order. +static void generate(const std::string& fixture, const fs::path& header, + const fs::path& source) { + std::vector args = {GENERATOR, fixture, "-o", header.string()}; + if (!source.empty()) { + args.push_back("--source-output"); + args.push_back(source.string()); + } + CommandResult res = run_command(args); + if (res.status != 0) + throw std::runtime_error("generator exited " + std::to_string(res.status) + ":\n" + + res.stdout_text + res.stderr_text); +} + +// --------------------------------------------------------------------------- +// Compile + run helpers (cc -std=c11 -Wall -Wextra -Werror, linking the +// generated source against cJSON and OpenSSL like the Python test did). +// --------------------------------------------------------------------------- + +static void compile_and_run(const fs::path& dir, const fs::path& generated_c, + const fs::path& check_c, const fs::path& executable) { + std::vector args = { + "cc", "-std=c11", "-Wall", "-Wextra", "-Werror", + "-I", dir.string(), + "-I", ROOT + "/include", + "-I", CJSON_INCLUDE, + "-I", OPENSSL_INCLUDE, + generated_c.string(), check_c.string(), + "-L", CJSON_LIB, "-lcjson", + OPENSSL_CRYPTO, + "-o", executable.string(), + }; + CommandResult compile = run_command(args); + if (compile.status != 0) + throw std::runtime_error("cc failed:\n" + compile.stdout_text + compile.stderr_text); + + std::vector env = { + "DYLD_LIBRARY_PATH=" + CJSON_LIB, + "LD_LIBRARY_PATH=" + CJSON_LIB, + }; + CommandResult run = run_command({executable.string()}, env); + if (run.status != 0) + throw std::runtime_error("check binary exited " + std::to_string(run.status) + ":\n" + + run.stdout_text + run.stderr_text); +} + +// --------------------------------------------------------------------------- +// Assertion helpers. +// --------------------------------------------------------------------------- + +static void assert_true(bool condition, const std::string& message) { + if (!condition) + throw std::runtime_error(message); +} + +static void assert_contains(const std::string& haystack, const std::string& needle, + const std::string& what) { + if (haystack.find(needle) == std::string::npos) + throw std::runtime_error(what + ": generated output is missing:\n" + needle); +} + +// --------------------------------------------------------------------------- +// Naming helpers mirroring the Python generator (needed to compute the +// endpoint symbols asserted against include/wolfram/atproto_lex.h). +// --------------------------------------------------------------------------- + +static const std::set C_KEYWORDS = { + "auto", "break", "case", "char", "const", "continue", "default", + "do", "double", "else", "enum", "extern", "float", "for", "goto", + "if", "inline", "int", "long", "register", "restrict", "return", + "short", "signed", "sizeof", "static", "struct", "switch", "typedef", + "union", "unsigned", "void", "volatile", "while", "_Alignas", + "_Alignof", "_Atomic", "_Bool", "_Complex", "_Generic", "_Imaginary", + "_Noreturn", "_Static_assert", "_Thread_local", +}; + +static const std::set PY_KEYWORDS = { + "False", "None", "True", "and", "as", "assert", "async", "await", + "break", "class", "continue", "def", "del", "elif", "else", "except", + "finally", "for", "from", "global", "if", "import", "in", "is", + "lambda", "nonlocal", "not", "or", "pass", "raise", "return", "try", + "while", "with", "yield", +}; + +static bool is_lower_or_digit(char c) { + unsigned char uc = static_cast(c); + return std::islower(uc) != 0 || std::isdigit(uc) != 0; +} + +static std::string snake(const std::string& value) { + std::string camel; + for (size_t i = 0; i < value.size(); ++i) { + char c = value[i]; + if (i > 0 && is_lower_or_digit(value[i - 1]) && + std::isupper(static_cast(c)) != 0) + camel += '_'; + camel += c; + } + std::string out; + bool pending = false; + for (char c : camel) { + if (std::isalnum(static_cast(c)) != 0) { + if (pending) { + out += '_'; + pending = false; + } + out += static_cast(std::tolower(static_cast(c))); + } else { + pending = true; + } + } + if (out.empty()) + out = "value"; + if (std::isdigit(static_cast(out[0])) != 0 || + C_KEYWORDS.count(out) != 0 || PY_KEYWORDS.count(out) != 0) + out += '_'; + return out; +} + +static std::string type_name(const std::string& nsid, const std::string& suffix) { + return "wf_lex_" + snake(nsid) + (suffix.empty() ? "" : "_" + snake(suffix)); +} + +// --------------------------------------------------------------------------- +// The ported test methods. +// --------------------------------------------------------------------------- + +static void test_generates_endpoint_and_named_types() { + std::string header = generate_to_stdout( + {ROOT + "/test/fixtures/lexicons/com.example.echo.json"}); + assert_contains(header, "#define WF_LEX_COM_EXAMPLE_ECHO_NSID \"com.example.echo\"", + "NSID define"); + assert_contains(header, "#define WF_LEX_COM_EXAMPLE_ECHO_KIND \"procedure\"", + "KIND define"); + assert_contains(header, "wf_lex_com_example_echo_main_input", "main input type"); + assert_contains(header, "const char * message;", "message field"); + assert_contains(header, "int64_t attempts;", "attempts field"); + assert_contains(header, "bool has_enabled;", "has_enabled field"); + assert_contains(header, "bool has_tags;", "has_tags field"); + assert_contains(header, "WF_LEX_ARRAY(const char *) tags;", "tags field"); + assert_contains(header, "wf_lex_json metadata;", "metadata field"); + assert_contains(header, "wf_lex_com_example_echo_named", "named type"); +} + +static void test_output_is_deterministic() { + std::string fixture = ROOT + "/test/fixtures/lexicons/com.example.echo.json"; + assert_true(generate_to_stdout({fixture}) == generate_to_stdout({fixture}), + "generator output is not deterministic"); +} + +static void test_generated_client_covers_bundled_endpoint_lexicons() { + std::string header = read_file(ROOT + "/include/wolfram/atproto_lex.h"); + std::set subscriptions; + int endpoint_count = 0; + int file_count = 0; + for (const fs::directory_entry& entry : fs::recursive_directory_iterator(ROOT + "/lexicons")) { + if (!entry.is_regular_file() || entry.path().extension().string() != ".json") + continue; + ++file_count; + std::string text = read_file(entry.path()); + cJSON* document = cJSON_Parse(text.c_str()); + if (!document) + throw std::runtime_error("cannot parse " + entry.path().string()); + cJSON* id_item = cJSON_GetObjectItemCaseSensitive(document, "id"); + std::string nsid = (id_item && cJSON_IsString(id_item) && id_item->valuestring) + ? id_item->valuestring + : ""; + cJSON* defs = cJSON_GetObjectItemCaseSensitive(document, "defs"); + if (defs && cJSON_IsObject(defs)) { + for (cJSON* def = defs->child; def; def = def->next) { + if (!def->string) + continue; + cJSON* type_item = cJSON_GetObjectItemCaseSensitive(def, "type"); + std::string kind = (type_item && cJSON_IsString(type_item) && type_item->valuestring) + ? type_item->valuestring + : ""; + if (kind == "subscription") { + subscriptions.insert(nsid); + continue; + } + if (kind != "query" && kind != "procedure") + continue; + ++endpoint_count; + std::string symbol = type_name(nsid, def->string); + assert_contains(header, "wf_status " + symbol + "_call(", + "endpoint call " + symbol); + assert_contains(header, "wf_status " + symbol + "_call_auth(", + "endpoint auth call " + symbol); + } + } + cJSON_Delete(document); + } + assert_true(endpoint_count == 312, + "endpoint count is " + std::to_string(endpoint_count) + ", expected 312"); + std::set expected = { + "chat.bsky.moderation.subscribeModEvents", + "com.atproto.label.subscribeLabels", + "com.atproto.sync.subscribeRepos", + }; + assert_true(subscriptions == expected, "subscription NSID set does not match"); + std::string subscription_apis = + read_file(ROOT + "/include/wolfram/sync_subscribe.h") + + read_file(ROOT + "/include/wolfram/label.h") + + read_file(ROOT + "/include/wolfram/chat_typed.h"); + assert_contains(subscription_apis, "wf_subscribe_start(", "wf_subscribe_start"); + assert_contains(subscription_apis, "wf_label_subscribe_start(", "wf_label_subscribe_start"); + assert_contains(subscription_apis, "wf_agent_chat_subscribe_mod_events_typed(", + "wf_agent_chat_subscribe_mod_events_typed"); + std::cout << " bundled lexicons: " << file_count << " files, " << endpoint_count + << " endpoints\n"; +} + +static void test_generated_header_compiles_as_c11() { + TempDir dir; + fs::path header = dir.path / "generated.h"; + fs::path check = dir.path / "check.c"; + generate(ROOT + "/test/fixtures/lexicons/com.example.echo.json", header, {}); + write_file(check, "#include \"generated.h\"\n" + "int main(void) {\n" + " wf_lex_com_example_echo_main_input value = {0};\n" + " value.message = \"hello\";\n" + " return value.message == 0;\n" + "}\n"); + std::vector args = { + "cc", "-std=c11", "-Wall", "-Wextra", "-Werror", "-fsyntax-only", + "-I", ROOT + "/include", check.string(), + }; + CommandResult res = run_command(args); + if (res.status != 0) + throw std::runtime_error("header C11 compile failed:\n" + res.stdout_text + + res.stderr_text); +} + +static void test_union_array_cleanup_casts_away_borrowed_view_const() { + TempDir dir; + fs::path fixture = dir.path / "com.example.union-array.json"; + fs::path header = dir.path / "generated.h"; + fs::path source = dir.path / "generated.c"; + write_file(fixture, R"lexchk({ + "lexicon": 1, + "id": "com.example.unionArray", + "defs": { + "main": { + "type": "query", + "output": { + "encoding": "application/json", + "schema": { + "type": "object", + "required": ["items"], + "properties": { + "items": { + "type": "array", + "items": { + "type": "union", + "refs": ["#entry"], + "closed": true + } + } + } + } + } + }, + "entry": { + "type": "object", + "required": ["value"], + "properties": {"value": {"type": "string"}} + } + } +})lexchk"); + generate(fixture, header, source); + std::string generated = read_file(source); + std::string union_type = "wf_lex_com_example_union_array_main_output_items_item_union"; + assert_contains(generated, "wf_lex_clear_" + union_type + "((" + union_type + " *)&(", + "union-array cleanup cast"); +} + +static void test_inline_objects_are_dependency_safe_and_deterministic() { + std::string fixture = ROOT + "/test/fixtures/lexicons/com.example.inline.json"; + std::string header = generate_to_stdout({fixture}); + std::string nested = "typedef struct wf_lex_com_example_inline_main_input_config_nested {"; + std::string parent = "typedef struct wf_lex_com_example_inline_main_input_config {"; + size_t nested_pos = header.find(nested); + size_t parent_pos = header.find(parent); + assert_true(nested_pos != std::string::npos, "nested inline typedef missing"); + assert_true(parent_pos != std::string::npos, "parent inline typedef missing"); + assert_true(nested_pos < parent_pos, "nested typedef must precede parent typedef"); + assert_contains(header, + "WF_LEX_ARRAY(wf_lex_com_example_inline_main_input_entries_item) entries;", + "entries array type"); + assert_true(header == generate_to_stdout({fixture}), + "inline output is not deterministic"); +} + +static void test_inline_object_codecs_run() { + TempDir dir; + fs::path header = dir.path / "generated.h"; + fs::path generated = dir.path / "generated.c"; + fs::path check = dir.path / "check.c"; + fs::path executable = dir.path / "check"; + generate(ROOT + "/test/fixtures/lexicons/com.example.inline.json", header, generated); + write_file(check, R"lexchk(#include "generated.h" +#include +#include +wf_status wf_xrpc_procedure(wf_xrpc_client *client, const char *nsid, + const char *json, wf_response *out) { + (void)client; (void)nsid; (void)json; (void)out; return WF_OK; +} +wf_status wf_auth_client_procedure(wf_auth_client *client, const char *nsid, + const char *json, wf_response *out) { + (void)client; (void)nsid; (void)json; (void)out; return WF_OK; +} +int main(void) { + wf_lex_com_example_inline_main_input_entries_item entries[] = {{2}, {3}}; + wf_lex_com_example_inline_main_input input = {0}; + input.config.name = "demo"; input.config.nested.enabled = true; + input.entries.items = entries; input.entries.count = 2; + char *json = NULL; + assert(wf_lex_com_example_inline_main_input_encode_json(&input, &json) == WF_OK); + assert(strcmp(json, "{\"config\":{\"name\":\"demo\",\"nested\":{\"enabled\":true}},\"entries\":[{\"count\":2},{\"count\":3}]}") == 0); + wf_lex_com_example_inline_main_json_free(json); + const char body[] = "{\"result\":{\"values\":[{\"label\":\"one\"},{\"label\":\"two\"}]}}"; + wf_lex_com_example_inline_main_output *output = NULL; + assert(wf_lex_com_example_inline_main_output_decode_json(body, strlen(body), &output) == WF_OK); + assert(output->result.values.count == 2); + assert(strcmp(output->result.values.items[1].label, "two") == 0); + wf_lex_com_example_inline_main_output_free(output); + return 0; +} +)lexchk"); + compile_and_run(dir.path, generated, check, executable); +} + +static void test_referenced_inputs_and_all_json_value_kinds_run() { + TempDir dir; + fs::path header = dir.path / "generated.h"; + fs::path generated = dir.path / "generated.c"; + fs::path check = dir.path / "check.c"; + fs::path executable = dir.path / "check"; + generate(ROOT + "/test/fixtures/lexicons/com.example.refs.json", header, generated); + write_file(check, R"lexchk(#include "generated.h" +#include +#include + +wf_status wf_xrpc_procedure(wf_xrpc_client *client, const char *nsid, + const char *json, wf_response *out) { + (void)client; (void)nsid; (void)json; (void)out; return WF_OK; +} +wf_status wf_auth_client_procedure(wf_auth_client *client, const char *nsid, + const char *json, wf_response *out) { + (void)client; (void)nsid; (void)json; (void)out; return WF_OK; +} + +int main(void) { + const wf_lex_json metadata = {"{\"n\":1}", 7}; + const wf_lex_com_example_refs_config first = {"one", metadata}; + const wf_lex_com_example_refs_config second = {"two", metadata}; + const wf_lex_com_example_refs_config *configs[] = {&first, &second}; + const char *names[] = {"alice", "bob"}; + const int64_t counts[] = {-1, 2}; + const bool switches[] = {true, false}; + const uint8_t payload[] = {1, 2, 3}; + wf_lex_com_example_refs_main_input input = {0}; + input.config = &first; + input.configs.items = configs; input.configs.count = 2; + input.names.items = names; input.names.count = 2; + input.mode = "fast"; input.token = "com.example.refs#token"; + input.counts.items = counts; input.counts.count = 2; + input.switches.items = switches; input.switches.count = 2; + input.payload.data = payload; input.payload.length = sizeof(payload); + input.link.cid = "bafy-link"; + input.blob = (wf_lex_blob){"bafy-blob", "image/png", 42}; + char *json = NULL; + assert(wf_lex_com_example_refs_main_input_encode_json(&input, &json) == WF_OK); + assert(strcmp(json, "{\"config\":{\"name\":\"one\",\"metadata\":{\"n\":1}},\"configs\":[{\"name\":\"one\",\"metadata\":{\"n\":1}},{\"name\":\"two\",\"metadata\":{\"n\":1}}],\"names\":[\"alice\",\"bob\"],\"mode\":\"fast\",\"token\":\"com.example.refs#token\",\"counts\":[-1,2],\"switches\":[true,false],\"payload\":{\"$bytes\":\"AQID\"},\"link\":{\"$link\":\"bafy-link\"},\"blob\":{\"$type\":\"blob\",\"ref\":{\"$link\":\"bafy-blob\"},\"mimeType\":\"image/png\",\"size\":42}}") == 0); + wf_lex_com_example_refs_main_json_free(json); + input.config = NULL; + assert(wf_lex_com_example_refs_main_input_encode_json(&input, &json) == WF_ERR_INVALID_ARG); + return 0; +} +)lexchk"); + compile_and_run(dir.path, generated, check, executable); +} + +static void test_generates_json_codec_and_transport_wrapper() { + TempDir dir; + fs::path header = dir.path / "generated.h"; + fs::path source = dir.path / "generated.c"; + generate(ROOT + "/test/fixtures/lexicons/com.example.echo.json", header, source); + std::string generated = read_file(source); + assert_contains(generated, "cJSON_PrintUnformatted", "cJSON_PrintUnformatted"); + assert_contains(generated, "_output_decode_json", "output decode json"); + assert_contains(generated, "cJSON_GetObjectItemCaseSensitive(item, \"$bytes\")", + "$bytes lookup"); + assert_contains(generated, "cJSON_GetObjectItemCaseSensitive(item, \"$link\")", + "$link lookup"); + assert_contains(generated, "wf_xrpc_procedure(client, \"com.example.echo\"", + "procedure wrapper"); + generate(ROOT + "/test/fixtures/lexicons/com.example.echo.json", header, source); + assert_true(generated == read_file(source), "regenerated source differs"); +} + +static void test_query_wrapper_uses_encoded_xrpc_parameters() { + TempDir dir; + fs::path header = dir.path / "generated.h"; + fs::path source = dir.path / "generated.c"; + generate(ROOT + "/test/fixtures/lexicons/com.example.get.json", header, source); + std::string generated = read_file(source); + assert_contains(generated, "wf_xrpc_query_params(client, \"com.example.get\"", + "query params wrapper"); + assert_contains(generated, + "encoded[count++] = (wf_xrpc_param){\"limit\", number_values[number_count++]}", + "limit param"); + assert_contains(generated, + "encoded[count++] = (wf_xrpc_param){\"dids\", params->dids.items[i]}", + "dids param"); + assert_contains(generated, + "encoded[count++] = (wf_xrpc_param){\"ids\", number_values[number_count++]}", + "ids param"); + assert_contains(generated, + "encoded[count++] = (wf_xrpc_param){\"flags\", (params->flags.items[i] ? \"true\" : \"false\")}", + "flags param"); +} + +static void test_query_array_wrapper_runs_with_repeated_keys() { + TempDir dir; + fs::path header = dir.path / "generated.h"; + fs::path generated = dir.path / "generated.c"; + fs::path check = dir.path / "check.c"; + fs::path executable = dir.path / "check"; + generate(ROOT + "/test/fixtures/lexicons/com.example.get.json", header, generated); + write_file(check, R"lexchk(#include "generated.h" +#include +#include + +wf_status wf_xrpc_query_params(wf_xrpc_client *client, const char *nsid, + const wf_xrpc_param *params, size_t count, + wf_response *out) { + const char *names[] = {"name", "limit", "dids", "dids", "ids", "ids", "flags", "flags"}; + const char *values[] = {"alice", "42", "did:plc:a", "did:plc:b", "-7", "9", "true", "false"}; + assert(client && out && strcmp(nsid, "com.example.get") == 0 && count == 8); + for (size_t i = 0; i < count; ++i) { + assert(strcmp(params[i].name, names[i]) == 0); + assert(strcmp(params[i].value, values[i]) == 0); + } + return WF_OK; +} +wf_status wf_auth_client_query_params(wf_auth_client *client, const char *nsid, + const wf_xrpc_param *params, size_t count, + wf_response *out) { + (void)client; (void)nsid; (void)params; (void)count; (void)out; return WF_OK; +} + +int main(void) { + const char *dids[] = {"did:plc:a", "did:plc:b"}; + int64_t ids[] = {-7, 9}; + bool flags[] = {true, false}; + wf_lex_com_example_get_main_params params = {0}; + params.name = "alice"; params.dids.items = dids; params.dids.count = 2; + params.has_limit = true; params.limit = 42; + params.has_ids = true; params.ids.items = ids; params.ids.count = 2; + params.has_flags = true; params.flags.items = flags; params.flags.count = 2; + wf_response response = {0}; + assert(wf_lex_com_example_get_main_call((wf_xrpc_client *)1, ¶ms, &response) == WF_OK); + dids[1] = NULL; + assert(wf_lex_com_example_get_main_call((wf_xrpc_client *)1, ¶ms, &response) == WF_ERR_INVALID_ARG); + dids[1] = "did:plc:b"; + params.dids.items = NULL; + assert(wf_lex_com_example_get_main_call((wf_xrpc_client *)1, ¶ms, &response) == WF_ERR_INVALID_ARG); + return 0; +} +)lexchk"); + compile_and_run(dir.path, generated, check, executable); +} + +static void test_generated_codec_and_wrapper_run() { + TempDir dir; + fs::path header = dir.path / "generated.h"; + fs::path generated = dir.path / "generated.c"; + fs::path check = dir.path / "check.c"; + fs::path executable = dir.path / "check"; + generate(ROOT + "/test/fixtures/lexicons/com.example.echo.json", header, generated); + write_file(check, R"lexchk(#include "generated.h" +#include +#include +#include + +static int called; +wf_status wf_xrpc_procedure(wf_xrpc_client *client, const char *nsid, + const char *json, wf_response *out) { + assert(client != NULL); + assert(strcmp(nsid, "com.example.echo") == 0); + assert(strcmp(json, "{\"message\":\"hello\",\"attempts\":2,\"enabled\":true,\"tags\":[\"a\",\"b\"]}") == 0); + called = 1; out->status = 200; out->body = NULL; out->body_len = 0; return WF_OK; +} +wf_status wf_auth_client_procedure(wf_auth_client *client, const char *nsid, + const char *json, wf_response *out) { + (void)client; (void)nsid; (void)json; (void)out; return WF_OK; +} +int main(void) { + const char *tags[] = {"a", "b"}; + wf_lex_com_example_echo_main_input input = {0}; + input.message = "hello"; input.attempts = 2; + input.has_enabled = true; input.enabled = true; + input.has_tags = true; input.tags.items = tags; input.tags.count = 2; + char *json = NULL; + assert(wf_lex_com_example_echo_main_input_encode_json(&input, &json) == WF_OK); + assert(json != NULL); wf_lex_com_example_echo_main_json_free(json); + wf_response response = {0}; + assert(wf_lex_com_example_echo_main_call((wf_xrpc_client *)1, &input, &response) == WF_OK); + assert(called); + const char output_json[] = "{\"value\":\"ok\",\"items\":[{\"id\":\"one\",\"payload\":{\"$bytes\":\"AQID\"}}],\"raw\":{\"$bytes\":\"AAE=\"},\"link\":{\"$link\":\"bafytest\"},\"blob\":{\"$type\":\"blob\",\"ref\":{\"$link\":\"bafyblob\"},\"mimeType\":\"image/png\",\"size\":42},\"extra\":{\"x\":1},\"flags\":[true,false]}"; + wf_lex_com_example_echo_main_output *output = NULL; + assert(wf_lex_com_example_echo_main_output_decode_json(output_json, strlen(output_json), &output) == WF_OK); + assert(output && strcmp(output->value, "ok") == 0); + assert(output->items.count == 1 && strcmp(output->items.items[0]->id, "one") == 0); + assert(output->items.items[0]->payload.length == 3 && output->items.items[0]->payload.data[2] == 3); + assert(output->raw.length == 2 && output->raw.data[1] == 1); + assert(strcmp(output->link.cid, "bafytest") == 0); + assert(strcmp(output->blob.cid, "bafyblob") == 0 && output->blob.size == 42); + assert(output->has_flags && output->flags.count == 2 && output->flags.items[0]); + assert(output->extra.length == 7 && strcmp(output->extra.data, "{\"x\":1}") == 0); + wf_lex_com_example_echo_main_output_free(output); + output = (void *)1; + assert(wf_lex_com_example_echo_main_output_decode_json("{}", 2, &output) == WF_ERR_INVALID_ARG); + assert(output == NULL); + return 0; +} +)lexchk"); + compile_and_run(dir.path, generated, check, executable); +} + +// --------------------------------------------------------------------------- +// Main: run every ported test method and report. +// --------------------------------------------------------------------------- + +struct TestEntry { + const char* name; + void (*run)(); +}; + +int main() { + const TestEntry tests[] = { + {"generates_endpoint_and_named_types", test_generates_endpoint_and_named_types}, + {"output_is_deterministic", test_output_is_deterministic}, + {"generated_client_covers_bundled_endpoint_lexicons", + test_generated_client_covers_bundled_endpoint_lexicons}, + {"generated_header_compiles_as_c11", test_generated_header_compiles_as_c11}, + {"union_array_cleanup_casts_away_borrowed_view_const", + test_union_array_cleanup_casts_away_borrowed_view_const}, + {"inline_objects_are_dependency_safe_and_deterministic", + test_inline_objects_are_dependency_safe_and_deterministic}, + {"inline_object_codecs_run", test_inline_object_codecs_run}, + {"referenced_inputs_and_all_json_value_kinds_run", + test_referenced_inputs_and_all_json_value_kinds_run}, + {"generates_json_codec_and_transport_wrapper", + test_generates_json_codec_and_transport_wrapper}, + {"query_wrapper_uses_encoded_xrpc_parameters", + test_query_wrapper_uses_encoded_xrpc_parameters}, + {"query_array_wrapper_runs_with_repeated_keys", + test_query_array_wrapper_runs_with_repeated_keys}, + {"generated_codec_and_wrapper_run", test_generated_codec_and_wrapper_run}, + }; + + int failures = 0; + for (const TestEntry& test : tests) { + try { + test.run(); + std::cout << "ok " << test.name << "\n"; + } catch (const std::exception& error) { + std::cerr << "FAIL " << test.name << ": " << error.what() << "\n"; + ++failures; + } + } + if (failures != 0) { + std::cerr << failures << " test(s) failed\n"; + return 1; + } + std::cout << "all lexgen tests passed (" << sizeof(tests) / sizeof(tests[0]) + << " tests)\n"; + return 0; +} diff --git a/test/test_lexgen.py b/test/test_lexgen.py deleted file mode 100644 index ca3e348..0000000 --- a/test/test_lexgen.py +++ /dev/null @@ -1,486 +0,0 @@ -import json -import subprocess -import sys -import tempfile -import unittest -import os -import shlex -from pathlib import Path - -sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "tools")) -from wf_lexgen import type_name - - -ROOT = Path(__file__).resolve().parents[1] -GENERATOR = ROOT / "tools" / "wf_lexgen.py" -FIXTURE = ROOT / "test" / "fixtures" / "lexicons" / "com.example.echo.json" -QUERY_FIXTURE = ROOT / "test" / "fixtures" / "lexicons" / "com.example.get.json" -INLINE_FIXTURE = ROOT / "test" / "fixtures" / "lexicons" / "com.example.inline.json" -REFS_FIXTURE = ROOT / "test" / "fixtures" / "lexicons" / "com.example.refs.json" - - -def openssl_flags(): - """Return compiler/linker flags from pkg-config or the CMake cache.""" - pkg_config = subprocess.run( - ["pkg-config", "--cflags", "--libs", "openssl"], - capture_output=True, text=True, - ) - if pkg_config.returncode == 0: - return shlex.split(pkg_config.stdout) - - cache = ROOT / "build" / "CMakeCache.txt" - if cache.exists(): - values = {} - for line in cache.read_text(encoding="utf-8").splitlines(): - if line.startswith(("OPENSSL_INCLUDE_DIR:", - "OPENSSL_CRYPTO_LIBRARY:")): - key, value = line.split("=", 1) - values[key.split(":", 1)[0]] = value - include = values.get("OPENSSL_INCLUDE_DIR") - crypto = values.get("OPENSSL_CRYPTO_LIBRARY") - if include and crypto: - return [f"-I{include}", crypto] - - raise RuntimeError( - "OpenSSL compiler flags unavailable from pkg-config or CMake cache") - - -class LexgenTests(unittest.TestCase): - def generate(self, *inputs: Path) -> str: - result = subprocess.run( - [sys.executable, str(GENERATOR), *(str(path) for path in inputs)], - check=True, capture_output=True, text=True, - ) - return result.stdout - - def test_generates_endpoint_and_named_types(self): - header = self.generate(FIXTURE) - self.assertIn('#define WF_LEX_COM_EXAMPLE_ECHO_NSID "com.example.echo"', header) - self.assertIn('#define WF_LEX_COM_EXAMPLE_ECHO_KIND "procedure"', header) - self.assertIn("wf_lex_com_example_echo_main_input", header) - self.assertIn("const char * message;", header) - self.assertIn("int64_t attempts;", header) - self.assertIn("bool has_enabled;", header) - self.assertIn("bool has_tags;", header) - self.assertIn("WF_LEX_ARRAY(const char *) tags;", header) - self.assertIn("wf_lex_json metadata;", header) - self.assertIn("wf_lex_com_example_echo_named", header) - - def test_output_is_deterministic(self): - self.assertEqual(self.generate(FIXTURE), self.generate(FIXTURE)) - - def test_generated_client_covers_bundled_endpoint_lexicons(self): - header = (ROOT / "include" / "wolfram" / "atproto_lex.h").read_text( - encoding="utf-8") - subscriptions = set() - endpoint_count = 0 - for path in (ROOT / "lexicons").rglob("*.json"): - document = json.loads(path.read_text(encoding="utf-8")) - for name, definition in document.get("defs", {}).items(): - kind = definition.get("type") - if kind == "subscription": - subscriptions.add(document["id"]) - continue - if kind not in ("query", "procedure"): - continue - endpoint_count += 1 - symbol = type_name(document["id"], name) - self.assertIn(f"wf_status {symbol}_call(", header) - self.assertIn(f"wf_status {symbol}_call_auth(", header) - - self.assertEqual(endpoint_count, 312) - self.assertEqual(subscriptions, { - "chat.bsky.moderation.subscribeModEvents", - "com.atproto.label.subscribeLabels", - "com.atproto.sync.subscribeRepos", - }) - subscription_apis = ( - (ROOT / "include" / "wolfram" / "sync_subscribe.h").read_text() + - (ROOT / "include" / "wolfram" / "label.h").read_text() + - (ROOT / "include" / "wolfram" / "chat_typed.h").read_text()) - self.assertIn("wf_subscribe_start(", subscription_apis) - self.assertIn("wf_label_subscribe_start(", subscription_apis) - self.assertIn("wf_agent_chat_subscribe_mod_events_typed(", - subscription_apis) - - def test_generated_header_compiles_as_c11(self): - with tempfile.TemporaryDirectory() as directory: - directory = Path(directory) - header = directory / "generated.h" - source = directory / "check.c" - subprocess.run( - [sys.executable, str(GENERATOR), str(FIXTURE), "-o", str(header)], - check=True, - ) - source.write_text( - '#include "generated.h"\n' - "int main(void) {\n" - " wf_lex_com_example_echo_main_input value = {0};\n" - " value.message = \"hello\";\n" - " return value.message == 0;\n" - "}\n", - encoding="utf-8", - ) - subprocess.run( - ["cc", "-std=c11", "-Wall", "-Wextra", "-Werror", "-fsyntax-only", - "-I", str(ROOT / "include"), str(source)], check=True, - ) - - def test_union_array_cleanup_casts_away_borrowed_view_const(self): - with tempfile.TemporaryDirectory() as directory: - directory = Path(directory) - fixture = directory / "com.example.union-array.json" - header = directory / "generated.h" - source = directory / "generated.c" - fixture.write_text(json.dumps({ - "lexicon": 1, - "id": "com.example.unionArray", - "defs": { - "main": { - "type": "query", - "output": { - "encoding": "application/json", - "schema": { - "type": "object", - "required": ["items"], - "properties": { - "items": { - "type": "array", - "items": { - "type": "union", - "refs": ["#entry"], - "closed": True, - }, - }, - }, - }, - }, - }, - "entry": { - "type": "object", - "required": ["value"], - "properties": {"value": {"type": "string"}}, - }, - }, - }), encoding="utf-8") - subprocess.run([ - sys.executable, str(GENERATOR), str(fixture), - "-o", str(header), - "--source-output", str(source), - ], check=True, capture_output=True, text=True) - generated = source.read_text(encoding="utf-8") - union_type = ( - "wf_lex_com_example_union_array_main_output_items_item_union" - ) - self.assertIn( - f"wf_lex_clear_{union_type}(({union_type} *)&(", generated) - - def test_inline_objects_are_dependency_safe_and_deterministic(self): - header = self.generate(INLINE_FIXTURE) - nested = "typedef struct wf_lex_com_example_inline_main_input_config_nested {" - parent = "typedef struct wf_lex_com_example_inline_main_input_config {" - self.assertLess(header.index(nested), header.index(parent)) - self.assertIn("WF_LEX_ARRAY(wf_lex_com_example_inline_main_input_entries_item) entries;", header) - self.assertEqual(header, self.generate(INLINE_FIXTURE)) - - def test_inline_object_codecs_run(self): - cjson_include = ROOT / "build" / "_deps" / "cjson-src" - cjson_lib = ROOT / "build" / "_deps" / "cjson-build" - if not (cjson_include / "cJSON.h").exists(): - self.skipTest("configured cJSON dependency is not available") - with tempfile.TemporaryDirectory() as directory: - directory = Path(directory) - header = directory / "generated.h" - generated = directory / "generated.c" - check = directory / "check.c" - executable = directory / "check" - subprocess.run([sys.executable, str(GENERATOR), str(INLINE_FIXTURE), - "-o", str(header), "--source-output", str(generated)], check=True) - check.write_text(r'''#include "generated.h" -#include -#include -wf_status wf_xrpc_procedure(wf_xrpc_client *client, const char *nsid, - const char *json, wf_response *out) { - (void)client; (void)nsid; (void)json; (void)out; return WF_OK; -} -wf_status wf_auth_client_procedure(wf_auth_client *client, const char *nsid, - const char *json, wf_response *out) { - (void)client; (void)nsid; (void)json; (void)out; return WF_OK; -} -int main(void) { - wf_lex_com_example_inline_main_input_entries_item entries[] = {{2}, {3}}; - wf_lex_com_example_inline_main_input input = {0}; - input.config.name = "demo"; input.config.nested.enabled = true; - input.entries.items = entries; input.entries.count = 2; - char *json = NULL; - assert(wf_lex_com_example_inline_main_input_encode_json(&input, &json) == WF_OK); - assert(strcmp(json, "{\"config\":{\"name\":\"demo\",\"nested\":{\"enabled\":true}},\"entries\":[{\"count\":2},{\"count\":3}]}") == 0); - wf_lex_com_example_inline_main_json_free(json); - const char body[] = "{\"result\":{\"values\":[{\"label\":\"one\"},{\"label\":\"two\"}]}}"; - wf_lex_com_example_inline_main_output *output = NULL; - assert(wf_lex_com_example_inline_main_output_decode_json(body, strlen(body), &output) == WF_OK); - assert(output->result.values.count == 2); - assert(strcmp(output->result.values.items[1].label, "two") == 0); - wf_lex_com_example_inline_main_output_free(output); - return 0; -} -''', encoding="utf-8") - openssl = openssl_flags() - subprocess.run(["cc", "-std=c11", "-Wall", "-Wextra", "-Werror", - "-I", str(directory), "-I", str(ROOT / "include"), - "-I", str(cjson_include), str(generated), str(check), - "-L", str(cjson_lib), "-lcjson", *openssl, - "-o", str(executable)], check=True) - env = os.environ.copy() - env["DYLD_LIBRARY_PATH"] = str(cjson_lib) - env["LD_LIBRARY_PATH"] = str(cjson_lib) - subprocess.run([str(executable)], check=True, env=env) - - def test_referenced_inputs_and_all_json_value_kinds_run(self): - cjson_include = ROOT / "build" / "_deps" / "cjson-src" - cjson_lib = ROOT / "build" / "_deps" / "cjson-build" - if not (cjson_include / "cJSON.h").exists(): - self.skipTest("configured cJSON dependency is not available") - with tempfile.TemporaryDirectory() as directory: - directory = Path(directory) - header = directory / "generated.h" - generated = directory / "generated.c" - check = directory / "check.c" - executable = directory / "check" - subprocess.run([sys.executable, str(GENERATOR), str(REFS_FIXTURE), - "-o", str(header), "--source-output", str(generated)], - check=True) - check.write_text(r'''#include "generated.h" -#include -#include - -wf_status wf_xrpc_procedure(wf_xrpc_client *client, const char *nsid, - const char *json, wf_response *out) { - (void)client; (void)nsid; (void)json; (void)out; return WF_OK; -} -wf_status wf_auth_client_procedure(wf_auth_client *client, const char *nsid, - const char *json, wf_response *out) { - (void)client; (void)nsid; (void)json; (void)out; return WF_OK; -} - -int main(void) { - const wf_lex_json metadata = {"{\"n\":1}", 7}; - const wf_lex_com_example_refs_config first = {"one", metadata}; - const wf_lex_com_example_refs_config second = {"two", metadata}; - const wf_lex_com_example_refs_config *configs[] = {&first, &second}; - const char *names[] = {"alice", "bob"}; - const int64_t counts[] = {-1, 2}; - const bool switches[] = {true, false}; - const uint8_t payload[] = {1, 2, 3}; - wf_lex_com_example_refs_main_input input = {0}; - input.config = &first; - input.configs.items = configs; input.configs.count = 2; - input.names.items = names; input.names.count = 2; - input.mode = "fast"; input.token = "com.example.refs#token"; - input.counts.items = counts; input.counts.count = 2; - input.switches.items = switches; input.switches.count = 2; - input.payload.data = payload; input.payload.length = sizeof(payload); - input.link.cid = "bafy-link"; - input.blob = (wf_lex_blob){"bafy-blob", "image/png", 42}; - char *json = NULL; - assert(wf_lex_com_example_refs_main_input_encode_json(&input, &json) == WF_OK); - assert(strcmp(json, "{\"config\":{\"name\":\"one\",\"metadata\":{\"n\":1}},\"configs\":[{\"name\":\"one\",\"metadata\":{\"n\":1}},{\"name\":\"two\",\"metadata\":{\"n\":1}}],\"names\":[\"alice\",\"bob\"],\"mode\":\"fast\",\"token\":\"com.example.refs#token\",\"counts\":[-1,2],\"switches\":[true,false],\"payload\":{\"$bytes\":\"AQID\"},\"link\":{\"$link\":\"bafy-link\"},\"blob\":{\"$type\":\"blob\",\"ref\":{\"$link\":\"bafy-blob\"},\"mimeType\":\"image/png\",\"size\":42}}") == 0); - wf_lex_com_example_refs_main_json_free(json); - input.config = NULL; - assert(wf_lex_com_example_refs_main_input_encode_json(&input, &json) == WF_ERR_INVALID_ARG); - return 0; -} -''', encoding="utf-8") - openssl = openssl_flags() - subprocess.run(["cc", "-std=c11", "-Wall", "-Wextra", "-Werror", - "-I", str(directory), "-I", str(ROOT / "include"), - "-I", str(cjson_include), str(generated), str(check), - "-L", str(cjson_lib), "-lcjson", *openssl, - "-o", str(executable)], check=True) - env = os.environ.copy() - env["DYLD_LIBRARY_PATH"] = str(cjson_lib) - env["LD_LIBRARY_PATH"] = str(cjson_lib) - subprocess.run([str(executable)], check=True, env=env) - - def test_generates_json_codec_and_transport_wrapper(self): - with tempfile.TemporaryDirectory() as directory: - directory = Path(directory) - header = directory / "generated.h" - source = directory / "generated.c" - subprocess.run( - [sys.executable, str(GENERATOR), str(FIXTURE), "-o", str(header), - "--source-output", str(source)], check=True, - ) - generated = source.read_text(encoding="utf-8") - self.assertIn("cJSON_PrintUnformatted", generated) - self.assertIn("_output_decode_json", generated) - self.assertIn('cJSON_GetObjectItemCaseSensitive(item, "$bytes")', generated) - self.assertIn('cJSON_GetObjectItemCaseSensitive(item, "$link")', generated) - self.assertIn('wf_xrpc_procedure(client, "com.example.echo"', generated) - subprocess.run( - [sys.executable, str(GENERATOR), str(FIXTURE), "-o", str(header), - "--source-output", str(source)], check=True, - ) - self.assertEqual(generated, source.read_text(encoding="utf-8")) - - def test_query_wrapper_uses_encoded_xrpc_parameters(self): - with tempfile.TemporaryDirectory() as directory: - directory = Path(directory) - header = directory / "generated.h" - source = directory / "generated.c" - subprocess.run( - [sys.executable, str(GENERATOR), str(QUERY_FIXTURE), "-o", str(header), - "--source-output", str(source)], check=True, - ) - generated = source.read_text(encoding="utf-8") - self.assertIn('wf_xrpc_query_params(client, "com.example.get"', generated) - self.assertIn('encoded[count++] = (wf_xrpc_param){"limit", number_values[number_count++]}', generated) - self.assertIn('encoded[count++] = (wf_xrpc_param){"dids", params->dids.items[i]}', generated) - self.assertIn('encoded[count++] = (wf_xrpc_param){"ids", number_values[number_count++]}', generated) - self.assertIn('encoded[count++] = (wf_xrpc_param){"flags", (params->flags.items[i] ? "true" : "false")}', generated) - - def test_query_array_wrapper_runs_with_repeated_keys(self): - cjson_include = ROOT / "build" / "_deps" / "cjson-src" - cjson_lib = ROOT / "build" / "_deps" / "cjson-build" - if not (cjson_include / "cJSON.h").exists(): - self.skipTest("configured cJSON dependency is not available") - with tempfile.TemporaryDirectory() as directory: - directory = Path(directory) - header = directory / "generated.h" - generated = directory / "generated.c" - check = directory / "check.c" - executable = directory / "check" - subprocess.run( - [sys.executable, str(GENERATOR), str(QUERY_FIXTURE), "-o", str(header), - "--source-output", str(generated)], check=True, - ) - check.write_text(r'''#include "generated.h" -#include -#include - -wf_status wf_xrpc_query_params(wf_xrpc_client *client, const char *nsid, - const wf_xrpc_param *params, size_t count, - wf_response *out) { - const char *names[] = {"name", "limit", "dids", "dids", "ids", "ids", "flags", "flags"}; - const char *values[] = {"alice", "42", "did:plc:a", "did:plc:b", "-7", "9", "true", "false"}; - assert(client && out && strcmp(nsid, "com.example.get") == 0 && count == 8); - for (size_t i = 0; i < count; ++i) { - assert(strcmp(params[i].name, names[i]) == 0); - assert(strcmp(params[i].value, values[i]) == 0); - } - return WF_OK; -} -wf_status wf_auth_client_query_params(wf_auth_client *client, const char *nsid, - const wf_xrpc_param *params, size_t count, - wf_response *out) { - (void)client; (void)nsid; (void)params; (void)count; (void)out; return WF_OK; -} - -int main(void) { - const char *dids[] = {"did:plc:a", "did:plc:b"}; - int64_t ids[] = {-7, 9}; - bool flags[] = {true, false}; - wf_lex_com_example_get_main_params params = {0}; - params.name = "alice"; params.dids.items = dids; params.dids.count = 2; - params.has_limit = true; params.limit = 42; - params.has_ids = true; params.ids.items = ids; params.ids.count = 2; - params.has_flags = true; params.flags.items = flags; params.flags.count = 2; - wf_response response = {0}; - assert(wf_lex_com_example_get_main_call((wf_xrpc_client *)1, ¶ms, &response) == WF_OK); - dids[1] = NULL; - assert(wf_lex_com_example_get_main_call((wf_xrpc_client *)1, ¶ms, &response) == WF_ERR_INVALID_ARG); - dids[1] = "did:plc:b"; - params.dids.items = NULL; - assert(wf_lex_com_example_get_main_call((wf_xrpc_client *)1, ¶ms, &response) == WF_ERR_INVALID_ARG); - return 0; -} -''', encoding="utf-8") - openssl = openssl_flags() - subprocess.run([ - "cc", "-std=c11", "-Wall", "-Wextra", "-Werror", - "-I", str(directory), "-I", str(ROOT / "include"), - "-I", str(cjson_include), str(generated), str(check), - "-L", str(cjson_lib), "-lcjson", *openssl, "-o", str(executable), - ], check=True) - env = os.environ.copy() - env["DYLD_LIBRARY_PATH"] = str(cjson_lib) - env["LD_LIBRARY_PATH"] = str(cjson_lib) - subprocess.run([str(executable)], check=True, env=env) - - def test_generated_codec_and_wrapper_run(self): - cjson_include = ROOT / "build" / "_deps" / "cjson-src" - cjson_lib = ROOT / "build" / "_deps" / "cjson-build" - if not (cjson_include / "cJSON.h").exists(): - self.skipTest("configured cJSON dependency is not available") - with tempfile.TemporaryDirectory() as directory: - directory = Path(directory) - header = directory / "generated.h" - generated = directory / "generated.c" - check = directory / "check.c" - executable = directory / "check" - subprocess.run( - [sys.executable, str(GENERATOR), str(FIXTURE), "-o", str(header), - "--source-output", str(generated)], check=True, - ) - check.write_text(r'''#include "generated.h" -#include -#include -#include - -static int called; -wf_status wf_xrpc_procedure(wf_xrpc_client *client, const char *nsid, - const char *json, wf_response *out) { - assert(client != NULL); - assert(strcmp(nsid, "com.example.echo") == 0); - assert(strcmp(json, "{\"message\":\"hello\",\"attempts\":2,\"enabled\":true,\"tags\":[\"a\",\"b\"]}") == 0); - called = 1; out->status = 200; out->body = NULL; out->body_len = 0; return WF_OK; -} -wf_status wf_auth_client_procedure(wf_auth_client *client, const char *nsid, - const char *json, wf_response *out) { - (void)client; (void)nsid; (void)json; (void)out; return WF_OK; -} -int main(void) { - const char *tags[] = {"a", "b"}; - wf_lex_com_example_echo_main_input input = {0}; - input.message = "hello"; input.attempts = 2; - input.has_enabled = true; input.enabled = true; - input.has_tags = true; input.tags.items = tags; input.tags.count = 2; - char *json = NULL; - assert(wf_lex_com_example_echo_main_input_encode_json(&input, &json) == WF_OK); - assert(json != NULL); wf_lex_com_example_echo_main_json_free(json); - wf_response response = {0}; - assert(wf_lex_com_example_echo_main_call((wf_xrpc_client *)1, &input, &response) == WF_OK); - assert(called); - const char output_json[] = "{\"value\":\"ok\",\"items\":[{\"id\":\"one\",\"payload\":{\"$bytes\":\"AQID\"}}],\"raw\":{\"$bytes\":\"AAE=\"},\"link\":{\"$link\":\"bafytest\"},\"blob\":{\"$type\":\"blob\",\"ref\":{\"$link\":\"bafyblob\"},\"mimeType\":\"image/png\",\"size\":42},\"extra\":{\"x\":1},\"flags\":[true,false]}"; - wf_lex_com_example_echo_main_output *output = NULL; - assert(wf_lex_com_example_echo_main_output_decode_json(output_json, strlen(output_json), &output) == WF_OK); - assert(output && strcmp(output->value, "ok") == 0); - assert(output->items.count == 1 && strcmp(output->items.items[0]->id, "one") == 0); - assert(output->items.items[0]->payload.length == 3 && output->items.items[0]->payload.data[2] == 3); - assert(output->raw.length == 2 && output->raw.data[1] == 1); - assert(strcmp(output->link.cid, "bafytest") == 0); - assert(strcmp(output->blob.cid, "bafyblob") == 0 && output->blob.size == 42); - assert(output->has_flags && output->flags.count == 2 && output->flags.items[0]); - assert(output->extra.length == 7 && strcmp(output->extra.data, "{\"x\":1}") == 0); - wf_lex_com_example_echo_main_output_free(output); - output = (void *)1; - assert(wf_lex_com_example_echo_main_output_decode_json("{}", 2, &output) == WF_ERR_INVALID_ARG); - assert(output == NULL); - return 0; -} -''', encoding="utf-8") - openssl = openssl_flags() - subprocess.run([ - "cc", "-std=c11", "-Wall", "-Wextra", "-Werror", - "-I", str(directory), "-I", str(ROOT / "include"), - "-I", str(cjson_include), str(generated), str(check), - "-L", str(cjson_lib), "-lcjson", *openssl, "-o", str(executable), - ], check=True) - env = os.environ.copy() - env["DYLD_LIBRARY_PATH"] = str(cjson_lib) - env["LD_LIBRARY_PATH"] = str(cjson_lib) - subprocess.run([str(executable)], check=True, env=env) - - -if __name__ == "__main__": - unittest.main() diff --git a/tools/wf_gen_unspecced_wrappers.cpp b/tools/wf_gen_unspecced_wrappers.cpp new file mode 100644 index 0000000..2ee44ac --- /dev/null +++ b/tools/wf_gen_unspecced_wrappers.cpp @@ -0,0 +1,200 @@ +// C++ replacement for tools/wf_gen_unspecced_wrappers.py +// Generate XRPC-level convenience wrappers for remaining app.bsky.unspecced endpoints + +#include +#include +#include +#include +#include +#include +#include +#include + +namespace fs = std::filesystem; + +static const char* HEADER_PATH = "include/wolfram/atproto_lex.h"; + +// Endpoints that already have wrappers in unspecced_typed.c (skip these) +static const std::set EXISTING = { + "get_age_assurance_state", + "get_config", + "get_onboarding_suggested_starter_packs", + "get_onboarding_suggested_starter_packs_skeleton", + "get_suggestions_skeleton", + "get_tagged_suggestions", + "get_trending_topics", + "search_starter_packs_skeleton", +}; + +std::string snake_to_nsid_method(const std::string& snake) { + std::string result; + bool first_segment = true; + for (size_t i = 0; i < snake.size();) { + size_t end = snake.find('_', i); + if (end == std::string::npos) + end = snake.size(); + std::string part = snake.substr(i, end - i); + if (first_segment) { + result += part; + first_segment = false; + } else { + result += static_cast(std::toupper(static_cast(part[0]))); + for (size_t j = 1; j < part.size(); ++j) + result += static_cast(std::tolower(static_cast(part[j]))); + } + i = end + 1; + } + return result; +} + +std::vector read_endpoints() { + std::ifstream file(HEADER_PATH); + if (!file) { + std::cerr << "error: cannot open " << HEADER_PATH << "\n"; + std::exit(1); + } + std::string content((std::istreambuf_iterator(file)), std::istreambuf_iterator()); + + std::set found; + size_t pos = 0; + const std::string marker = "wf_status wf_lex_app_bsky_unspecced_"; + const std::string suffix = "_main_call"; + while (true) { + pos = content.find(marker, pos); + if (pos == std::string::npos) + break; + size_t name_start = pos + marker.size(); + size_t name_end = name_start; + while (name_end < content.size() && + (std::isalnum(static_cast(content[name_end])) || + content[name_end] == '_')) + ++name_end; + if (content.compare(name_end, 1, "(") == 0 && + name_end >= name_start + suffix.size() && + content.compare(name_end - suffix.size(), suffix.size(), suffix) == 0) { + std::string name = content.substr(name_start, name_end - name_start - suffix.size()); + if (EXISTING.find(name) == EXISTING.end()) + found.insert(name); + } + pos = name_end; + } + + return std::vector(found.begin(), found.end()); +} + +void output_header(const std::vector& endpoints) { + std::cout << "/* ------------------------------------------------------------------ */\n"; + std::cout << "/* Unspecced — XRPC-level convenience wrappers */\n"; + std::cout << "/* ------------------------------------------------------------------ */\n"; + std::cout << "\n"; + for (const auto& ep : endpoints) { + std::string method = snake_to_nsid_method(ep); + std::string define_name = "WF_UNSPECCED_" + [&method]() { + std::string upper = method; + for (auto& c : upper) + c = static_cast(std::toupper(static_cast(c))); + return upper; + }() + "_NSID"; + std::string nsid = "app.bsky.unspecced." + method; + std::cout << "#define " << std::left << std::setw(60) << define_name << " \"" << nsid << "\"\n"; + } + std::cout << "\n"; + for (const auto& ep : endpoints) { + std::string lex_name = "app_bsky_unspecced_" + ep; + std::string method = snake_to_nsid_method(ep); + bool has_input = ep == "init_age_assurance"; + bool has_output = ep != "init_age_assurance"; + std::string params_type = has_input ? "wf_lex_" + lex_name + "_main_input" + : "wf_lex_" + lex_name + "_main_params"; + std::string func = "wf_unspecced_" + ep; + if (has_input) { + std::cout << "wf_status " << func + << "(wf_xrpc_client *client, const " << params_type + << " *input, wf_response *out);\n"; + } else { + std::cout << "wf_status " << func + << "(wf_xrpc_client *client, const " << params_type + << " *params, wf_response *out);\n"; + } + if (has_output) { + std::string output_type = "wf_lex_" + lex_name + "_main_output"; + std::cout << "wf_status " << func << "_parse(const wf_response *resp, " + << output_type << " **out);\n"; + } + std::cout << "\n"; + } +} + +void output_source(const std::vector& endpoints) { + std::cout << "/* ================================================================== */\n"; + std::cout << "/* Unspecced — generated XRPC-level convenience wrappers */\n"; + std::cout << "/* ================================================================== */\n"; + std::cout << "\n"; + + for (const auto& ep : endpoints) { + std::string lex_name = "app_bsky_unspecced_" + ep; + bool has_input = ep == "init_age_assurance"; + bool has_output = ep != "init_age_assurance"; + std::string params_type = has_input ? "wf_lex_" + lex_name + "_main_input" + : "wf_lex_" + lex_name + "_main_params"; + std::string call_func = "wf_lex_" + lex_name + "_main_call"; + std::string func = "wf_unspecced_" + ep; + std::string arg_name = has_input ? "input" : "params"; + + std::cout << "wf_status " << func << "(\n"; + std::cout << " wf_xrpc_client *client,\n"; + std::cout << " const " << params_type << " *" << arg_name << ",\n"; + std::cout << " wf_response *out) {\n"; + std::cout << " if (!client || !" << arg_name << " || !out) {\n"; + std::cout << " return WF_ERR_INVALID_ARG;\n"; + std::cout << " }\n"; + std::cout << " return " << call_func << "(client, " << arg_name << ", out);\n"; + std::cout << "}\n"; + std::cout << "\n"; + + if (has_output) { + std::string output_type = "wf_lex_" + lex_name + "_main_output"; + std::string decode_func = "wf_lex_" + lex_name + "_main_output_decode_json"; + std::cout << "wf_status " << func << "_parse(\n"; + std::cout << " const wf_response *resp,\n"; + std::cout << " " << output_type << " **out) {\n"; + std::cout << " if (!resp || !out) {\n"; + std::cout << " return WF_ERR_INVALID_ARG;\n"; + std::cout << " }\n"; + std::cout << " *out = NULL;\n"; + std::cout << " if (!resp->body) {\n"; + std::cout << " return WF_ERR_INVALID_ARG;\n"; + std::cout << " }\n"; + std::cout << " return " << decode_func << "(\n"; + std::cout << " resp->body, resp->body_len, out);\n"; + std::cout << "}\n"; + std::cout << "\n"; + } + } +} + +int main(int argc, char* argv[]) { + std::vector endpoints = read_endpoints(); + if (endpoints.empty()) { + std::cerr << "No remaining unspecced endpoints found.\n"; + return 1; + } + + bool header_mode = false; + bool source_mode = false; + for (int i = 1; i < argc; ++i) { + std::string arg = argv[i]; + if (arg == "--header") header_mode = true; + if (arg == "--source") source_mode = true; + } + + if (header_mode) { + output_header(endpoints); + } else if (source_mode) { + output_source(endpoints); + } else { + std::cerr << "Usage: " << argv[0] << " [--header | --source]\n"; + std::cerr << "Found " << endpoints.size() << " endpoints\n"; + } + return 0; +} diff --git a/tools/wf_gen_unspecced_wrappers.py b/tools/wf_gen_unspecced_wrappers.py deleted file mode 100644 index 44e38ba..0000000 --- a/tools/wf_gen_unspecced_wrappers.py +++ /dev/null @@ -1,155 +0,0 @@ -#!/usr/bin/env python3 -"""Generate XRPC-level convenience wrappers for remaining app.bsky.unspecced endpoints. - -Usage: - python3 tools/wf_gen_unspecced_wrappers.py --header # header declarations - python3 tools/wf_gen_unspecced_wrappers.py --source # source implementations -""" - -import re -import sys - -HEADER_PATH = "include/wolfram/atproto_lex.h" - -# Endpoints that already have wrappers in unspecced_typed.c (skip these) -EXISTING = frozenset({ - "get_age_assurance_state", - "get_config", - "get_onboarding_suggested_starter_packs", - "get_onboarding_suggested_starter_packs_skeleton", - "get_suggestions_skeleton", - "get_tagged_suggestions", - "get_trending_topics", - "search_starter_packs_skeleton", -}) - - -def snake_to_nsid_method(snake): - """Convert snake_case method name to NSID method name (camelCase). - E.g. 'get_popular_feed_generators' -> 'getPopularFeedGenerators' - """ - parts = snake.split("_") - result = parts[0] - for p in parts[1:]: - result += p.capitalize() - return result - - -def read_endpoints(): - """Read all unspecced _main_call endpoint names from atproto_lex.h.""" - with open(HEADER_PATH, "r") as f: - content = f.read() - endpoints = set() - for m in re.finditer( - r'wf_status\s+wf_lex_app_bsky_unspecced_(\w+)_main_call\s*\(', content - ): - name = m.group(1) - if name not in EXISTING: - endpoints.add(name) - return sorted(endpoints) - - -def output_header(endpoints): - """Output header declarations.""" - print("/* ------------------------------------------------------------------ */") - print("/* Unspecced — XRPC-level convenience wrappers */") - print("/* ------------------------------------------------------------------ */") - print() - # NSID defines - for ep in endpoints: - method = snake_to_nsid_method(ep) - define_name = f"WF_UNSPECCED_{method.upper()}_NSID" - nsid = f"app.bsky.unspecced.{method}" - print(f"#define {define_name:60s} \"{nsid}\"") - print() - # Wrapper declarations - for ep in endpoints: - lex_name = f"app_bsky_unspecced_{ep}" - method = snake_to_nsid_method(ep) - has_input = ep == "init_age_assurance" - has_output = ep != "init_age_assurance" - if has_input: - params_type = f"wf_lex_{lex_name}_main_input" - else: - params_type = f"wf_lex_{lex_name}_main_params" - func = f"wf_unspecced_{ep}" - if has_input: - print(f"wf_status {func}(wf_xrpc_client *client, const {params_type} *input, wf_response *out);") - else: - print(f"wf_status {func}(wf_xrpc_client *client, const {params_type} *params, wf_response *out);") - if has_output: - output_type = f"wf_lex_{lex_name}_main_output" - print(f"wf_status {func}_parse(const wf_response *resp, {output_type} **out);") - print() - - -def output_source(endpoints): - """Output source implementations.""" - print("/* ================================================================== */") - print("/* Unspecced — generated XRPC-level convenience wrappers */") - print("/* ================================================================== */") - print() - - for ep in endpoints: - lex_name = f"app_bsky_unspecced_{ep}" - has_input = ep == "init_age_assurance" - has_output = ep != "init_age_assurance" - - if has_input: - params_type = f"wf_lex_{lex_name}_main_input" - call_func = f"wf_lex_{lex_name}_main_call" - else: - params_type = f"wf_lex_{lex_name}_main_params" - call_func = f"wf_lex_{lex_name}_main_call" - - func = f"wf_unspecced_{ep}" - - # Call function - arg_name = "input" if has_input else "params" - print(f"wf_status {func}(") - print(f" wf_xrpc_client *client,") - print(f" const {params_type} *{arg_name},") - print(f" wf_response *out) {{") - print(f" if (!client || !{arg_name} || !out) {{") - print(f" return WF_ERR_INVALID_ARG;") - print(f" }}") - print(f" return {call_func}(client, {arg_name}, out);") - print(f"}}") - print() - - # Parse function - if has_output: - output_type = f"wf_lex_{lex_name}_main_output" - decode_func = f"wf_lex_{lex_name}_main_output_decode_json" - print(f"wf_status {func}_parse(") - print(f" const wf_response *resp,") - print(f" {output_type} **out) {{") - print(f" if (!resp || !out) {{") - print(f" return WF_ERR_INVALID_ARG;") - print(f" }}") - print(f" *out = NULL;") - print(f" if (!resp->body) {{") - print(f" return WF_ERR_INVALID_ARG;") - print(f" }}") - print(f" return {decode_func}(") - print(f" resp->body, resp->body_len, out);") - print(f"}}") - print() - - -def main(): - endpoints = read_endpoints() - if not endpoints: - print("No remaining unspecced endpoints found.", file=sys.stderr) - sys.exit(1) - - if "--header" in sys.argv: - output_header(endpoints) - elif "--source" in sys.argv: - output_source(endpoints) - else: - print(f"Usage: {sys.argv[0]} [--header | --source]", file=sys.stderr) - print(f"Found {len(endpoints)} endpoints", file=sys.stderr) - -if __name__ == "__main__": - main() diff --git a/tools/wf_lexgen.cpp b/tools/wf_lexgen.cpp new file mode 100644 index 0000000..20447ba --- /dev/null +++ b/tools/wf_lexgen.cpp @@ -0,0 +1,1916 @@ +// wf_lexgen — C++ port of tools/wf_lexgen.py. +// +// Generate C declarations and optional implementations from Lexicon JSON. +// Generated endpoint wrappers delegate all transport to wolfram/xrpc.h. +// +// The C++ tool is the successor to the Python generator. It links against the +// vendored cJSON (target `cjson`) and mirrors the Python generator's output +// byte for byte, so regenerating the checked-in atproto_lex.h/atproto_lex.c +// with either tool yields the same files. +// +// Usage: +// wf_lexgen [-o output.h] +// [--source-output output.c] [--guard GUARD] [--header-rel H] + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +#include + +namespace fs = std::filesystem; + +// --------------------------------------------------------------------------- +// Helpers mirroring the Python generator. +// --------------------------------------------------------------------------- + +static const std::set C_KEYWORDS = { + "auto", "break", "case", "char", "const", "continue", "default", + "do", "double", "else", "enum", "extern", "float", "for", "goto", + "if", "inline", "int", "long", "register", "restrict", "return", + "short", "signed", "sizeof", "static", "struct", "switch", "typedef", + "union", "unsigned", "void", "volatile", "while", "_Alignas", + "_Alignof", "_Atomic", "_Bool", "_Complex", "_Generic", "_Imaginary", + "_Noreturn", "_Static_assert", "_Thread_local", +}; + +static const std::set CPP_KEYWORDS = { + "alignas", "alignof", "and", "and_eq", "asm", "bitand", "bitor", + "bool", "catch", "char8_t", "char16_t", "char32_t", "class", + "compl", "concept", "consteval", "constexpr", "constinit", "const_cast", + "co_await", "co_return", "co_yield", "decltype", "delete", "dynamic_cast", + "explicit", "export", "false", "friend", "mutable", "namespace", "new", + "noexcept", "not", "not_eq", "nullptr", "operator", "or", "or_eq", + "private", "protected", "public", "reflexpr", "reinterpret_cast", + "requires", "static_assert", "static_cast", "synchronized", "template", + "this", "thread_local", "throw", "true", "try", "typeid", "typename", + "using", "virtual", "wchar_t", "xor", "xor_eq", +}; + +static const std::set PY_KEYWORDS = { + "False", "None", "True", "and", "as", "assert", "async", "await", + "break", "class", "continue", "def", "del", "elif", "else", "except", + "finally", "for", "from", "global", "if", "import", "in", "is", + "lambda", "nonlocal", "not", "or", "pass", "raise", "return", "try", + "while", "with", "yield", +}; + +static bool is_lower_or_digit(char c) { + unsigned char uc = static_cast(c); + return std::islower(uc) || std::isdigit(uc); +} + +static std::string snake(const std::string& value) { + // re.sub(r"([a-z0-9])([A-Z])", r"\1_\2", value) + std::string camel; + for (size_t i = 0; i < value.size(); ++i) { + char c = value[i]; + if (i > 0 && is_lower_or_digit(value[i - 1]) && std::isupper(static_cast(c))) + camel += '_'; + camel += c; + } + // re.sub(r"[^A-Za-z0-9]+", "_", value).strip("_").lower() + std::string out; + bool pending = false; + for (char c : camel) { + if (std::isalnum(static_cast(c))) { + if (pending) { + out += '_'; + pending = false; + } + out += static_cast(std::tolower(static_cast(c))); + } else { + pending = true; + } + } + if (out.empty()) + out = "value"; + if (std::isdigit(static_cast(out[0])) || + C_KEYWORDS.count(out) || PY_KEYWORDS.count(out)) + out += '_'; + return out; +} + +static std::string member_name(const std::string& value) { + std::string s = snake(value); + if (CPP_KEYWORDS.count(s)) + s += '_'; + return s; +} + +static std::string type_name(const std::string& nsid, const std::string& suffix) { + return "wf_lex_" + snake(nsid) + (suffix.empty() ? "" : "_" + snake(suffix)); +} + +static std::string ref_type(const std::string& nsid, const std::string& ref) { + if (!ref.empty() && ref[0] == '#') + return type_name(nsid, ref.substr(1)); + size_t hash = ref.find('#'); + if (hash != std::string::npos) + return type_name(ref.substr(0, hash), ref.substr(hash + 1)); + return type_name(ref, "main"); +} + +static std::vector comment(const std::string& text, + const std::string& indent = "") { + std::string clean; + { + std::string with_slash = text; + // replace "*/" with "* /" + size_t pos = 0; + while ((pos = with_slash.find("*/", pos)) != std::string::npos) { + with_slash.replace(pos, 2, "* /"); + pos += 3; + } + std::istringstream stream(with_slash); + std::string word; + std::vector words; + while (stream >> word) + words.push_back(word); + for (size_t i = 0; i < words.size(); ++i) { + if (i) + clean += ' '; + clean += words[i]; + } + } + if (clean.empty()) + return {}; + return {indent + "/** " + clean + " */"}; +} + +// --------------------------------------------------------------------------- +// cJSON helpers. +// --------------------------------------------------------------------------- + +static std::string schema_type(cJSON* schema) { + if (!schema) + return ""; + cJSON* t = cJSON_GetObjectItemCaseSensitive(schema, "type"); + if (t && cJSON_IsString(t) && t->valuestring) + return t->valuestring; + return ""; +} + +static bool is_required(cJSON* schema, const std::string& wire) { + cJSON* req = cJSON_GetObjectItemCaseSensitive(schema, "required"); + if (!req || !cJSON_IsArray(req)) + return false; + int n = cJSON_GetArraySize(req); + for (int i = 0; i < n; ++i) { + cJSON* item = cJSON_GetArrayItem(req, i); + if (cJSON_IsString(item) && item->valuestring && wire == item->valuestring) + return true; + } + return false; +} + +static std::set required_set(cJSON* schema) { + std::set out; + cJSON* req = cJSON_GetObjectItemCaseSensitive(schema, "required"); + if (req && cJSON_IsArray(req)) { + int n = cJSON_GetArraySize(req); + for (int i = 0; i < n; ++i) { + cJSON* item = cJSON_GetArrayItem(req, i); + if (cJSON_IsString(item) && item->valuestring) + out.insert(item->valuestring); + } + } + return out; +} + +// A document is a single parsed Lexicon file. `raw` is owned by the Doc and +// freed in the destructor; `defs` borrows from `raw`. +struct Doc { + std::string id; + cJSON* raw; + cJSON* defs; + Doc() : raw(nullptr), defs(nullptr) {} + ~Doc() { + if (raw) + cJSON_Delete(raw); + } + Doc(const Doc&) = delete; + Doc& operator=(const Doc&) = delete; + Doc(Doc&& other) noexcept : id(std::move(other.id)), raw(other.raw), defs(other.defs) { + other.raw = nullptr; + other.defs = nullptr; + } + Doc& operator=(Doc&& other) noexcept { + if (this != &other) { + if (raw) + cJSON_Delete(raw); + id = std::move(other.id); + raw = other.raw; + defs = other.defs; + other.raw = nullptr; + other.defs = nullptr; + } + return *this; + } +}; + +class Generator { +public: + std::vector docs; // sorted by id + std::string guard; + + Generator(std::vector parsed, std::string g) : docs(std::move(parsed)), guard(std::move(g)) { + std::sort(docs.begin(), docs.end(), + [](const Doc& a, const Doc& b) { return a.id < b.id; }); + } + + // ----------------------------------------------------------------------- + // Schema resolution. + // ----------------------------------------------------------------------- + + // Returns (document_id, definition) for a ref, or nullopt. + std::optional> resolve_ref(const std::string& nsid, + const std::string& ref) const { + std::string document_id, fragment; + if (!ref.empty() && ref[0] == '#') { + document_id = nsid; + fragment = ref.substr(1); + } else { + size_t hash = ref.find('#'); + if (hash != std::string::npos) { + document_id = ref.substr(0, hash); + fragment = ref.substr(hash + 1); + } else { + document_id = ref; + fragment = "main"; + } + } + for (const auto& doc : docs) { + if (doc.id != document_id) + continue; + cJSON* definition = cJSON_GetObjectItemCaseSensitive(doc.defs, fragment.c_str()); + if (!definition) + return std::nullopt; + if (schema_type(definition) == "record") { + cJSON* record = cJSON_GetObjectItemCaseSensitive(definition, "record"); + if (record) + definition = record; + } + return std::make_pair(document_id, definition); + } + return std::nullopt; + } + + // ----------------------------------------------------------------------- + // Naming. + // ----------------------------------------------------------------------- + + static std::string field_name(cJSON* schema, const std::string& wire) { + std::string base = member_name(wire); + std::set presence; + cJSON* props = cJSON_GetObjectItemCaseSensitive(schema, "properties"); + if (props) { + for (cJSON* p = props->child; p; p = p->next) { + if (!p->string) + continue; + std::string name = p->string; + if (!is_required(schema, name)) + presence.insert("has_" + member_name(name)); + } + } + return presence.count(base) ? base + "_value" : base; + } + + std::pair c_type(const std::string& nsid, const std::string& owner, + const std::string& field, cJSON* schema) const { + std::string kind = schema_type(schema); + if (kind == "string") + return {"const char *", false}; + if (kind == "integer") + return {"int64_t", true}; + if (kind == "boolean") + return {"bool", true}; + if (kind == "bytes") + return {"wf_lex_bytes", false}; + if (kind == "cid-link") + return {"wf_lex_cid_link", false}; + if (kind == "blob") + return {"wf_lex_blob", false}; + if (kind == "ref") { + cJSON* ref = cJSON_GetObjectItemCaseSensitive(schema, "ref"); + if (ref && cJSON_IsString(ref) && ref->valuestring) { + auto resolved = resolve_ref(nsid, ref->valuestring); + if (resolved && schema_type(resolved->second) != "object") { + if (schema_type(resolved->second) == "token") + return {"const char *", false}; + return c_type(resolved->first, owner, field, resolved->second); + } + // Borrowed pointers support recursive and cross-document refs + // without imposing declaration order on generated headers. + return {"const " + ref_type(nsid, ref->valuestring) + " *", false}; + } + } + if (kind == "object") + return {owner + "_" + snake(field), false}; + if (kind == "union") + return {union_name(owner, field), false}; + if (kind == "unknown") + return {"wf_lex_json", false}; + if (kind == "array") { + cJSON* items = cJSON_GetObjectItemCaseSensitive(schema, "items"); + auto item = c_type(nsid, owner, field + "_item", + items ? items : cJSON_CreateObject()); // temp fallback + return {"WF_LEX_ARRAY(" + item.first + ")", false}; + } + return {"wf_lex_json", false}; + } + + std::string union_name(const std::string& owner, const std::string& field) const { + return owner + "_" + snake(field) + "_union"; + } + + // ----------------------------------------------------------------------- + // Object / union catalogs. + // ----------------------------------------------------------------------- + + // schemas(doc): named and inline schema roots for a document, in def order. + std::vector> schemas(const Doc& doc) const { + std::vector> found; + for (cJSON* def = doc.defs->child; def; def = def->next) { + if (!def->string) + continue; + std::string def_name = def->string; + std::string kind = schema_type(def); + std::string base = type_name(doc.id, def_name); + if (kind == "object") { + found.emplace_back(base, def); + } else if (kind == "record") { + cJSON* record = cJSON_GetObjectItemCaseSensitive(def, "record"); + if (record && schema_type(record) == "object") + found.emplace_back(base, record); + } else if (kind == "query" || kind == "procedure") { + cJSON* params = cJSON_GetObjectItemCaseSensitive(def, "parameters"); + if (params && schema_type(params) == "params") + found.emplace_back(base + "_params", params); + for (const char* label : {"input", "output"}) { + cJSON* value = cJSON_GetObjectItemCaseSensitive(def, label); + cJSON* schema = value ? cJSON_GetObjectItemCaseSensitive(value, "schema") : nullptr; + if (schema && schema_type(schema) == "object") + found.emplace_back(base + "_" + label, schema); + } + } + } + return found; + } + + // Ordered catalog of named and inline objects. `order` preserves the + // dependency-safe insertion order the Python generator produced. + std::map>& object_catalog() const { + if (!_objects_ready) { + build_object_catalog(); + _objects_ready = true; + } + return _objects; + } + + std::vector& object_order() const { + if (!_objects_ready) { + build_object_catalog(); + _objects_ready = true; + } + return _object_order; + } + + std::map>& collect_unions() const { + if (_unions_ready) + return _unions; + _unions.clear(); + auto& catalog = object_catalog(); + + // Local recursive walk, mirroring the Python closure. + std::function walk; + walk = [&](const std::string& nsid, const std::string& owner, cJSON* schema, + const std::string& path) { + std::string kind = schema_type(schema); + if (kind == "union") { + _unions[union_name(owner, path)] = std::make_pair(nsid, schema); + } else if (kind == "array") { + cJSON* items = cJSON_GetObjectItemCaseSensitive(schema, "items"); + walk(nsid, owner, items ? items : cJSON_CreateObject(), path + "_item"); + } else if (kind == "object") { + cJSON* props = cJSON_GetObjectItemCaseSensitive(schema, "properties"); + if (props) { + for (cJSON* p = props->child; p; p = p->next) { + if (!p->string) + continue; + walk(nsid, owner, p, field_name(schema, p->string)); + } + } + } else if (kind == "ref") { + cJSON* ref = cJSON_GetObjectItemCaseSensitive(schema, "ref"); + if (ref && cJSON_IsString(ref) && ref->valuestring) { + auto resolved = resolve_ref(nsid, ref->valuestring); + if (!resolved) + return; + cJSON* target = resolved->second; + if (schema_type(target) == "object") + return; // Borrowed pointer; walked via catalog. + cJSON walked_schema; + memset(&walked_schema, 0, sizeof(walked_schema)); + if (schema_type(target) == "token") { + cJSON_DeleteItemFromObject(&walked_schema, "type"); + } + // Mirror `walked = {"type": "string"} if token else target`. + walk(resolved->first, owner, target, path); + } + } + }; + + for (const auto& name : _object_order) { + auto it = _objects.find(name); + if (it != _objects.end()) + walk(it->second.first, name, it->second.second, ""); + } + _unions_ready = true; + return _unions; + } + + // union_members: (index, full_$type, c_type, member_name) for resolvable + // object members. + std::vector> + union_members(const std::string& nsid, cJSON* schema) const { + std::vector> members; + std::set seen; + cJSON* refs = cJSON_GetObjectItemCaseSensitive(schema, "refs"); + if (!refs || !cJSON_IsArray(refs)) + return members; + int n = cJSON_GetArraySize(refs); + for (int i = 0; i < n; ++i) { + cJSON* ref_item = cJSON_GetArrayItem(refs, i); + if (!cJSON_IsString(ref_item) || !ref_item->valuestring) + continue; + std::string ref = ref_item->valuestring; + auto resolved = resolve_ref(nsid, ref); + if (!resolved || schema_type(resolved->second) != "object") + continue; + std::string full = ref[0] == '#' ? nsid + ref : ref; + std::string frag; + size_t hash = ref.find('#'); + if (hash != std::string::npos) + frag = ref.substr(hash + 1); + else + frag = ref.substr(ref.rfind('.') + 1); + std::string mname = member_name(frag); + if (seen.count(mname)) + mname += "_" + std::to_string(i); + seen.insert(mname); + members.emplace_back(i, full, ref_type(nsid, ref), mname); + } + return members; + } + + std::set referenced_types() const { + std::set found; + std::function visit; + visit = [&](const std::string& nsid, cJSON* value) { + if (!value) + return; + if (value->type == cJSON_Object) { + if (value->type == cJSON_Object) { + cJSON* type_item = cJSON_GetObjectItemCaseSensitive(value, "type"); + cJSON* ref_item = cJSON_GetObjectItemCaseSensitive(value, "ref"); + if (type_item && cJSON_IsString(type_item) && + strcmp(type_item->valuestring, "ref") == 0 && + ref_item && cJSON_IsString(ref_item) && ref_item->valuestring) { + std::string ref = ref_item->valuestring; + auto resolved = resolve_ref(nsid, ref); + if (!resolved || schema_type(resolved->second) == "object") + found.insert(ref_type(nsid, ref)); + } + for (cJSON* child = value->child; child; child = child->next) + visit(nsid, child); + } + } else if (value->type == cJSON_Array) { + for (cJSON* child = value->child; child; child = child->next) + visit(nsid, child); + } + }; + for (const auto& doc : docs) + visit(doc.id, doc.defs); + return found; + } + + // ----------------------------------------------------------------------- + // Input type helpers. + // ----------------------------------------------------------------------- + + static std::optional json_input_type(const std::string& nsid, + const std::string& base, + cJSON* schema) { + (void)nsid; + std::string kind = schema_type(schema); + static const std::set kinds = { + "array", "blob", "boolean", "bytes", "cid-link", "integer", + "object", "ref", "string", "union", "unknown"}; + if (kinds.count(kind)) + return base + "_input"; + return std::nullopt; + } + + std::string json_input_alias_type(const std::string& nsid, const std::string& base, + cJSON* schema) const { + if (schema_type(schema) == "ref") { + cJSON* ref = cJSON_GetObjectItemCaseSensitive(schema, "ref"); + if (ref && cJSON_IsString(ref) && ref->valuestring) { + auto resolved = resolve_ref(nsid, ref->valuestring); + if (resolved && schema_type(resolved->second) == "object") + return ref_type(nsid, ref->valuestring); + } + } + return c_type(nsid, base, "input", schema).first; + } + + // ----------------------------------------------------------------------- + // Header emission. + // ----------------------------------------------------------------------- + + std::vector emit_union_struct(const std::string& name, const std::string& nsid, + cJSON* schema) const { + std::vector lines = { + "typedef struct " + name + " {", + " int kind;", + " /* Retained raw JSON (mirrors wf_lex_json) for re-encoding", + " * and open-union/unknown $type fallback. */", + " const char *data;", + " size_t length;", + " union {"}; + auto members = union_members(nsid, schema); + if (!members.empty()) { + for (const auto& m : members) + lines.push_back(" const " + std::get<2>(m) + " *" + std::get<3>(m) + ";"); + } else { + lines.push_back(" unsigned char _unused;"); + } + lines.push_back(" } value;"); + lines.push_back("} " + name + ";"); + lines.push_back(""); + return lines; + } + + std::vector emit_union_decoder(const std::string& name, const std::string& nsid, + cJSON* schema) const { + std::vector lines = { + "static wf_status wf_lex_decode_" + name + "(cJSON *node, " + name + " *value) {", + " if (!node || !value) return WF_ERR_INVALID_ARG;", + " memset(value, 0, sizeof(*value));", + " value->kind = -1;", + " char *raw = cJSON_PrintUnformatted(node);", + " if (!raw) return WF_ERR_ALLOC;", + " value->data = raw; value->length = strlen(raw);", + " cJSON *type = cJSON_GetObjectItemCaseSensitive(node, \"$type\");", + " const char *t = (type && cJSON_IsString(type)) ? type->valuestring : NULL;", + " if (t) {"}; + for (const auto& m : union_members(nsid, schema)) { + int idx = std::get<0>(m); + const std::string& full = std::get<1>(m); + const std::string& ctype = std::get<2>(m); + const std::string& mname = std::get<3>(m); + lines.push_back(" if (strcmp(t, \"" + full + "\") == 0) {"); + lines.push_back(" value->kind = " + std::to_string(idx) + ";"); + lines.push_back(" " + ctype + " *m = calloc(1, sizeof(*m));"); + lines.push_back(" if (!m) { wf_lex_clear_" + name + "(value); return WF_ERR_ALLOC; }"); + lines.push_back(" wf_status status = wf_lex_decode_" + ctype + "(node, m);"); + lines.push_back(" if (status != WF_OK) {"); + lines.push_back(" free(m); value->kind = -1;"); + lines.push_back(" } else {"); + lines.push_back(" value->value." + mname + " = m;"); + lines.push_back(" }"); + lines.push_back(" }"); + } + lines.push_back(" }"); + lines.push_back(" return WF_OK;"); + lines.push_back("}"); + lines.push_back(""); + return lines; + } + + std::vector emit_union_clear(const std::string& name, const std::string& nsid, + cJSON* schema) const { + std::vector lines = { + "static void wf_lex_clear_" + name + "(" + name + " *value) {", + " if (!value) return;", + " free((void *)value->data);", + " switch (value->kind) {"}; + for (const auto& m : union_members(nsid, schema)) { + int idx = std::get<0>(m); + const std::string& ctype = std::get<2>(m); + const std::string& mname = std::get<3>(m); + lines.push_back(" case " + std::to_string(idx) + ":"); + lines.push_back(" if (value->value." + mname + ") {"); + lines.push_back(" wf_lex_clear_" + ctype + "((" + ctype + " *)value->value." + mname + ");"); + lines.push_back(" free((void *)value->value." + mname + ");"); + lines.push_back(" }"); + lines.push_back(" break;"); + } + lines.push_back(" default: break;"); + lines.push_back(" }"); + lines.push_back(" memset(value, 0, sizeof(*value));"); + lines.push_back("}"); + lines.push_back(""); + return lines; + } + + std::vector emit_object(const std::string& nsid, const std::string& name, + cJSON* schema) const { + std::vector lines; + cJSON* desc = cJSON_GetObjectItemCaseSensitive(schema, "description"); + if (desc && cJSON_IsString(desc) && desc->valuestring) { + auto c = comment(desc->valuestring); + lines.insert(lines.end(), c.begin(), c.end()); + } + lines.push_back("typedef struct " + name + " {"); + cJSON* props = cJSON_GetObjectItemCaseSensitive(schema, "properties"); + bool any = props && props->child; + if (!any) + lines.push_back(" unsigned char _unused;"); + if (props) { + for (cJSON* p = props->child; p; p = p->next) { + if (!p->string) + continue; + std::string wire = p->string; + std::string field = field_name(schema, wire); + auto type = c_type(nsid, name, field, p); + cJSON* pdesc = cJSON_GetObjectItemCaseSensitive(p, "description"); + if (pdesc && cJSON_IsString(pdesc) && pdesc->valuestring) { + auto c = comment(pdesc->valuestring, " "); + lines.insert(lines.end(), c.begin(), c.end()); + } + if (!is_required(schema, wire)) + lines.push_back(" bool has_" + field + ";"); + lines.push_back(" " + type.first + " " + field + ";"); + } + } + lines.push_back("} " + name + ";"); + return lines; + } + + // ----------------------------------------------------------------------- + // Source emission: encoders. + // ----------------------------------------------------------------------- + + void add_encoded_value(std::vector& lines, const std::string& nsid, + const std::string& owner, cJSON* prop, const std::string& target, + const std::string& result, const std::string& path, + const std::string& indent) const { + std::string kind = schema_type(prop); + std::optional expr; + if (kind == "string") { + lines.push_back(indent + "if (!" + target + ") goto invalid;"); + expr = "cJSON_CreateString(" + target + ")"; + } else if (kind == "integer") { + expr = "cJSON_CreateNumber((double)" + target + ")"; + } else if (kind == "boolean") { + expr = "cJSON_CreateBool(" + target + ")"; + } else if (kind == "union" || kind == "unknown") { + lines.push_back(indent + "if (!" + target + ".data) goto invalid;"); + lines.push_back(indent + result + " = cJSON_ParseWithLength(" + target + + ".data, " + target + ".length);"); + lines.push_back(indent + "if (!" + result + ") goto invalid;"); + } else if (kind == "bytes") { + lines.push_back(indent + "status = wf_lex_bytes_encode(&" + target + ", &" + result + ");"); + lines.push_back(indent + "if (status != WF_OK) goto status_fail;"); + } else if (kind == "cid-link") { + lines.push_back(indent + "status = wf_lex_cid_encode(&" + target + ", &" + result + ");"); + lines.push_back(indent + "if (status != WF_OK) goto status_fail;"); + } else if (kind == "blob") { + lines.push_back(indent + "status = wf_lex_blob_encode(&" + target + ", &" + result + ");"); + lines.push_back(indent + "if (status != WF_OK) goto status_fail;"); + } else if (kind == "object") { + std::string inline_name = owner + "_" + snake(path); + lines.push_back(indent + "status = wf_lex_encode_" + inline_name + "(&" + target + ", &" + result + ");"); + lines.push_back(indent + "if (status != WF_OK) goto status_fail;"); + } else if (kind == "ref") { + cJSON* ref = cJSON_GetObjectItemCaseSensitive(prop, "ref"); + if (!ref || !cJSON_IsString(ref) || !ref->valuestring) + throw std::runtime_error("cannot encode unresolved ref in " + owner); + std::string ref_s = ref->valuestring; + auto resolved = resolve_ref(nsid, ref_s); + if (!resolved) + throw std::runtime_error("cannot encode unresolved ref " + ref_s + " in " + owner); + cJSON* target_schema = resolved->second; + if (schema_type(target_schema) == "object") { + std::string referenced = ref_type(nsid, ref_s); + lines.push_back(indent + "if (!" + target + ") goto invalid;"); + lines.push_back(indent + "status = wf_lex_encode_" + referenced + "(" + target + ", &" + result + ");"); + lines.push_back(indent + "if (status != WF_OK) goto status_fail;"); + } else { + if (schema_type(target_schema) == "token") { + cJSON temp; + memset(&temp, 0, sizeof(temp)); + temp.type = cJSON_String; + cJSON_AddItemToObject(&temp, "type", cJSON_CreateString("string")); + add_encoded_value(lines, resolved->first, owner, &temp, target, result, path, indent); + cJSON_Delete(temp.child); + } else { + add_encoded_value(lines, resolved->first, owner, target_schema, target, result, path, indent); + } + } + } else if (kind == "array") { + cJSON* items = cJSON_GetObjectItemCaseSensitive(prop, "items"); + cJSON* item_schema = items ? items : cJSON_CreateObject(); + lines.push_back(indent + "if (" + target + ".count && !" + target + ".items) goto invalid;"); + lines.push_back(indent + result + " = cJSON_CreateArray();"); + lines.push_back(indent + "if (!" + result + ") goto fail;"); + std::string suffix = snake(path); + std::string index = "i_" + suffix; + std::string element = "element_" + suffix; + lines.push_back(indent + "for (size_t " + index + " = 0; " + index + " < " + target + ".count; ++" + index + ") {"); + lines.push_back(indent + " cJSON *" + element + " = NULL;"); + add_encoded_value(lines, nsid, owner, item_schema, target + ".items[" + index + "]", + element, path + "_item", indent + " "); + lines.push_back(indent + " if (!cJSON_AddItemToArray(" + result + ", " + element + ")) {"); + lines.push_back(indent + " cJSON_Delete(" + element + "); goto fail;"); + lines.push_back(indent + " }"); + lines.push_back(indent + "}"); + } else { + throw std::runtime_error("JSON encoding is not supported for " + path + " (" + kind + ")"); + } + if (expr) { + lines.push_back(indent + result + " = " + *expr + ";"); + lines.push_back(indent + "if (!" + result + ") goto fail;"); + } + } + + void add_value(std::vector& lines, const std::string& nsid, const std::string& owner, + cJSON* prop, const std::string& field, const std::string& wire_name, + const std::string& indent = " ") const { + lines.push_back(indent + "item = NULL;"); + add_encoded_value(lines, nsid, owner, prop, "value->" + field, "item", field, indent); + lines.push_back(indent + "if (!item || !cJSON_AddItemToObject(root, \"" + wire_name + + "\", item)) { cJSON_Delete(item); item = NULL; goto fail; }"); + lines.push_back(indent + "item = NULL;"); + } + + std::vector emit_object_encoder(const std::string& nsid, const std::string& name, + cJSON* schema) const { + std::vector lines = { + "static wf_status wf_lex_encode_" + name + "(const " + name + " *value, cJSON **out) {", + " if (!value || !out) return WF_ERR_INVALID_ARG;", + " *out = NULL;", + " wf_status status = WF_OK;", + " (void)status;", + " cJSON *root = cJSON_CreateObject();", + " cJSON *item = NULL;", + " (void)item;", + " if (!root) return WF_ERR_ALLOC;"}; + cJSON* props = cJSON_GetObjectItemCaseSensitive(schema, "properties"); + if (props) { + for (cJSON* p = props->child; p; p = p->next) { + if (!p->string) + continue; + std::string wire = p->string; + std::string field = field_name(schema, wire); + if (!is_required(schema, wire)) { + lines.push_back(" if (value->has_" + field + ") {"); + add_value(lines, nsid, name, p, field, wire, " "); + lines.push_back(" }"); + } else { + add_value(lines, nsid, name, p, field, wire); + } + } + } + lines.push_back(" *out = root; return WF_OK;"); + bool has_invalid = false, has_fail = false, has_status = false; + for (const auto& l : lines) { + if (l.find("goto invalid;") != std::string::npos) + has_invalid = true; + if (l.find("goto fail;") != std::string::npos) + has_fail = true; + if (l.find("goto status_fail;") != std::string::npos) + has_status = true; + } + if (has_invalid) + lines.insert(lines.end(), {"invalid:", " cJSON_Delete(item); cJSON_Delete(root); return WF_ERR_INVALID_ARG;"}); + if (has_fail) + lines.insert(lines.end(), {"fail:", " cJSON_Delete(item); cJSON_Delete(root); return WF_ERR_ALLOC;"}); + if (has_status) + lines.insert(lines.end(), {"status_fail:", " cJSON_Delete(item); cJSON_Delete(root); return status;"}); + lines.push_back("}"); + lines.push_back(""); + return lines; + } + + // ----------------------------------------------------------------------- + // Source emission: decoders. + // ----------------------------------------------------------------------- + + void emit_decode_value(std::vector& lines, const std::string& nsid, + const std::string& owner, cJSON* schema, const std::string& target, + const std::string& source, const std::string& indent, + const std::string& field = "") const { + std::string kind = schema_type(schema); + if (kind == "string") { + lines.push_back(indent + "if (!cJSON_IsString(" + source + ")) { status = WF_ERR_INVALID_ARG; goto cleanup; }"); + lines.push_back(indent + target + " = wf_lex_strdup(" + source + "->valuestring);"); + lines.push_back(indent + "if (!" + target + ") { status = WF_ERR_ALLOC; goto cleanup; }"); + } else if (kind == "integer") { + lines.push_back(indent + "if (!wf_lex_json_integer(" + source + ", &" + target + ")) { status = WF_ERR_INVALID_ARG; goto cleanup; }"); + } else if (kind == "boolean") { + lines.push_back(indent + "if (!cJSON_IsBool(" + source + ")) { status = WF_ERR_INVALID_ARG; goto cleanup; }"); + lines.push_back(indent + target + " = cJSON_IsTrue(" + source + ");"); + } else if (kind == "union") { + std::string name = union_name(owner, field); + lines.push_back(indent + "status = wf_lex_decode_" + name + "(" + source + ", &(" + target + "));"); + lines.push_back(indent + "if (status != WF_OK) goto cleanup;"); + } else if (kind == "unknown") { + lines.push_back(indent + "status = wf_lex_json_copy(" + source + ", &" + target + ");"); + lines.push_back(indent + "if (status != WF_OK) goto cleanup;"); + } else if (kind == "bytes") { + lines.push_back(indent + "status = wf_lex_bytes_decode(" + source + ", &" + target + ");"); + lines.push_back(indent + "if (status != WF_OK) goto cleanup;"); + } else if (kind == "cid-link") { + lines.push_back(indent + "status = wf_lex_cid_decode(" + source + ", &" + target + ");"); + lines.push_back(indent + "if (status != WF_OK) goto cleanup;"); + } else if (kind == "blob") { + lines.push_back(indent + "status = wf_lex_blob_decode(" + source + ", &" + target + ");"); + lines.push_back(indent + "if (status != WF_OK) goto cleanup;"); + } else if (kind == "object") { + std::string inline_name = owner + "_" + snake(target_suffix_name(target)); + if (target_ends_with(target, "items[i]")) + inline_name = owner + "_array_item"; + lines.push_back(indent + "status = wf_lex_decode_" + inline_name + "(" + source + ", &" + target + ");"); + lines.push_back(indent + "if (status != WF_OK) goto cleanup;"); + } else if (kind == "ref") { + cJSON* ref = cJSON_GetObjectItemCaseSensitive(schema, "ref"); + if (!ref || !cJSON_IsString(ref) || !ref->valuestring) + throw std::runtime_error("cannot decode unresolved ref in " + owner); + std::string ref_s = ref->valuestring; + std::string refn = ref_type(nsid, ref_s); + auto resolved = resolve_ref(nsid, ref_s); + if (resolved && schema_type(resolved->second) != "object") { + cJSON* target_schema = resolved->second; + if (schema_type(target_schema) == "token") { + cJSON temp; + memset(&temp, 0, sizeof(temp)); + cJSON_AddItemToObject(&temp, "type", cJSON_CreateString("string")); + emit_decode_value(lines, resolved->first, owner, &temp, target, source, indent, field); + cJSON_Delete(temp.child); + } else { + emit_decode_value(lines, resolved->first, owner, target_schema, target, source, indent, field); + } + } else { + if (!object_catalog().count(refn)) + throw std::runtime_error("cannot decode unresolved ref " + ref_s + " in " + owner); + lines.push_back(indent + target + " = calloc(1, sizeof(*" + target + "));"); + lines.push_back(indent + "if (!" + target + ") { status = WF_ERR_ALLOC; goto cleanup; }"); + lines.push_back(indent + "status = wf_lex_decode_" + refn + "(" + source + ", (" + refn + " *)" + target + ");"); + lines.push_back(indent + "if (status != WF_OK) goto cleanup;"); + } + } else if (kind == "array") { + cJSON* items = cJSON_GetObjectItemCaseSensitive(schema, "items"); + cJSON* item_schema = items ? items : cJSON_CreateObject(); + std::string array_field = target_field(target); + std::string item_field = array_field + "_item"; + auto item = c_type(nsid, owner, item_field, item_schema); + std::string item_type = item.first; + lines.push_back(indent + "if (!cJSON_IsArray(" + source + ")) { status = WF_ERR_INVALID_ARG; goto cleanup; }"); + lines.push_back(indent + target + ".count = (size_t)cJSON_GetArraySize(" + source + ");"); + lines.push_back(indent + "if (" + target + ".count) {"); + lines.push_back(indent + " " + item_type + " *items = calloc(" + target + ".count, sizeof(*items));"); + lines.push_back(indent + " if (!items) { status = WF_ERR_ALLOC; goto cleanup; }"); + lines.push_back(indent + " " + target + ".items = items;"); + lines.push_back(indent + " for (size_t i = 0; i < " + target + ".count; ++i) {"); + lines.push_back(indent + " cJSON *element = cJSON_GetArrayItem(" + source + ", (int)i);"); + if (schema_type(item_schema) == "object") { + lines.push_back(indent + " status = wf_lex_decode_" + item_type + "(element, &items[i]);"); + lines.push_back(indent + " if (status != WF_OK) goto cleanup;"); + } else { + emit_decode_value(lines, nsid, owner, item_schema, "items[i]", "element", + indent + " ", item_field); + } + lines.push_back(indent + " }"); + lines.push_back(indent + "}"); + } else { + throw std::runtime_error("JSON decoding is not supported for " + owner + " (" + kind + ")"); + } + } + + void emit_clear_value(std::vector& lines, const std::string& nsid, + const std::string& owner, cJSON* schema, const std::string& target, + const std::string& indent, const std::string& field = "") const { + std::string kind = schema_type(schema); + if (kind == "string") { + lines.push_back(indent + "free((void *)" + target + ");"); + } else if (kind == "union") { + std::string union_type = union_name(owner, field); + if (target_ends_with(target, "items[i]")) + lines.push_back(indent + "wf_lex_clear_" + union_type + "((" + union_type + " *)&(" + target + "));"); + else + lines.push_back(indent + "wf_lex_clear_" + union_type + "(&(" + target + "));"); + } else if (kind == "unknown" || kind == "bytes") { + lines.push_back(indent + "free((void *)" + target + ".data);"); + } else if (kind == "cid-link") { + lines.push_back(indent + "free((void *)" + target + ".cid);"); + } else if (kind == "blob") { + lines.push_back(indent + "free((void *)" + target + ".cid);"); + lines.push_back(indent + "free((void *)" + target + ".mime_type);"); + } else if (kind == "object") { + std::string field_name2 = target_field(target); + std::string inline_name = owner + "_" + snake(field_name2); + if (target_ends_with(target, "items[i]")) + inline_name = owner + "_array_item"; + lines.push_back(indent + "wf_lex_clear_" + inline_name + "(&" + target + ");"); + } else if (kind == "ref") { + cJSON* ref = cJSON_GetObjectItemCaseSensitive(schema, "ref"); + if (!ref || !cJSON_IsString(ref) || !ref->valuestring) + throw std::runtime_error("cannot clear unresolved ref in " + owner); + std::string ref_s = ref->valuestring; + std::string refn = ref_type(nsid, ref_s); + auto resolved = resolve_ref(nsid, ref_s); + if (resolved && schema_type(resolved->second) != "object") { + cJSON* target_schema = resolved->second; + if (schema_type(target_schema) == "token") { + cJSON temp; + memset(&temp, 0, sizeof(temp)); + cJSON_AddItemToObject(&temp, "type", cJSON_CreateString("string")); + emit_clear_value(lines, resolved->first, owner, &temp, target, indent, field); + cJSON_Delete(temp.child); + } else { + emit_clear_value(lines, resolved->first, owner, target_schema, target, indent, field); + } + } else { + lines.push_back(indent + "if (" + target + ") { wf_lex_clear_" + refn + "((" + refn + + " *)" + target + "); free((void *)" + target + "); }"); + } + } else if (kind == "array") { + cJSON* items = cJSON_GetObjectItemCaseSensitive(schema, "items"); + cJSON* item = items ? items : cJSON_CreateObject(); + std::string array_field = target_field(target); + std::string item_field = array_field + "_item"; + auto item_type = c_type(nsid, owner, item_field, item).first; + lines.push_back(indent + "for (size_t i = 0; i < " + target + ".count; ++i) {"); + if (schema_type(item) == "object") { + lines.push_back(indent + " wf_lex_clear_" + item_type + "((" + item_type + " *)&" + target + ".items[i]);"); + } else { + emit_clear_value(lines, nsid, owner, item, target + ".items[i]", indent + " ", item_field); + } + lines.push_back(indent + "}"); + lines.push_back(indent + "free((void *)" + target + ".items);"); + } + } + + std::vector emit_object_decoder(const std::string& nsid, const std::string& name, + cJSON* schema) const { + std::vector lines = { + "static void wf_lex_clear_" + name + "(" + name + " *value) {", + " if (!value) return;"}; + cJSON* props = cJSON_GetObjectItemCaseSensitive(schema, "properties"); + if (props) { + for (cJSON* p = props->child; p; p = p->next) { + if (!p->string) + continue; + std::string wire = p->string; + std::string field = field_name(schema, wire); + emit_clear_value(lines, nsid, name, p, "value->" + field, " ", field); + } + } + lines.push_back(" memset(value, 0, sizeof(*value));"); + lines.push_back("}"); + lines.push_back(""); + lines.push_back("static wf_status wf_lex_decode_" + name + "(cJSON *node, " + name + " *value) {"); + lines.push_back(" wf_status status = WF_OK;"); + lines.push_back(" (void)status;"); + lines.push_back(" if (!cJSON_IsObject(node) || !value) return WF_ERR_INVALID_ARG;"); + bool has_cleanup = false; + if (props) { + for (cJSON* p = props->child; p; p = p->next) { + if (!p->string) + continue; + std::string wire = p->string; + std::string field = field_name(schema, wire); + lines.push_back(" {"); + lines.push_back(" cJSON *member = cJSON_GetObjectItemCaseSensitive(node, \"" + wire + "\");"); + std::string indent; + if (is_required(schema, wire)) { + lines.push_back(" if (!member) { status = WF_ERR_INVALID_ARG; goto cleanup; }"); + has_cleanup = true; + indent = " "; + } else { + lines.push_back(" if (member) {"); + lines.push_back(" value->has_" + field + " = true;"); + indent = " "; + } + emit_decode_value(lines, nsid, name, p, "value->" + field, "member", indent, field); + if (!is_required(schema, wire)) + lines.push_back(" }"); + lines.push_back(" }"); + } + } + lines.push_back(" return WF_OK;"); + for (const auto& l : lines) + if (l.find("goto cleanup;") != std::string::npos) + has_cleanup = true; + if (has_cleanup) { + lines.push_back("cleanup:"); + lines.push_back(" wf_lex_clear_" + name + "(value);"); + lines.push_back(" return status;"); + } + lines.push_back("}"); + lines.push_back(""); + return lines; + } + + // ----------------------------------------------------------------------- + // Encoder catalog for source generation. + // ----------------------------------------------------------------------- + + // encoder_objects: subset of the object catalog reachable from endpoint + // inputs, in catalog insertion order. + std::vector>> encoder_objects() { + auto& catalog = object_catalog(); + auto& order = object_order(); + std::set wanted; + + std::function visit_object; + std::function visit_schema; + + visit_object = [&](const std::string& nsid, const std::string& name, cJSON* schema) { + if (wanted.count(name)) + return; + wanted.insert(name); + cJSON* props = cJSON_GetObjectItemCaseSensitive(schema, "properties"); + if (props) { + for (cJSON* p = props->child; p; p = p->next) { + if (!p->string) + continue; + visit_schema(nsid, name, p, field_name(schema, p->string)); + } + } + }; + + visit_schema = [&](const std::string& nsid, const std::string& owner, cJSON* schema, + const std::string& path) { + std::string kind = schema_type(schema); + if (kind == "array") { + cJSON* items = cJSON_GetObjectItemCaseSensitive(schema, "items"); + visit_schema(nsid, owner, items ? items : cJSON_CreateObject(), path + "_item"); + } else if (kind == "object") { + std::string child = owner + "_" + snake(path); + visit_object(nsid, child, catalog.at(child).second); + } else if (kind == "ref") { + cJSON* ref = cJSON_GetObjectItemCaseSensitive(schema, "ref"); + if (!ref || !cJSON_IsString(ref) || !ref->valuestring) + throw std::runtime_error("cannot encode unresolved ref in " + owner); + std::string ref_s = ref->valuestring; + auto resolved = resolve_ref(nsid, ref_s); + if (!resolved) + throw std::runtime_error("cannot encode unresolved ref " + ref_s + " in " + owner); + cJSON* target = resolved->second; + if (schema_type(target) == "object") { + std::string target_name = ref_type(nsid, ref_s); + if (!catalog.count(target_name)) + throw std::runtime_error("cannot encode ref " + ref_s + " in " + owner); + visit_object(resolved->first, target_name, catalog.at(target_name).second); + } else if (schema_type(target) != "token") { + visit_schema(resolved->first, owner, target, path); + } + } + }; + + for (const auto& doc : docs) { + for (cJSON* def = doc.defs->child; def; def = def->next) { + if (!def->string) + continue; + std::string kind = schema_type(def); + if (kind != "query" && kind != "procedure") + continue; + cJSON* input = cJSON_GetObjectItemCaseSensitive(def, "input"); + cJSON* schema = input ? cJSON_GetObjectItemCaseSensitive(input, "schema") : nullptr; + if (!schema) + schema = cJSON_CreateObject(); + if (schema_type(schema) == "object") { + std::string root = type_name(doc.id, def->string) + "_input"; + visit_object(doc.id, root, schema); + } else if (schema->type == cJSON_Object && schema_type(schema) != "") { + visit_schema(doc.id, type_name(doc.id, def->string), schema, "input"); + } + } + } + + std::vector>> result; + for (const auto& name : order) + if (wanted.count(name)) + result.emplace_back(name, catalog.at(name)); + return result; + } + + // ----------------------------------------------------------------------- + // Top-level generation. + // ----------------------------------------------------------------------- + + std::string generate() { + std::vector out = { + "/* Generated by tools/wf_lexgen.cpp; do not edit. */", + "#ifndef " + guard, + "#define " + guard, + "", + "#include ", + "#include ", + "#include ", + "#include ", + "#include ", + "", + "#ifdef __cplusplus", + "extern \"C\" {", + "#endif", + "", + "/** Encoded JSON view. Decoded outputs own data; input values borrow it. */", + "typedef struct wf_lex_json { const char *data; size_t length; } wf_lex_json;", + "/** Byte sequence view. Decoded outputs own data; input values borrow it. */", + "typedef struct wf_lex_bytes { const uint8_t *data; size_t length; } wf_lex_bytes;", + "/** CID link string view. */", + "typedef struct wf_lex_cid_link { const char *cid; } wf_lex_cid_link;", + "/** Typed AT Protocol blob reference. */", + "typedef struct wf_lex_blob { const char *cid; const char *mime_type; int64_t size; } wf_lex_blob;", + "#define WF_LEX_ARRAY(type_) struct { type_ const *items; size_t count; }", + "", + }; + + for (const auto& doc : docs) { + const std::string& nsid = doc.id; + cJSON* main = cJSON_GetObjectItemCaseSensitive(doc.defs, "main"); + std::string kind = main ? schema_type(main) : "definition"; + if (kind.empty()) + kind = "definition"; + std::string symbol = snake(nsid); + std::string upper; + for (char c : symbol) + upper += static_cast(std::toupper(static_cast(c))); + if (main) { + cJSON* desc = cJSON_GetObjectItemCaseSensitive(main, "description"); + if (desc && cJSON_IsString(desc) && desc->valuestring) { + auto c = comment(desc->valuestring); + out.insert(out.end(), c.begin(), c.end()); + } + } + out.push_back("#define WF_LEX_" + upper + "_NSID \"" + nsid + "\""); + out.push_back("#define WF_LEX_" + upper + "_KIND \"" + kind + "\""); + out.push_back(""); + } + + auto& objects = object_catalog(); + auto& unions = collect_unions(); + std::set declarations; + for (const auto& kv : objects) + declarations.insert(kv.first); + auto refs = referenced_types(); + declarations.insert(refs.begin(), refs.end()); + for (const auto& kv : unions) + declarations.insert(kv.first); + for (const auto& name : declarations) { + out.push_back("typedef struct " + name + " " + name + ";"); + } + if (!declarations.empty()) + out.push_back(""); + + for (const auto& kv : unions) { + auto u = emit_union_struct(kv.first, kv.second.first, kv.second.second); + out.insert(out.end(), u.begin(), u.end()); + } + for (const auto& name : object_order()) { + auto it = objects.find(name); + if (it == objects.end()) + continue; + auto o = emit_object(it->second.first, name, it->second.second); + out.insert(out.end(), o.begin(), o.end()); + out.push_back(""); + } + + for (const auto& doc : docs) { + const std::string& nsid = doc.id; + for (cJSON* def = doc.defs->child; def; def = def->next) { + if (!def->string) + continue; + std::string kind = schema_type(def); + if (kind != "query" && kind != "procedure") + continue; + cJSON* input = cJSON_GetObjectItemCaseSensitive(def, "input"); + cJSON* schema = input ? cJSON_GetObjectItemCaseSensitive(input, "schema") : nullptr; + if (!schema || schema_type(schema) == "object") + continue; + std::string base = type_name(nsid, def->string); + auto it = json_input_type(nsid, base, schema); + if (it) { + std::string alias = json_input_alias_type(nsid, base, schema); + out.push_back("typedef " + alias + " " + base + "_input;"); + out.push_back(""); + } + } + } + + for (const auto& doc : docs) { + const std::string& nsid = doc.id; + for (cJSON* def = doc.defs->child; def; def = def->next) { + if (!def->string) + continue; + std::string kind = schema_type(def); + if (kind != "query" && kind != "procedure") + continue; + std::string base = type_name(nsid, def->string); + cJSON* input = cJSON_GetObjectItemCaseSensitive(def, "input"); + cJSON* input_schema = input ? cJSON_GetObjectItemCaseSensitive(input, "schema") : nullptr; + auto input_type = input_schema ? json_input_type(nsid, base, input_schema) : std::optional(); + if (input_type) { + out.push_back("wf_status " + base + "_input_encode_json("); + out.push_back(" const " + *input_type + " *value, char **out_json);"); + out.push_back("/** Free JSON returned by the matching encoder. */"); + out.push_back("void " + base + "_json_free(char *json);"); + } + cJSON* output = cJSON_GetObjectItemCaseSensitive(def, "output"); + cJSON* output_schema = output ? cJSON_GetObjectItemCaseSensitive(output, "schema") : nullptr; + if (output_schema && schema_type(output_schema) == "object") { + out.push_back("/** Decode an owning output value; free it with the matching function. */"); + out.push_back("wf_status " + base + "_output_decode_json("); + out.push_back(" const char *json, size_t length, " + base + "_output **out_value);"); + out.push_back("void " + base + "_output_free(" + base + "_output *value);"); + } + if (kind == "procedure") { + if (input_type) { + out.push_back("wf_status " + base + "_call(wf_xrpc_client *client,"); + out.push_back(" const " + *input_type + " *input, wf_response *out);"); + out.push_back("wf_status " + base + "_call_auth(wf_auth_client *client,"); + out.push_back(" const " + *input_type + " *input, wf_response *out);"); + } else { + out.push_back("wf_status " + base + "_call(wf_xrpc_client *client, wf_response *out);"); + out.push_back("wf_status " + base + "_call_auth(wf_auth_client *client, wf_response *out);"); + } + } else if (kind == "query") { + cJSON* params = cJSON_GetObjectItemCaseSensitive(def, "parameters"); + bool has_params = params && schema_type(params) == "params"; + std::string arg = has_params ? "const " + base + "_params *params, " : ""; + out.push_back("wf_status " + base + "_call(wf_xrpc_client *client,"); + out.push_back(" " + arg + "wf_response *out);"); + out.push_back("wf_status " + base + "_call_auth(wf_auth_client *client,"); + out.push_back(" " + arg + "wf_response *out);"); + } + out.push_back(""); + } + } + + out.push_back("#undef WF_LEX_ARRAY"); + out.push_back(""); + out.push_back("#ifdef __cplusplus"); + out.push_back("}"); + out.push_back("#endif"); + out.push_back(""); + out.push_back("#endif /* " + guard + " */"); + out.push_back(""); + return join_lines(out); + } + + std::string generate_source(const std::string& header_name) { + std::vector out = { + "/* Generated by tools/wf_lexgen.cpp; do not edit. */", + "#include \"" + header_name + "\"", + "#include ", + "#include ", + "#include ", + "#include ", + "#include ", + "#include ", + "#include ", + "#include ", + "", + "#if defined(__GNUC__) || defined(__clang__)", + "#define WF_LEX_UNUSED __attribute__((unused))", + "#else", + "#define WF_LEX_UNUSED", + "#endif", + "", + "static WF_LEX_UNUSED char *wf_lex_strdup(const char *source) {", + " size_t length = strlen(source) + 1; char *copy = malloc(length);", + " if (copy) memcpy(copy, source, length);", + " return copy;", + "}", + "", + "static WF_LEX_UNUSED bool wf_lex_json_integer(cJSON *item, int64_t *out) {", + " if (!cJSON_IsNumber(item) || !isfinite(item->valuedouble) ||", + " item->valuedouble < -9007199254740991.0 || item->valuedouble > 9007199254740991.0 ||", + " (double)(int64_t)item->valuedouble != item->valuedouble) return false;", + " *out = (int64_t)item->valuedouble; return true;", + "}", + "", + "static WF_LEX_UNUSED wf_status wf_lex_json_copy(cJSON *item, wf_lex_json *out) {", + " char *json = cJSON_PrintUnformatted(item); if (!json) return WF_ERR_ALLOC;", + " out->data = json; out->length = strlen(json); return WF_OK;", + "}", + "", + "static WF_LEX_UNUSED wf_status wf_lex_bytes_encode(const wf_lex_bytes *value, cJSON **out) {", + " if (!value || !out || (value->length && !value->data) || value->length > INT_MAX)", + " return WF_ERR_INVALID_ARG;", + " *out = NULL;", + " if (value->length > (SIZE_MAX / 4) * 3 - 2) return WF_ERR_INVALID_ARG;", + " size_t length = 4 * ((value->length + 2) / 3);", + " char *encoded = malloc(length + 1); if (!encoded) return WF_ERR_ALLOC;", + " const uint8_t *data = value->length ? value->data : (const uint8_t *)\"\";", + " int written = EVP_EncodeBlock((unsigned char *)encoded, data, (int)value->length);", + " if (written < 0 || (size_t)written != length) { free(encoded); return WF_ERR_INVALID_ARG; }", + " encoded[length] = '\\0'; cJSON *root = cJSON_CreateObject();", + " cJSON *tag = cJSON_CreateString(encoded); free(encoded);", + " if (!root || !tag || !cJSON_AddItemToObject(root, \"$bytes\", tag)) {", + " cJSON_Delete(tag); cJSON_Delete(root); return WF_ERR_ALLOC;", + " }", + " *out = root; return WF_OK;", + "}", + "", + "static WF_LEX_UNUSED wf_status wf_lex_cid_encode(const wf_lex_cid_link *value, cJSON **out) {", + " if (!value || !out || !value->cid || !value->cid[0]) return WF_ERR_INVALID_ARG;", + " *out = NULL;", + " cJSON *root = cJSON_CreateObject(); cJSON *link = cJSON_CreateString(value->cid);", + " if (!root || !link || !cJSON_AddItemToObject(root, \"$link\", link)) {", + " cJSON_Delete(link); cJSON_Delete(root); return WF_ERR_ALLOC;", + " }", + " *out = root; return WF_OK;", + "}", + "", + "static WF_LEX_UNUSED wf_status wf_lex_blob_encode(const wf_lex_blob *value, cJSON **out) {", + " if (!value || !out || !value->cid || !value->cid[0] || !value->mime_type || value->size < 0)", + " return WF_ERR_INVALID_ARG;", + " *out = NULL;", + " cJSON *root = cJSON_CreateObject(); cJSON *ref = NULL;", + " wf_lex_cid_link link = {value->cid};", + " if (!root) return WF_ERR_ALLOC;", + " wf_status status = wf_lex_cid_encode(&link, &ref);", + " if (status != WF_OK) { cJSON_Delete(root); return status; }", + " cJSON *type = cJSON_CreateString(\"blob\");", + " cJSON *mime = cJSON_CreateString(value->mime_type);", + " cJSON *size = cJSON_CreateNumber((double)value->size);", + " if (!type || !mime || !size) {", + " cJSON_Delete(type); cJSON_Delete(ref); cJSON_Delete(mime); cJSON_Delete(size);", + " cJSON_Delete(root); return WF_ERR_ALLOC;", + " }", + " if (!cJSON_AddItemToObject(root, \"$type\", type)) goto blob_fail;", + " type = NULL;", + " if (!cJSON_AddItemToObject(root, \"ref\", ref)) goto blob_fail;", + " ref = NULL;", + " if (!cJSON_AddItemToObject(root, \"mimeType\", mime)) goto blob_fail;", + " mime = NULL;", + " if (!cJSON_AddItemToObject(root, \"size\", size)) goto blob_fail;", + " size = NULL;", + " *out = root; return WF_OK;", + "blob_fail:", + " cJSON_Delete(type); cJSON_Delete(ref); cJSON_Delete(mime); cJSON_Delete(size);", + " cJSON_Delete(root); return WF_ERR_ALLOC;", + "}", + "", + "static WF_LEX_UNUSED wf_status wf_lex_cid_decode(cJSON *item, wf_lex_cid_link *out) {", + " cJSON *link = cJSON_IsObject(item) ? cJSON_GetObjectItemCaseSensitive(item, \"$link\") : NULL;", + " if (!cJSON_IsString(link) || !link->valuestring[0]) return WF_ERR_INVALID_ARG;", + " out->cid = wf_lex_strdup(link->valuestring); return out->cid ? WF_OK : WF_ERR_ALLOC;", + "}", + "", + "static WF_LEX_UNUSED wf_status wf_lex_bytes_decode(cJSON *item, wf_lex_bytes *out) {", + " cJSON *tag = cJSON_IsObject(item) ? cJSON_GetObjectItemCaseSensitive(item, \"$bytes\") : NULL;", + " if (!cJSON_IsString(tag)) return WF_ERR_INVALID_ARG;", + " size_t encoded = strlen(tag->valuestring);", + " if (encoded % 4 != 0 || encoded > (size_t)INT_MAX) return WF_ERR_INVALID_ARG;", + " size_t capacity = encoded / 4 * 3; uint8_t *data = capacity ? malloc(capacity) : NULL;", + " if (capacity && !data) return WF_ERR_ALLOC;", + " int decoded = encoded ? EVP_DecodeBlock(data, (const unsigned char *)tag->valuestring, (int)encoded) : 0;", + " if (decoded < 0) { free(data); return WF_ERR_INVALID_ARG; }", + " size_t padding = encoded && tag->valuestring[encoded - 1] == '=';", + " padding += encoded > 1 && tag->valuestring[encoded - 2] == '=';", + " out->data = data; out->length = (size_t)decoded - padding; return WF_OK;", + "}", + "", + "static WF_LEX_UNUSED wf_status wf_lex_blob_decode(cJSON *item, wf_lex_blob *out) {", + " if (!cJSON_IsObject(item)) return WF_ERR_INVALID_ARG;", + " cJSON *type = cJSON_GetObjectItemCaseSensitive(item, \"$type\");", + " cJSON *ref = cJSON_GetObjectItemCaseSensitive(item, \"ref\");", + " cJSON *mime = cJSON_GetObjectItemCaseSensitive(item, \"mimeType\");", + " cJSON *size = cJSON_GetObjectItemCaseSensitive(item, \"size\"); wf_lex_cid_link link = {0};", + " if (!cJSON_IsString(type) || strcmp(type->valuestring, \"blob\") != 0 ||", + " !cJSON_IsString(mime) || !wf_lex_json_integer(size, &out->size)) return WF_ERR_INVALID_ARG;", + " wf_status status = wf_lex_cid_decode(ref, &link); if (status != WF_OK) return status;", + " out->mime_type = wf_lex_strdup(mime->valuestring);", + " if (!out->mime_type) { free((void *)link.cid); return WF_ERR_ALLOC; }", + " out->cid = link.cid; return WF_OK;", + "}", + "", + }; + + auto& catalog = object_catalog(); + auto encoders = encoder_objects(); + auto& unions = collect_unions(); + + for (const auto& enc : encoders) + out.push_back("static WF_LEX_UNUSED wf_status wf_lex_encode_" + enc.first + + "(const " + enc.first + " *value, cJSON **out);"); + for (const auto& kv : catalog) { + out.push_back("static WF_LEX_UNUSED void wf_lex_clear_" + kv.first + "(" + kv.first + " *value);"); + out.push_back("static WF_LEX_UNUSED wf_status wf_lex_decode_" + kv.first + "(cJSON *node, " + kv.first + " *value);"); + } + for (const auto& kv : unions) { + out.push_back("static WF_LEX_UNUSED void wf_lex_clear_" + kv.first + "(" + kv.first + " *value);"); + out.push_back("static WF_LEX_UNUSED wf_status wf_lex_decode_" + kv.first + "(cJSON *node, " + kv.first + " *value);"); + } + if (!catalog.empty() || !unions.empty()) + out.push_back(""); + + for (const auto& enc : encoders) { + auto e = emit_object_encoder(enc.second.first, enc.first, enc.second.second); + out.insert(out.end(), e.begin(), e.end()); + } + for (const auto& kv : catalog) { + auto d = emit_object_decoder(kv.second.first, kv.first, kv.second.second); + out.insert(out.end(), d.begin(), d.end()); + } + for (const auto& kv : unions) { + auto u = emit_union_decoder(kv.first, kv.second.first, kv.second.second); + out.insert(out.end(), u.begin(), u.end()); + auto uc = emit_union_clear(kv.first, kv.second.first, kv.second.second); + out.insert(out.end(), uc.begin(), uc.end()); + } + + for (const auto& doc : docs) { + const std::string& nsid = doc.id; + for (cJSON* def = doc.defs->child; def; def = def->next) { + if (!def->string) + continue; + std::string kind = schema_type(def); + if (kind != "query" && kind != "procedure") + continue; + std::string base = type_name(nsid, def->string); + cJSON* output = cJSON_GetObjectItemCaseSensitive(def, "output"); + cJSON* output_schema = output ? cJSON_GetObjectItemCaseSensitive(output, "schema") : nullptr; + if (output_schema && schema_type(output_schema) == "object") { + out.push_back("wf_status " + base + "_output_decode_json("); + out.push_back(" const char *json, size_t length, " + base + "_output **out_value) {"); + out.push_back(" if (!json || !out_value) return WF_ERR_INVALID_ARG;"); + out.push_back(" *out_value = NULL; cJSON *root = cJSON_ParseWithLength(json, length);"); + out.push_back(" if (!root) return WF_ERR_INVALID_ARG;"); + out.push_back(" " + base + "_output *value = calloc(1, sizeof(*value));"); + out.push_back(" if (!value) { cJSON_Delete(root); return WF_ERR_ALLOC; }"); + out.push_back(" wf_status status = wf_lex_decode_" + base + "_output(root, value);"); + out.push_back(" cJSON_Delete(root);"); + out.push_back(" if (status != WF_OK) { free(value); return status; }"); + out.push_back(" *out_value = value; return WF_OK;"); + out.push_back("}"); + out.push_back(""); + out.push_back("void " + base + "_output_free(" + base + "_output *value) {"); + out.push_back(" wf_lex_clear_" + base + "_output(value); free(value);"); + out.push_back("}"); + out.push_back(""); + } + cJSON* input = cJSON_GetObjectItemCaseSensitive(def, "input"); + cJSON* schema = input ? cJSON_GetObjectItemCaseSensitive(input, "schema") : nullptr; + auto input_type = schema ? json_input_type(nsid, base, schema) : std::optional(); + if (input_type) { + std::string encoder; + if (schema_type(schema) == "object") { + encoder = base + "_input"; + } else { + cJSON* ref = cJSON_GetObjectItemCaseSensitive(schema, "ref"); + if (ref && cJSON_IsString(ref) && ref->valuestring) { + auto resolved = resolve_ref(nsid, ref->valuestring); + if (resolved && schema_type(resolved->second) == "object") + encoder = ref_type(nsid, ref->valuestring); + } + } + std::vector function = { + "wf_status " + base + "_input_encode_json(", + " const " + *input_type + " *value, char **out_json) {", + " if (!value || !out_json) return WF_ERR_INVALID_ARG;", + " *out_json = NULL; cJSON *root = NULL;", + " wf_status status = WF_OK;", + " (void)status;"}; + if (!encoder.empty()) { + function.push_back(" status = wf_lex_encode_" + encoder + "(value, &root);"); + function.push_back(" if (status != WF_OK) return status;"); + } else { + add_encoded_value(function, nsid, base, schema, "*value", "root", "input", " "); + } + function.push_back(" *out_json = cJSON_PrintUnformatted(root);"); + function.push_back(" cJSON_Delete(root);"); + function.push_back(" return *out_json ? WF_OK : WF_ERR_ALLOC;"); + bool has_invalid = false, has_fail = false, has_status = false; + for (const auto& l : function) { + if (l.find("goto invalid;") != std::string::npos) + has_invalid = true; + if (l.find("goto fail;") != std::string::npos) + has_fail = true; + if (l.find("goto status_fail;") != std::string::npos) + has_status = true; + } + if (has_invalid) { + function.push_back("invalid:"); + function.push_back(" cJSON_Delete(root); return WF_ERR_INVALID_ARG;"); + } + if (has_fail) { + function.push_back("fail:"); + function.push_back(" cJSON_Delete(root); return WF_ERR_ALLOC;"); + } + if (has_status) { + function.push_back("status_fail:"); + function.push_back(" cJSON_Delete(root); return status;"); + } + out.insert(out.end(), function.begin(), function.end()); + out.push_back("}"); + out.push_back(""); + out.push_back("void " + base + "_json_free(char *json) { cJSON_free(json); }"); + out.push_back(""); + } + if (kind == "procedure" && !def_has_params(def)) { + if (input_type) { + out.push_back("wf_status " + base + "_call(wf_xrpc_client *client,"); + out.push_back(" const " + *input_type + " *input, wf_response *response) {"); + out.push_back(" if (!client || !input || !response) return WF_ERR_INVALID_ARG;"); + out.push_back(" char *json = NULL;"); + out.push_back(" wf_status status = " + base + "_input_encode_json(input, &json);"); + out.push_back(" if (status != WF_OK) return status;"); + out.push_back(" status = wf_xrpc_procedure(client, \"" + nsid + "\", json, response);"); + out.push_back(" cJSON_free(json);"); + out.push_back(" return status;"); + out.push_back("}"); + out.push_back(""); + out.push_back("wf_status " + base + "_call_auth(wf_auth_client *client,"); + out.push_back(" const " + *input_type + " *input, wf_response *response) {"); + out.push_back(" if (!client || !input || !response) return WF_ERR_INVALID_ARG;"); + out.push_back(" char *json = NULL;"); + out.push_back(" wf_status status = " + base + "_input_encode_json(input, &json);"); + out.push_back(" if (status != WF_OK) return status;"); + out.push_back(" status = wf_auth_client_procedure(client, \"" + nsid + "\", json, response);"); + out.push_back(" cJSON_free(json);"); + out.push_back(" return status;"); + out.push_back("}"); + out.push_back(""); + } else { + out.push_back("wf_status " + base + "_call(wf_xrpc_client *client, wf_response *response) {"); + out.push_back(" if (!client || !response) return WF_ERR_INVALID_ARG;"); + out.push_back(" return wf_xrpc_procedure(client, \"" + nsid + "\", NULL, response);"); + out.push_back("}"); + out.push_back(""); + out.push_back("wf_status " + base + "_call_auth(wf_auth_client *client, wf_response *response) {"); + out.push_back(" if (!client || !response) return WF_ERR_INVALID_ARG;"); + out.push_back(" return wf_auth_client_procedure(client, \"" + nsid + "\", NULL, response);"); + out.push_back("}"); + out.push_back(""); + } + } + } + + cJSON* main = cJSON_GetObjectItemCaseSensitive(doc.defs, "main"); + if (main && schema_type(main) == "query") { + std::string base = type_name(nsid, "main"); + cJSON* params = cJSON_GetObjectItemCaseSensitive(main, "parameters"); + if (!params) { + out.push_back("wf_status " + base + "_call(wf_xrpc_client *client, wf_response *response) {"); + out.push_back(" if (!client || !response) return WF_ERR_INVALID_ARG;"); + out.push_back(" return wf_xrpc_query(client, \"" + nsid + "\", NULL, response);"); + out.push_back("}"); + out.push_back(""); + out.push_back("wf_status " + base + "_call_auth(wf_auth_client *client, wf_response *response) {"); + out.push_back(" if (!client || !response) return WF_ERR_INVALID_ARG;"); + out.push_back(" return wf_auth_client_query(client, \"" + nsid + "\", NULL, response);"); + out.push_back("}"); + out.push_back(""); + } else if (schema_type(params) == "params") { + cJSON* props = cJSON_GetObjectItemCaseSensitive(params, "properties"); + // Validate parameter kinds. + std::set required = required_set(params); + if (props) { + for (cJSON* p = props->child; p; p = p->next) { + if (!p->string) + continue; + std::string wire = p->string; + std::string kind = schema_type(p); + if (kind == "array") { + cJSON* items = cJSON_GetObjectItemCaseSensitive(p, "items"); + std::string item_kind = schema_type(items); + if (item_kind != "string" && item_kind != "integer" && item_kind != "boolean") + throw std::runtime_error("query parameter " + wire + + " has unsupported array item type " + item_kind); + } else if (kind != "string" && kind != "integer" && kind != "boolean") { + throw std::runtime_error("query parameter " + wire + " has unsupported type " + kind); + } + } + } + std::vector call_body = { + " if (!params || !response) return WF_ERR_INVALID_ARG;", + " size_t encoded_capacity = 0, number_capacity = 0;"}; + if (props) { + for (cJSON* p = props->child; p; p = p->next) { + if (!p->string) + continue; + std::string wire = p->string; + std::string field = field_name(params, wire); + std::optional condition; + if (!required.count(wire)) + condition = "params->has_" + field; + std::string indent2 = " "; + if (condition) { + call_body.push_back(" if (" + *condition + ") {"); + indent2 = " "; + } + std::string kind = schema_type(p); + if (kind == "array") { + cJSON* items = cJSON_GetObjectItemCaseSensitive(p, "items"); + std::string item_kind = schema_type(items); + call_body.push_back(indent2 + "if (params->" + field + ".count && !params->" + field + ".items) return WF_ERR_INVALID_ARG;"); + if (item_kind == "string") { + call_body.push_back(indent2 + "for (size_t i = 0; i < params->" + field + ".count; ++i)"); + call_body.push_back(indent2 + " if (!params->" + field + ".items[i]) return WF_ERR_INVALID_ARG;"); + } + call_body.push_back(indent2 + "if (params->" + field + ".count > SIZE_MAX - encoded_capacity) return WF_ERR_INVALID_ARG;"); + call_body.push_back(indent2 + "encoded_capacity += params->" + field + ".count;"); + if (item_kind == "integer") { + call_body.push_back(indent2 + "if (params->" + field + ".count > SIZE_MAX - number_capacity) return WF_ERR_INVALID_ARG;"); + call_body.push_back(indent2 + "number_capacity += params->" + field + ".count;"); + } + } else { + if (kind == "string") + call_body.push_back(indent2 + "if (!params->" + field + ") return WF_ERR_INVALID_ARG;"); + call_body.push_back(indent2 + "if (encoded_capacity == SIZE_MAX) return WF_ERR_INVALID_ARG;"); + call_body.push_back(indent2 + "++encoded_capacity;"); + if (kind == "integer") { + call_body.push_back(indent2 + "if (number_capacity == SIZE_MAX) return WF_ERR_INVALID_ARG;"); + call_body.push_back(indent2 + "++number_capacity;"); + } + } + if (condition) + call_body.push_back(" }"); + } + } + call_body.push_back(" if (encoded_capacity > SIZE_MAX / sizeof(wf_xrpc_param) ||"); + call_body.push_back(" number_capacity > SIZE_MAX / sizeof(char[32])) return WF_ERR_INVALID_ARG;"); + call_body.push_back(" wf_xrpc_param *encoded = encoded_capacity ? calloc(encoded_capacity, sizeof(*encoded)) : NULL;"); + call_body.push_back(" char (*number_values)[32] = number_capacity ? malloc(number_capacity * sizeof(*number_values)) : NULL;"); + call_body.push_back(" if ((encoded_capacity && !encoded) || (number_capacity && !number_values)) {"); + call_body.push_back(" free(encoded); free(number_values); return WF_ERR_ALLOC;"); + call_body.push_back(" }"); + call_body.push_back(" size_t count = 0, number_count = 0;"); + call_body.push_back(" (void)number_count;"); + if (props) { + for (cJSON* p = props->child; p; p = p->next) { + if (!p->string) + continue; + std::string wire = p->string; + std::string field = field_name(params, wire); + std::optional condition; + if (!required.count(wire)) + condition = "params->has_" + field; + std::string indent2 = " "; + if (condition) { + call_body.push_back(" if (" + *condition + ") {"); + indent2 = " "; + } + std::string kind = schema_type(p); + std::string value; + if (kind == "string") { + value = "params->" + field; + } else if (kind == "boolean") { + value = "(params->" + field + " ? \"true\" : \"false\")"; + } else if (kind == "integer") { + call_body.push_back(indent2 + "snprintf(number_values[number_count], sizeof(number_values[number_count]), \"%\" PRId64, params->" + field + ");"); + value = "number_values[number_count++]"; + } else { + cJSON* items = cJSON_GetObjectItemCaseSensitive(p, "items"); + std::string item_kind = schema_type(items); + call_body.push_back(indent2 + "for (size_t i = 0; i < params->" + field + ".count; ++i) {"); + if (item_kind == "string") { + value = "params->" + field + ".items[i]"; + } else if (item_kind == "boolean") { + value = "(params->" + field + ".items[i] ? \"true\" : \"false\")"; + } else { + call_body.push_back(indent2 + " snprintf(number_values[number_count], sizeof(number_values[number_count]), \"%\" PRId64, params->" + field + ".items[i]);"); + value = "number_values[number_count++]"; + } + call_body.push_back(indent2 + " encoded[count++] = (wf_xrpc_param){\"" + wire + "\", " + value + "};"); + call_body.push_back(indent2 + "}"); + if (condition) + call_body.push_back(" }"); + continue; + } + call_body.push_back(indent2 + "encoded[count++] = (wf_xrpc_param){\"" + wire + "\", " + value + "};"); + if (condition) + call_body.push_back(" }"); + } + } + // _call variant + out.push_back("wf_status " + base + "_call(wf_xrpc_client *client,"); + out.push_back(" const " + base + "_params *params, wf_response *response) {"); + out.push_back(" if (!client) return WF_ERR_INVALID_ARG;"); + out.insert(out.end(), call_body.begin(), call_body.end()); + out.push_back(" wf_status status = wf_xrpc_query_params(client, \"" + nsid + "\", encoded, count, response);"); + out.push_back(" free(encoded); free(number_values); return status;"); + out.push_back("}"); + out.push_back(""); + // _call_auth variant + out.push_back("wf_status " + base + "_call_auth(wf_auth_client *client,"); + out.push_back(" const " + base + "_params *params, wf_response *response) {"); + out.push_back(" if (!client) return WF_ERR_INVALID_ARG;"); + out.insert(out.end(), call_body.begin(), call_body.end()); + out.push_back(" wf_status status = wf_auth_client_query_params(client, \"" + nsid + "\", encoded, count, response);"); + out.push_back(" free(encoded); free(number_values); return status;"); + out.push_back("}"); + out.push_back(""); + } + } + } + + return join_lines(out); + } + +private: + mutable std::map> _objects; + mutable bool _objects_ready = false; + mutable std::vector _object_order; + mutable std::map> _unions; + mutable bool _unions_ready = false; + + void build_object_catalog() const { + _objects.clear(); + _object_order.clear(); + std::set visiting; + + std::function collect; + collect = [&](const std::string& nsid, const std::string& name, cJSON* schema) { + auto it = _objects.find(name); + if (it != _objects.end()) { + if (it->second.first != nsid || it->second.second != schema) + throw std::runtime_error("inline object name collision for " + name); + return; + } + if (visiting.count(name)) + throw std::runtime_error("recursive inline object in " + name); + visiting.insert(name); + + std::function visit; + visit = [&](cJSON* value, const std::string& path) { + std::string kind = schema_type(value); + if (kind == "array") { + cJSON* items = cJSON_GetObjectItemCaseSensitive(value, "items"); + visit(items ? items : cJSON_CreateObject(), path + "_item"); + } else if (kind == "object") { + std::string child = name + "_" + snake(path); + collect(nsid, child, value); + } + }; + cJSON* props = cJSON_GetObjectItemCaseSensitive(schema, "properties"); + if (props) { + for (cJSON* p = props->child; p; p = p->next) { + if (!p->string) + continue; + visit(p, field_name(schema, p->string)); + } + } + visiting.erase(name); + // Children are inserted first, so direct embedded fields are complete. + _objects[name] = std::make_pair(nsid, schema); + _object_order.push_back(name); + }; + + for (const auto& doc : docs) { + for (auto& pair : schemas(doc)) + collect(doc.id, pair.first, pair.second); + } + } + + static std::string join_lines(const std::vector& lines) { + std::string result; + for (size_t i = 0; i < lines.size(); ++i) { + if (i) + result += '\n'; + result += lines[i]; + } + return result; + } + + static bool target_ends_with(const std::string& target, const char* suffix) { + size_t n = strlen(suffix); + return target.size() >= n && target.compare(target.size() - n, n, suffix) == 0; + } + + static std::string target_field(const std::string& target) { + // rsplit("->", 1)[-1].split(".")[-1] + size_t arrow = target.rfind("->"); + std::string rest = arrow == std::string::npos ? target : target.substr(arrow + 2); + size_t dot = rest.rfind('.'); + return dot == std::string::npos ? rest : rest.substr(dot + 1); + } + + static std::string target_suffix_name(const std::string& target) { + return target_field(target); + } + + static bool def_has_params(cJSON* def) { + cJSON* params = cJSON_GetObjectItemCaseSensitive(def, "parameters"); + return params != nullptr; + } +}; + +// --------------------------------------------------------------------------- +// CLI. +// --------------------------------------------------------------------------- + +static Doc load(const fs::path& path) { + std::ifstream stream(path); + if (!stream) + throw std::runtime_error(path.string() + ": cannot open"); + std::string text((std::istreambuf_iterator(stream)), std::istreambuf_iterator()); + cJSON* raw = cJSON_Parse(text.c_str()); + if (!raw) + throw std::runtime_error(path.string() + ": invalid JSON"); + Doc doc; + doc.raw = raw; + cJSON* lexicon = cJSON_GetObjectItemCaseSensitive(raw, "lexicon"); + cJSON* id = cJSON_GetObjectItemCaseSensitive(raw, "id"); + cJSON* defs = cJSON_GetObjectItemCaseSensitive(raw, "defs"); + if (!lexicon || !cJSON_IsNumber(lexicon) || (int)lexicon->valuedouble != 1 || + !id || !cJSON_IsString(id) || !id->valuestring) + throw std::runtime_error(path.string() + ": expected a Lexicon 1 document with an id"); + if (!defs || !cJSON_IsObject(defs)) + throw std::runtime_error(path.string() + ": expected a defs object"); + doc.id = id->valuestring; + doc.defs = defs; + return doc; +} + +static void write_file(const fs::path& path, const std::string& content) { + if (path.has_parent_path()) + fs::create_directories(path.parent_path()); + std::ofstream out(path); + if (!out) + throw std::runtime_error(path.string() + ": cannot write"); + out << content; +} + +int main(int argc, char* argv[]) { + std::vector lexicons; + std::string output; + std::string source_output; + std::string guard = "WOLFRAM_GENERATED_LEXICONS_H"; + std::string header_rel; + + for (int i = 1; i < argc; ++i) { + std::string arg = argv[i]; + if (arg == "-o" || arg == "--output") { + if (i + 1 < argc) + output = argv[++i]; + } else if (arg == "--source-output") { + if (i + 1 < argc) + source_output = argv[++i]; + } else if (arg == "--guard") { + if (i + 1 < argc) + guard = argv[++i]; + } else if (arg == "--header-rel") { + if (i + 1 < argc) + header_rel = argv[++i]; + } else { + lexicons.push_back(arg); + } + } + + if (lexicons.empty()) { + std::cerr << "wf_lexgen: no lexicon files given\n"; + return 1; + } + + try { + std::vector parsed; + for (const auto& path : lexicons) + parsed.push_back(load(path)); + Generator generator(std::move(parsed), guard); + + std::string result = generator.generate(); + if (!output.empty()) { + write_file(output, result); + } else { + std::cout << result; + } + + if (!source_output.empty()) { + if (output.empty()) + throw std::runtime_error("--source-output requires --output"); + std::string header = header_rel.empty() ? fs::path(output).filename().string() : header_rel; + write_file(source_output, generator.generate_source(header)); + } + } catch (const std::exception& error) { + std::cerr << "wf_lexgen: error: " << error.what() << "\n"; + return 1; + } + return 0; +} diff --git a/tools/wf_lexgen.py b/tools/wf_lexgen.py deleted file mode 100644 index a526df8..0000000 --- a/tools/wf_lexgen.py +++ /dev/null @@ -1,1172 +0,0 @@ -#!/usr/bin/env python3 -"""Generate C declarations and optional implementations from Lexicon JSON. - -Generated endpoint wrappers delegate all transport to wolfram/xrpc.h. -""" - -from __future__ import annotations - -import argparse -import json -import keyword -import re -import sys -from pathlib import Path -from typing import Any - - -C_KEYWORDS = { - "auto", "break", "case", "char", "const", "continue", "default", - "do", "double", "else", "enum", "extern", "float", "for", "goto", - "if", "inline", "int", "long", "register", "restrict", "return", - "short", "signed", "sizeof", "static", "struct", "switch", "typedef", - "union", "unsigned", "void", "volatile", "while", "_Alignas", - "_Alignof", "_Atomic", "_Bool", "_Complex", "_Generic", "_Imaginary", - "_Noreturn", "_Static_assert", "_Thread_local", -} - -CPP_KEYWORDS = { - "alignas", "alignof", "and", "and_eq", "asm", "bitand", "bitor", - "bool", "catch", "char8_t", "char16_t", "char32_t", "class", - "compl", "concept", "consteval", "constexpr", "constinit", "const_cast", - "co_await", "co_return", "co_yield", "decltype", "delete", "dynamic_cast", - "explicit", "export", "false", "friend", "mutable", "namespace", "new", - "noexcept", "not", "not_eq", "nullptr", "operator", "or", "or_eq", - "private", "protected", "public", "reflexpr", "reinterpret_cast", - "requires", "static_assert", "static_cast", "synchronized", "template", - "this", "thread_local", "throw", "true", "try", "typeid", "typename", - "using", "virtual", "wchar_t", "xor", "xor_eq", -} - - -def snake(value: str) -> str: - value = re.sub(r"([a-z0-9])([A-Z])", r"\1_\2", value) - value = re.sub(r"[^A-Za-z0-9]+", "_", value).strip("_").lower() - if not value: - value = "value" - if value[0].isdigit() or value in C_KEYWORDS or keyword.iskeyword(value): - value += "_" - return value - - -def member_name(value: str) -> str: - """Return an identifier safe as a struct/union member in C and C++.""" - value = snake(value) - return value + "_" if value in CPP_KEYWORDS else value - - -def type_name(nsid: str, suffix: str) -> str: - return "wf_lex_" + snake(nsid) + ("_" + snake(suffix) if suffix else "") - - -def ref_type(nsid: str, ref: str) -> str: - if ref.startswith("#"): - return type_name(nsid, ref[1:]) - if "#" in ref: - doc, fragment = ref.split("#", 1) - return type_name(doc, fragment) - return type_name(ref, "main") - - -def comment(text: str, indent: str = "") -> list[str]: - clean = " ".join(text.replace("*/", "* /").split()) - return [f"{indent}/** {clean} */"] if clean else [] - - -class Generator: - def __init__(self, docs: list[dict[str, Any]], guard: str): - self.docs = sorted(docs, key=lambda doc: doc["id"]) - self.guard = guard - self._objects: dict[str, tuple[str, dict[str, Any]]] | None = None - self._unions: dict[str, tuple[str, dict[str, Any]]] | None = None - - def resolve_ref(self, nsid: str, ref: str) -> tuple[str, dict[str, Any]] | None: - document_id, fragment = (nsid, ref[1:]) if ref.startswith("#") else ( - ref.split("#", 1) if "#" in ref else (ref, "main")) - for doc in self.docs: - if doc["id"] == document_id: - definition = doc.get("defs", {}).get(fragment) - if not isinstance(definition, dict): - return None - if definition.get("type") == "record": - definition = definition.get("record", definition) - return document_id, definition - return None - - @staticmethod - def field_name(schema: dict[str, Any], wire_name: str) -> str: - base = member_name(wire_name) - required = set(schema.get("required", [])) - presence = {"has_" + member_name(name) for name in schema.get("properties", {}) - if name not in required} - return base + "_value" if base in presence else base - - def c_type(self, nsid: str, owner: str, field: str, - schema: dict[str, Any]) -> tuple[str, bool]: - kind = schema.get("type") - if kind == "string": - return "const char *", False - if kind == "integer": - return "int64_t", True - if kind == "boolean": - return "bool", True - if kind == "bytes": - return "wf_lex_bytes", False - if kind == "cid-link": - return "wf_lex_cid_link", False - if kind == "blob": - return "wf_lex_blob", False - if kind == "ref": - resolved = self.resolve_ref(nsid, schema["ref"]) - if resolved and resolved[1].get("type") not in ("object",): - target_nsid, target = resolved - if target.get("type") == "token": - return "const char *", False - return self.c_type(target_nsid, owner, field, target) - # Borrowed pointers support recursive and cross-document refs - # without imposing declaration order on generated headers. - return "const " + ref_type(nsid, schema["ref"]) + " *", False - if kind == "object": - name = f"{owner}_{snake(field)}" - return name, False - # Closed/open unions decode into a generated typed struct keyed by the - # owner+field name (see collect_unions / emit_union_struct). The struct - # retains the raw JSON for re-encoding and open-union fallback. - if kind == "union": - return self.union_name(owner, field), False - # Unknown (free-form) values remain opaque JSON. - if kind == "unknown": - return "wf_lex_json", False - if kind == "array": - item_type, _ = self.c_type(nsid, owner, field + "_item", - schema.get("items", {"type": "unknown"})) - return f"WF_LEX_ARRAY({item_type})", False - return "wf_lex_json", False - - def union_name(self, owner: str, field: str) -> str: - return f"{owner}_{snake(field)}_union" - - def collect_unions(self) -> dict[str, tuple[str, dict[str, Any]]]: - """Register every union field (direct, array element, or reached - through a ref to an array/union) as a typed struct, keyed by the same - owner+field name the field's C type uses.""" - if self._unions is not None: - return self._unions - self._unions = {} - catalog = self.object_catalog() - - def walk(nsid: str, owner: str, schema: dict[str, Any], path: str) -> None: - kind = schema.get("type") - if kind == "union": - self._unions[self.union_name(owner, path)] = (nsid, schema) - elif kind == "array": - walk(nsid, owner, schema.get("items", {"type": "unknown"}), - path + "_item") - elif kind == "object": - for wire_name, prop in schema.get("properties", {}).items(): - walk(nsid, owner, prop, self.field_name(schema, wire_name)) - elif kind == "ref": - resolved = self.resolve_ref(nsid, schema["ref"]) - if not resolved: - return - target_nsid, target = resolved - if target.get("type") == "object": - # Borrowed pointer; its nested unions are walked via catalog. - return - walked = {"type": "string"} if target.get("type") == "token" else target - walk(target_nsid, owner, walked, path) - - for name, (nsid, schema) in catalog.items(): - walk(nsid, name, schema, "") - return self._unions - - def union_members(self, nsid: str, - schema: dict[str, Any]) -> list[tuple[int, str, str, str]]: - """Return (index, full_$type, c_type, member_name) for each resolvable - object member of the union. Non-object refs (external/knownValues) are - skipped; the raw JSON fallback covers them.""" - members: list[tuple[int, str, str, str]] = [] - seen: set[str] = set() - for i, ref in enumerate(schema.get("refs", [])): - resolved = self.resolve_ref(nsid, ref) - if not resolved or resolved[1].get("type") != "object": - continue - full = ref if not ref.startswith("#") else nsid + ref - frag = ref.split("#", 1)[1] if "#" in ref else ref.rsplit(".", 1)[-1] - mname = member_name(frag) - if mname in seen: - mname = f"{mname}_{i}" - seen.add(mname) - members.append((i, full, ref_type(nsid, ref), mname)) - return members - - def emit_union_struct(self, name: str, nsid: str, - schema: dict[str, Any]) -> list[str]: - lines = [f"typedef struct {name} {{", - " int kind;", - " /* Retained raw JSON (mirrors wf_lex_json) for re-encoding", - " * and open-union/unknown $type fallback. */", - " const char *data;", - " size_t length;", - " union {"] - members = self.union_members(nsid, schema) - if members: - for _idx, _full, ctype, mname in members: - lines.append(f" const {ctype} *{mname};") - else: - lines.append(" unsigned char _unused;") - lines += [" } value;", f"}} {name};", ""] - return lines - - def emit_union_decoder(self, name: str, nsid: str, - schema: dict[str, Any]) -> list[str]: - lines = [ - f"static wf_status wf_lex_decode_{name}(cJSON *node, {name} *value) {{", - " if (!node || !value) return WF_ERR_INVALID_ARG;", - " memset(value, 0, sizeof(*value));", - " value->kind = -1;", - " char *raw = cJSON_PrintUnformatted(node);", - " if (!raw) return WF_ERR_ALLOC;", - " value->data = raw; value->length = strlen(raw);", - ' cJSON *type = cJSON_GetObjectItemCaseSensitive(node, "$type");', - " const char *t = (type && cJSON_IsString(type)) ? type->valuestring : NULL;", - " if (t) {", - ] - for idx, full, ctype, mname in self.union_members(nsid, schema): - lines.append(f' if (strcmp(t, "{full}") == 0) {{') - lines.append(f" value->kind = {idx};") - lines.append(f" {ctype} *m = calloc(1, sizeof(*m));") - lines.append(f" if (!m) {{ wf_lex_clear_{name}(value); return WF_ERR_ALLOC; }}") - lines.append(f" wf_status status = wf_lex_decode_{ctype}(node, m);") - lines.append(f" if (status != WF_OK) {{") - lines.append(f" free(m); value->kind = -1;") - lines.append(f" }} else {{") - lines.append(f" value->value.{mname} = m;") - lines.append(f" }}") - lines.append(" }") - lines += [" }", " return WF_OK;", "}", ""] - return lines - - def emit_union_clear(self, name: str, nsid: str, - schema: dict[str, Any]) -> list[str]: - lines = [f"static void wf_lex_clear_{name}({name} *value) {{", - " if (!value) return;", - " free((void *)value->data);", - " switch (value->kind) {"] - for idx, _full, ctype, mname in self.union_members(nsid, schema): - lines.append(f" case {idx}:") - lines.append(f" if (value->value.{mname}) {{") - lines.append(f" wf_lex_clear_{ctype}(({ctype} *)value->value.{mname});") - lines.append(f" free((void *)value->value.{mname});") - lines.append(" }") - lines.append(" break;") - lines += [" default: break;", " }", - " memset(value, 0, sizeof(*value));", "}", ""] - return lines - - def json_input_type(self, nsid: str, base: str, - schema: dict[str, Any]) -> str | None: - """Return the named type accepted by a JSON endpoint input.""" - if schema.get("type") in { - "array", "blob", "boolean", "bytes", "cid-link", "integer", - "object", "ref", "string", "union", "unknown"}: - return base + "_input" - return None - - def json_input_alias_type(self, nsid: str, base: str, - schema: dict[str, Any]) -> str: - """Return the C type aliased by a non-object endpoint input.""" - if schema.get("type") == "ref": - resolved = self.resolve_ref(nsid, schema["ref"]) - if resolved and resolved[1].get("type") == "object": - return ref_type(nsid, schema["ref"]) - return self.c_type(nsid, base, "input", schema)[0] - - def emit_object(self, nsid: str, name: str, - schema: dict[str, Any]) -> list[str]: - lines = comment(schema.get("description", "")) - lines.append(f"typedef struct {name} {{") - properties = schema.get("properties", {}) - required = set(schema.get("required", [])) - if not properties: - lines.append(" unsigned char _unused;") - for wire_name, prop in properties.items(): - field = self.field_name(schema, wire_name) - ctype, scalar = self.c_type(nsid, name, field, prop) - lines.extend(comment(prop.get("description", ""), " ")) - if wire_name not in required: - lines.append(f" bool has_{field};") - lines.append(f" {ctype} {field};") - lines.append(f"}} {name};") - return lines - - def schemas(self, doc: dict[str, Any]) -> list[tuple[str, dict[str, Any]]]: - nsid = doc["id"] - found: list[tuple[str, dict[str, Any]]] = [] - for def_name, definition in doc.get("defs", {}).items(): - kind = definition.get("type") - base = type_name(nsid, def_name) - if kind == "object": - found.append((base, definition)) - elif kind == "record" and definition.get("record", {}).get("type") == "object": - found.append((base, definition["record"])) - elif kind in ("query", "procedure"): - params = definition.get("parameters") - if params and params.get("type") == "params": - found.append((base + "_params", params)) - for label in ("input", "output"): - value = definition.get(label, {}) - schema = value.get("schema") - if schema and schema.get("type") == "object": - found.append((base + "_" + label, schema)) - return found - - def object_catalog(self) -> dict[str, tuple[str, dict[str, Any]]]: - """Return named and inline objects in dependency-safe deterministic order.""" - if self._objects is not None: - return self._objects - catalog: dict[str, tuple[str, dict[str, Any]]] = {} - visiting: set[str] = set() - - def collect(nsid: str, name: str, schema: dict[str, Any]) -> None: - previous = catalog.get(name) - if previous is not None: - if previous != (nsid, schema): - raise ValueError(f"inline object name collision for {name}") - return - if name in visiting: - raise ValueError(f"recursive inline object in {name}") - visiting.add(name) - - def visit(value: dict[str, Any], path: str) -> None: - kind = value.get("type") - if kind == "array": - visit(value.get("items", {"type": "unknown"}), path + "_item") - elif kind == "object": - child = f"{name}_{snake(path)}" - collect(nsid, child, value) - - for wire_name, prop in schema.get("properties", {}).items(): - visit(prop, self.field_name(schema, wire_name)) - visiting.remove(name) - # Children are inserted first, so direct embedded fields are complete. - catalog[name] = (nsid, schema) - - for doc in self.docs: - for name, schema in self.schemas(doc): - collect(doc["id"], name, schema) - self._objects = catalog - return catalog - - def encoder_objects(self) -> dict[str, tuple[str, dict[str, Any]]]: - catalog = self.object_catalog() - wanted: set[str] = set() - - def visit_object(nsid: str, name: str, schema: dict[str, Any]) -> None: - if name in wanted: - return - wanted.add(name) - for wire_name, prop in schema.get("properties", {}).items(): - field = self.field_name(schema, wire_name) - visit_schema(nsid, name, prop, field) - - def visit_schema(nsid: str, owner: str, schema: dict[str, Any], - path: str) -> None: - kind = schema.get("type") - if kind == "array": - visit_schema(nsid, owner, - schema.get("items", {"type": "unknown"}), - path + "_item") - elif kind == "object": - child = f"{owner}_{snake(path)}" - visit_object(nsid, child, catalog[child][1]) - elif kind == "ref": - resolved = self.resolve_ref(nsid, schema["ref"]) - if not resolved: - raise ValueError( - f"cannot encode unresolved ref {schema['ref']} in {owner}" - ) - target_nsid, target = resolved - if target.get("type") == "object": - target_name = ref_type(nsid, schema["ref"]) - if target_name not in catalog: - raise ValueError( - f"cannot encode ref {schema['ref']} in {owner}" - ) - visit_object(target_nsid, target_name, catalog[target_name][1]) - elif target.get("type") != "token": - visit_schema(target_nsid, owner, target, path) - - for doc in self.docs: - for def_name, definition in doc.get("defs", {}).items(): - schema = definition.get("input", {}).get("schema", {}) - if definition.get("type") not in ("query", "procedure"): - continue - if schema.get("type") == "object": - root = type_name(doc["id"], def_name) + "_input" - visit_object(doc["id"], root, schema) - elif schema: - visit_schema(doc["id"], type_name(doc["id"], def_name), - schema, "input") - return {name: value for name, value in catalog.items() if name in wanted} - - def referenced_types(self) -> set[str]: - found: set[str] = set() - - def visit(nsid: str, value: Any) -> None: - if isinstance(value, dict): - if value.get("type") == "ref" and isinstance(value.get("ref"), str): - resolved = self.resolve_ref(nsid, value["ref"]) - if not resolved or resolved[1].get("type") == "object": - found.add(ref_type(nsid, value["ref"])) - for child in value.values(): - visit(nsid, child) - elif isinstance(value, list): - for child in value: - visit(nsid, child) - - for doc in self.docs: - visit(doc["id"], doc.get("defs", {})) - return found - - def generate(self) -> str: - out = [ - "/* Generated by tools/wf_lexgen.py; do not edit. */", - f"#ifndef {self.guard}", f"#define {self.guard}", "", - "#include ", "#include ", "#include ", - "#include ", "#include ", "", - "#ifdef __cplusplus", 'extern "C" {', "#endif", "", - "/** Encoded JSON view. Decoded outputs own data; input values borrow it. */", - "typedef struct wf_lex_json { const char *data; size_t length; } wf_lex_json;", - "/** Byte sequence view. Decoded outputs own data; input values borrow it. */", - "typedef struct wf_lex_bytes { const uint8_t *data; size_t length; } wf_lex_bytes;", - "/** CID link string view. */", - "typedef struct wf_lex_cid_link { const char *cid; } wf_lex_cid_link;", - "/** Typed AT Protocol blob reference. */", - "typedef struct wf_lex_blob { const char *cid; const char *mime_type; int64_t size; } wf_lex_blob;", - "#define WF_LEX_ARRAY(type_) struct { type_ const *items; size_t count; }", "", - ] - for doc in self.docs: - nsid = doc["id"] - main = doc.get("defs", {}).get("main", {}) - kind = main.get("type", "definition") - symbol = snake(nsid).upper() - out.extend(comment(main.get("description", ""))) - out.append(f'#define WF_LEX_{symbol}_NSID "{nsid}"') - out.append(f'#define WF_LEX_{symbol}_KIND "{kind}"') - out.append("") - - # Forward declarations allow refs to definitions emitted later. - objects = self.object_catalog() - unions = self.collect_unions() - declarations = set(objects) | self.referenced_types() | set(unions) - for name in sorted(declarations): - out.append(f"typedef struct {name} {name};") - if declarations: - out.append("") - # Union structs only contain pointers to object members plus a raw JSON - # view, so they never need a complete object definition; emit them - # first so object structs that embed a union value see a complete type. - for name in sorted(unions): - nsid, schema = unions[name] - out.extend(self.emit_union_struct(name, nsid, schema)) - for name, (nsid, schema) in objects.items(): - out.extend(self.emit_object(nsid, name, schema)) - out.append("") - for doc in self.docs: - nsid = doc["id"] - for def_name, definition in doc.get("defs", {}).items(): - if definition.get("type") not in ("query", "procedure"): - continue - schema = definition.get("input", {}).get("schema") - if not schema or schema.get("type") == "object": - continue - base = type_name(nsid, def_name) - if self.json_input_type(nsid, base, schema): - alias = self.json_input_alias_type(nsid, base, schema) - out.append(f"typedef {alias} {base}_input;") - out.append("") - for doc in self.docs: - nsid = doc["id"] - for def_name, definition in doc.get("defs", {}).items(): - if definition.get("type") not in ("query", "procedure"): - continue - base = type_name(nsid, def_name) - input_schema = definition.get("input", {}).get("schema") - input_type = (self.json_input_type(nsid, base, input_schema) - if input_schema else None) - if input_type: - out.append(f"wf_status {base}_input_encode_json(") - out.append(f" const {input_type} *value, char **out_json);") - out.append("/** Free JSON returned by the matching encoder. */") - out.append(f"void {base}_json_free(char *json);") - output_schema = definition.get("output", {}).get("schema") - if output_schema and output_schema.get("type") == "object": - out.append("/** Decode an owning output value; free it with the matching function. */") - out.append(f"wf_status {base}_output_decode_json(") - out.append(f" const char *json, size_t length, {base}_output **out_value);") - out.append(f"void {base}_output_free({base}_output *value);") - if definition.get("type") == "procedure": - if input_type: - out.append(f"wf_status {base}_call(wf_xrpc_client *client,") - out.append(f" const {input_type} *input, wf_response *out);") - out.append(f"wf_status {base}_call_auth(wf_auth_client *client,") - out.append(f" const {input_type} *input, wf_response *out);") - else: - out.append(f"wf_status {base}_call(wf_xrpc_client *client, wf_response *out);") - out.append(f"wf_status {base}_call_auth(wf_auth_client *client, wf_response *out);") - elif definition.get("type") == "query": - params = definition.get("parameters") - has_params = params and params.get("type") == "params" - arg = f"const {base}_params *params, " if has_params else "" - out.append(f"wf_status {base}_call(wf_xrpc_client *client,") - out.append(f" {arg}wf_response *out);") - out.append(f"wf_status {base}_call_auth(wf_auth_client *client,") - out.append(f" {arg}wf_response *out);") - out.append("") - out.extend(["#undef WF_LEX_ARRAY", "", "#ifdef __cplusplus", "}", - "#endif", "", f"#endif /* {self.guard} */", ""]) - return "\n".join(out) - - def add_encoded_value(self, lines: list[str], nsid: str, owner: str, - prop: dict[str, Any], target: str, result: str, - path: str, indent: str) -> None: - kind = prop.get("type") - if kind == "string": - lines.append(f"{indent}if (!{target}) goto invalid;") - expr = f"cJSON_CreateString({target})" - elif kind == "integer": - expr = f"cJSON_CreateNumber((double){target})" - elif kind == "boolean": - expr = f"cJSON_CreateBool({target})" - elif kind == "union": - # Re-emit the retained raw JSON (see emit_union_struct). - lines.append(f"{indent}if (!{target}.data) goto invalid;") - lines.append(f'{indent}{result} = cJSON_ParseWithLength({target}.data, {target}.length);') - lines.append(f"{indent}if (!{result}) goto invalid;") - expr = None - elif kind == "unknown": - lines.append(f"{indent}if (!{target}.data) goto invalid;") - lines.append(f'{indent}{result} = cJSON_ParseWithLength({target}.data, {target}.length);') - lines.append(f"{indent}if (!{result}) goto invalid;") - expr = None - elif kind == "bytes": - lines.append(f"{indent}status = wf_lex_bytes_encode(&{target}, &{result});") - lines.append(f"{indent}if (status != WF_OK) goto status_fail;") - expr = None - elif kind == "cid-link": - lines.append(f"{indent}status = wf_lex_cid_encode(&{target}, &{result});") - lines.append(f"{indent}if (status != WF_OK) goto status_fail;") - expr = None - elif kind == "blob": - lines.append(f"{indent}status = wf_lex_blob_encode(&{target}, &{result});") - lines.append(f"{indent}if (status != WF_OK) goto status_fail;") - expr = None - elif kind == "object": - inline = f"{owner}_{snake(path)}" - lines.append(f"{indent}status = wf_lex_encode_{inline}(&{target}, &{result});") - lines.append(f"{indent}if (status != WF_OK) goto status_fail;") - expr = None - elif kind == "ref": - resolved = self.resolve_ref(nsid, prop["ref"]) - if not resolved: - raise ValueError(f"cannot encode unresolved ref {prop['ref']} in {owner}") - target_nsid, target_schema = resolved - if target_schema.get("type") == "object": - referenced = ref_type(nsid, prop["ref"]) - lines.append(f"{indent}if (!{target}) goto invalid;") - lines.append(f"{indent}status = wf_lex_encode_{referenced}({target}, &{result});") - lines.append(f"{indent}if (status != WF_OK) goto status_fail;") - else: - if target_schema.get("type") == "token": - target_schema = {"type": "string"} - self.add_encoded_value(lines, target_nsid, owner, target_schema, - target, result, path, indent) - expr = None - elif kind == "array": - item_schema = prop.get("items", {}) - lines.append(f"{indent}if ({target}.count && !{target}.items) goto invalid;") - lines.append(f"{indent}{result} = cJSON_CreateArray();") - lines.append(f"{indent}if (!{result}) goto fail;") - suffix = snake(path) - index = f"i_{suffix}" - element = f"element_{suffix}" - lines.append(f"{indent}for (size_t {index} = 0; {index} < {target}.count; ++{index}) {{") - lines.append(f"{indent} cJSON *{element} = NULL;") - self.add_encoded_value(lines, nsid, owner, item_schema, - f"{target}.items[{index}]", element, - path + "_item", indent + " ") - lines.append(f"{indent} if (!cJSON_AddItemToArray({result}, {element})) {{") - lines.append(f"{indent} cJSON_Delete({element}); goto fail;") - lines.append(f"{indent} }}") - lines.append(f"{indent}}}") - expr = None - else: - raise ValueError(f"JSON encoding is not supported for {path} ({kind})") - if expr: - lines.append(f"{indent}{result} = {expr};") - lines.append(f"{indent}if (!{result}) goto fail;") - - def add_value(self, lines: list[str], nsid: str, owner: str, - prop: dict[str, Any], field: str, wire_name: str, - indent: str = " ") -> None: - lines.append(f"{indent}item = NULL;") - self.add_encoded_value(lines, nsid, owner, prop, f"value->{field}", - "item", field, indent) - lines.append(f'{indent}if (!item || !cJSON_AddItemToObject(root, "{wire_name}", item)) {{ cJSON_Delete(item); item = NULL; goto fail; }}') - lines.append(f"{indent}item = NULL;") - - def emit_object_encoder(self, nsid: str, name: str, - schema: dict[str, Any]) -> list[str]: - lines = [f"static wf_status wf_lex_encode_{name}(const {name} *value, cJSON **out) {{", - " if (!value || !out) return WF_ERR_INVALID_ARG;", " *out = NULL;", - " wf_status status = WF_OK;", " (void)status;", - " cJSON *root = cJSON_CreateObject();", " cJSON *item = NULL;", - " (void)item;", " if (!root) return WF_ERR_ALLOC;"] - required = set(schema.get("required", [])) - for wire_name, prop in schema.get("properties", {}).items(): - field = self.field_name(schema, wire_name) - if wire_name not in required: - lines.append(f" if (value->has_{field}) {{") - self.add_value(lines, nsid, name, prop, field, wire_name, " ") - lines.append(" }") - else: - self.add_value(lines, nsid, name, prop, field, wire_name) - lines.append(" *out = root; return WF_OK;") - if any("goto invalid;" in line for line in lines): - lines += ["invalid:", " cJSON_Delete(item); cJSON_Delete(root); return WF_ERR_INVALID_ARG;"] - if any("goto fail;" in line for line in lines): - lines += ["fail:", " cJSON_Delete(item); cJSON_Delete(root); return WF_ERR_ALLOC;"] - if any("goto status_fail;" in line for line in lines): - lines += ["status_fail:", " cJSON_Delete(item); cJSON_Delete(root); return status;"] - lines += ["}", ""] - return lines - - def emit_decode_value(self, lines: list[str], nsid: str, owner: str, - schema: dict[str, Any], target: str, - source: str, indent: str, field: str = "") -> None: - kind = schema.get("type") - if kind == "string": - lines.append(f"{indent}if (!cJSON_IsString({source})) {{ status = WF_ERR_INVALID_ARG; goto cleanup; }}") - lines.append(f"{indent}{target} = wf_lex_strdup({source}->valuestring);") - lines.append(f"{indent}if (!{target}) {{ status = WF_ERR_ALLOC; goto cleanup; }}") - elif kind == "integer": - lines.append(f"{indent}if (!wf_lex_json_integer({source}, &{target})) {{ status = WF_ERR_INVALID_ARG; goto cleanup; }}") - elif kind == "boolean": - lines.append(f"{indent}if (!cJSON_IsBool({source})) {{ status = WF_ERR_INVALID_ARG; goto cleanup; }}") - lines.append(f"{indent}{target} = cJSON_IsTrue({source});") - elif kind == "union": - name = self.union_name(owner, field) - lines.append(f"{indent}status = wf_lex_decode_{name}({source}, &({target}));") - lines.append(f"{indent}if (status != WF_OK) goto cleanup;") - elif kind == "unknown": - lines.append(f"{indent}status = wf_lex_json_copy({source}, &{target});") - lines.append(f"{indent}if (status != WF_OK) goto cleanup;") - elif kind == "bytes": - lines.append(f"{indent}status = wf_lex_bytes_decode({source}, &{target});") - lines.append(f"{indent}if (status != WF_OK) goto cleanup;") - elif kind == "cid-link": - lines.append(f"{indent}status = wf_lex_cid_decode({source}, &{target});") - lines.append(f"{indent}if (status != WF_OK) goto cleanup;") - elif kind == "blob": - lines.append(f"{indent}status = wf_lex_blob_decode({source}, &{target});") - lines.append(f"{indent}if (status != WF_OK) goto cleanup;") - elif kind == "object": - inline = f"{owner}_{snake(target.rsplit('->', 1)[-1].split('.')[-1])}" - # Array elements pass an owner-specific item type through c_type. - if target.endswith("items[i]"): - inline = f"{owner}_array_item" - lines.append(f"{indent}status = wf_lex_decode_{inline}({source}, &{target});") - lines.append(f"{indent}if (status != WF_OK) goto cleanup;") - elif kind == "ref": - ref = ref_type(nsid, schema["ref"]) - resolved = self.resolve_ref(nsid, schema["ref"]) - if resolved and resolved[1].get("type") != "object": - target_nsid, target_schema = resolved - if target_schema.get("type") == "token": - target_schema = {"type": "string"} - self.emit_decode_value(lines, target_nsid, owner, target_schema, - target, source, indent, field) - elif ref not in self.object_catalog(): - raise ValueError(f"cannot decode unresolved ref {schema['ref']} in {owner}") - else: - lines.append(f"{indent}{target} = calloc(1, sizeof(*{target}));") - lines.append(f"{indent}if (!{target}) {{ status = WF_ERR_ALLOC; goto cleanup; }}") - lines.append(f"{indent}status = wf_lex_decode_{ref}({source}, ({ref} *){target});") - lines.append(f"{indent}if (status != WF_OK) goto cleanup;") - elif kind == "array": - item_schema = schema.get("items", {"type": "unknown"}) - array_field = target.rsplit("->", 1)[-1].split(".")[-1] - item_field = array_field + "_item" - item_type, _ = self.c_type(nsid, owner, item_field, item_schema) - lines.append(f"{indent}if (!cJSON_IsArray({source})) {{ status = WF_ERR_INVALID_ARG; goto cleanup; }}") - lines.append(f"{indent}{target}.count = (size_t)cJSON_GetArraySize({source});") - lines.append(f"{indent}if ({target}.count) {{") - lines.append(f"{indent} {item_type} *items = calloc({target}.count, sizeof(*items));") - lines.append(f"{indent} if (!items) {{ status = WF_ERR_ALLOC; goto cleanup; }}") - lines.append(f"{indent} {target}.items = items;") - lines.append(f"{indent} for (size_t i = 0; i < {target}.count; ++i) {{") - lines.append(f"{indent} cJSON *element = cJSON_GetArrayItem({source}, (int)i);") - if item_schema.get("type") == "object": - lines.append(f"{indent} status = wf_lex_decode_{item_type}(element, &items[i]);") - lines.append(f"{indent} if (status != WF_OK) goto cleanup;") - else: - self.emit_decode_value(lines, nsid, owner, item_schema, "items[i]", - "element", indent + " ", item_field) - lines.append(f"{indent} }}") - lines.append(f"{indent}}}") - else: - raise ValueError(f"JSON decoding is not supported for {owner} ({kind})") - - def emit_clear_value(self, lines: list[str], nsid: str, owner: str, - schema: dict[str, Any], target: str, - indent: str, field: str = "") -> None: - kind = schema.get("type") - if kind == "string": - lines.append(f"{indent}free((void *){target});") - elif kind == "union": - union_type = self.union_name(owner, field) - if target.endswith("items[i]"): - lines.append( - f"{indent}wf_lex_clear_{union_type}(({union_type} *)&({target}));" - ) - else: - lines.append(f"{indent}wf_lex_clear_{union_type}(&({target}));") - elif kind == "unknown": - lines.append(f"{indent}free((void *){target}.data);") - elif kind == "bytes": - lines.append(f"{indent}free((void *){target}.data);") - elif kind == "cid-link": - lines.append(f"{indent}free((void *){target}.cid);") - elif kind == "blob": - lines.append(f"{indent}free((void *){target}.cid);") - lines.append(f"{indent}free((void *){target}.mime_type);") - elif kind == "object": - field = target.rsplit("->", 1)[-1].split(".")[-1] - inline = f"{owner}_{snake(field)}" - if target.endswith("items[i]"): - inline = f"{owner}_array_item" - lines.append(f"{indent}wf_lex_clear_{inline}(&{target});") - elif kind == "ref": - ref = ref_type(nsid, schema["ref"]) - resolved = self.resolve_ref(nsid, schema["ref"]) - if resolved and resolved[1].get("type") != "object": - target_nsid, target_schema = resolved - if target_schema.get("type") == "token": - target_schema = {"type": "string"} - self.emit_clear_value(lines, target_nsid, owner, target_schema, - target, indent, field) - else: - lines.append(f"{indent}if ({target}) {{ wf_lex_clear_{ref}(({ref} *){target}); free((void *){target}); }}") - elif kind == "array": - item = schema.get("items", {"type": "unknown"}) - array_field = target.rsplit("->", 1)[-1].split(".")[-1] - item_field = array_field + "_item" - item_type, _ = self.c_type(nsid, owner, item_field, item) - lines.append(f"{indent}for (size_t i = 0; i < {target}.count; ++i) {{") - if item.get("type") == "object": - lines.append(f"{indent} wf_lex_clear_{item_type}(({item_type} *)&{target}.items[i]);") - else: - self.emit_clear_value(lines, nsid, owner, item, f"{target}.items[i]", - indent + " ", item_field) - lines.append(f"{indent}}}") - lines.append(f"{indent}free((void *){target}.items);") - - def emit_object_decoder(self, nsid: str, name: str, - schema: dict[str, Any]) -> list[str]: - lines = [f"static void wf_lex_clear_{name}({name} *value) {{", - " if (!value) return;"] - for wire_name, prop in schema.get("properties", {}).items(): - self.emit_clear_value(lines, nsid, name, prop, - f"value->{self.field_name(schema, wire_name)}", " ", - self.field_name(schema, wire_name)) - lines += [" memset(value, 0, sizeof(*value));", "}", "", - f"static wf_status wf_lex_decode_{name}(cJSON *node, {name} *value) {{", - " wf_status status = WF_OK;", - " (void)status;", - " if (!cJSON_IsObject(node) || !value) return WF_ERR_INVALID_ARG;"] - required = set(schema.get("required", [])) - for wire_name, prop in schema.get("properties", {}).items(): - field = self.field_name(schema, wire_name) - lines += [" {", f' cJSON *member = cJSON_GetObjectItemCaseSensitive(node, "{wire_name}");'] - if wire_name in required: - lines.append(" if (!member) { status = WF_ERR_INVALID_ARG; goto cleanup; }") - else: - lines.append(" if (member) {") - lines.append(f" value->has_{field} = true;") - indent = " " if wire_name not in required else " " - self.emit_decode_value(lines, nsid, name, prop, f"value->{field}", "member", indent, field) - if wire_name not in required: - lines.append(" }") - lines.append(" }") - lines.append(" return WF_OK;") - if any("goto cleanup;" in line for line in lines): - lines += ["cleanup:", f" wf_lex_clear_{name}(value);", - " return status;"] - lines += ["}", ""] - return lines - - def generate_source(self, header_name: str) -> str: - out = [ - "/* Generated by tools/wf_lexgen.py; do not edit. */", - f'#include "{header_name}"', "#include ", - "#include ", "#include ", - "#include ", "#include ", "#include ", - "#include ", "#include ", "", - "#if defined(__GNUC__) || defined(__clang__)", - "#define WF_LEX_UNUSED __attribute__((unused))", "#else", - "#define WF_LEX_UNUSED", "#endif", "", - "static WF_LEX_UNUSED char *wf_lex_strdup(const char *source) {", - " size_t length = strlen(source) + 1; char *copy = malloc(length);", - " if (copy) memcpy(copy, source, length);", - " return copy;", "}", "", - "static WF_LEX_UNUSED bool wf_lex_json_integer(cJSON *item, int64_t *out) {", - " if (!cJSON_IsNumber(item) || !isfinite(item->valuedouble) ||", - " item->valuedouble < -9007199254740991.0 || item->valuedouble > 9007199254740991.0 ||", - " (double)(int64_t)item->valuedouble != item->valuedouble) return false;", - " *out = (int64_t)item->valuedouble; return true;", "}", "", - "static WF_LEX_UNUSED wf_status wf_lex_json_copy(cJSON *item, wf_lex_json *out) {", - " char *json = cJSON_PrintUnformatted(item); if (!json) return WF_ERR_ALLOC;", - " out->data = json; out->length = strlen(json); return WF_OK;", "}", "", - "static WF_LEX_UNUSED wf_status wf_lex_bytes_encode(const wf_lex_bytes *value, cJSON **out) {", - " if (!value || !out || (value->length && !value->data) || value->length > INT_MAX)", - " return WF_ERR_INVALID_ARG;", - " *out = NULL;", - " if (value->length > (SIZE_MAX / 4) * 3 - 2) return WF_ERR_INVALID_ARG;", - " size_t length = 4 * ((value->length + 2) / 3);", - " char *encoded = malloc(length + 1); if (!encoded) return WF_ERR_ALLOC;", - " const uint8_t *data = value->length ? value->data : (const uint8_t *)\"\";", - " int written = EVP_EncodeBlock((unsigned char *)encoded, data, (int)value->length);", - " if (written < 0 || (size_t)written != length) { free(encoded); return WF_ERR_INVALID_ARG; }", - " encoded[length] = '\\0'; cJSON *root = cJSON_CreateObject();", - " cJSON *tag = cJSON_CreateString(encoded); free(encoded);", - " if (!root || !tag || !cJSON_AddItemToObject(root, \"$bytes\", tag)) {", - " cJSON_Delete(tag); cJSON_Delete(root); return WF_ERR_ALLOC;", - " }", - " *out = root; return WF_OK;", "}", "", - "static WF_LEX_UNUSED wf_status wf_lex_cid_encode(const wf_lex_cid_link *value, cJSON **out) {", - " if (!value || !out || !value->cid || !value->cid[0]) return WF_ERR_INVALID_ARG;", - " *out = NULL;", - " cJSON *root = cJSON_CreateObject(); cJSON *link = cJSON_CreateString(value->cid);", - " if (!root || !link || !cJSON_AddItemToObject(root, \"$link\", link)) {", - " cJSON_Delete(link); cJSON_Delete(root); return WF_ERR_ALLOC;", - " }", - " *out = root; return WF_OK;", "}", "", - "static WF_LEX_UNUSED wf_status wf_lex_blob_encode(const wf_lex_blob *value, cJSON **out) {", - " if (!value || !out || !value->cid || !value->cid[0] || !value->mime_type || value->size < 0)", - " return WF_ERR_INVALID_ARG;", - " *out = NULL;", - " cJSON *root = cJSON_CreateObject(); cJSON *ref = NULL;", - " wf_lex_cid_link link = {value->cid};", - " if (!root) return WF_ERR_ALLOC;", - " wf_status status = wf_lex_cid_encode(&link, &ref);", - " if (status != WF_OK) { cJSON_Delete(root); return status; }", - " cJSON *type = cJSON_CreateString(\"blob\");", - " cJSON *mime = cJSON_CreateString(value->mime_type);", - " cJSON *size = cJSON_CreateNumber((double)value->size);", - " if (!type || !mime || !size) {", - " cJSON_Delete(type); cJSON_Delete(ref); cJSON_Delete(mime); cJSON_Delete(size);", - " cJSON_Delete(root); return WF_ERR_ALLOC;", - " }", - " if (!cJSON_AddItemToObject(root, \"$type\", type)) goto blob_fail;", - " type = NULL;", - " if (!cJSON_AddItemToObject(root, \"ref\", ref)) goto blob_fail;", - " ref = NULL;", - " if (!cJSON_AddItemToObject(root, \"mimeType\", mime)) goto blob_fail;", - " mime = NULL;", - " if (!cJSON_AddItemToObject(root, \"size\", size)) goto blob_fail;", - " size = NULL;", - " *out = root; return WF_OK;", - "blob_fail:", - " cJSON_Delete(type); cJSON_Delete(ref); cJSON_Delete(mime); cJSON_Delete(size);", - " cJSON_Delete(root); return WF_ERR_ALLOC;", "}", "", - "static WF_LEX_UNUSED wf_status wf_lex_cid_decode(cJSON *item, wf_lex_cid_link *out) {", - " cJSON *link = cJSON_IsObject(item) ? cJSON_GetObjectItemCaseSensitive(item, \"$link\") : NULL;", - " if (!cJSON_IsString(link) || !link->valuestring[0]) return WF_ERR_INVALID_ARG;", - " out->cid = wf_lex_strdup(link->valuestring); return out->cid ? WF_OK : WF_ERR_ALLOC;", "}", "", - "static WF_LEX_UNUSED wf_status wf_lex_bytes_decode(cJSON *item, wf_lex_bytes *out) {", - " cJSON *tag = cJSON_IsObject(item) ? cJSON_GetObjectItemCaseSensitive(item, \"$bytes\") : NULL;", - " if (!cJSON_IsString(tag)) return WF_ERR_INVALID_ARG;", - " size_t encoded = strlen(tag->valuestring);", - " if (encoded % 4 != 0 || encoded > (size_t)INT_MAX) return WF_ERR_INVALID_ARG;", - " size_t capacity = encoded / 4 * 3; uint8_t *data = capacity ? malloc(capacity) : NULL;", - " if (capacity && !data) return WF_ERR_ALLOC;", - " int decoded = encoded ? EVP_DecodeBlock(data, (const unsigned char *)tag->valuestring, (int)encoded) : 0;", - " if (decoded < 0) { free(data); return WF_ERR_INVALID_ARG; }", - " size_t padding = encoded && tag->valuestring[encoded - 1] == '=';", - " padding += encoded > 1 && tag->valuestring[encoded - 2] == '=';", - " out->data = data; out->length = (size_t)decoded - padding; return WF_OK;", "}", "", - "static WF_LEX_UNUSED wf_status wf_lex_blob_decode(cJSON *item, wf_lex_blob *out) {", - " if (!cJSON_IsObject(item)) return WF_ERR_INVALID_ARG;", - " cJSON *type = cJSON_GetObjectItemCaseSensitive(item, \"$type\");", - " cJSON *ref = cJSON_GetObjectItemCaseSensitive(item, \"ref\");", - " cJSON *mime = cJSON_GetObjectItemCaseSensitive(item, \"mimeType\");", - " cJSON *size = cJSON_GetObjectItemCaseSensitive(item, \"size\"); wf_lex_cid_link link = {0};", - " if (!cJSON_IsString(type) || strcmp(type->valuestring, \"blob\") != 0 ||", - " !cJSON_IsString(mime) || !wf_lex_json_integer(size, &out->size)) return WF_ERR_INVALID_ARG;", - " wf_status status = wf_lex_cid_decode(ref, &link); if (status != WF_OK) return status;", - " out->mime_type = wf_lex_strdup(mime->valuestring);", - " if (!out->mime_type) { free((void *)link.cid); return WF_ERR_ALLOC; }", - " out->cid = link.cid; return WF_OK;", "}", "", - ] - catalog = self.object_catalog() - encoders = self.encoder_objects() - unions = self.collect_unions() - for name in encoders: - out.append(f"static WF_LEX_UNUSED wf_status wf_lex_encode_{name}(const {name} *value, cJSON **out);") - for name in sorted(catalog): - out.append(f"static WF_LEX_UNUSED void wf_lex_clear_{name}({name} *value);") - out.append(f"static WF_LEX_UNUSED wf_status wf_lex_decode_{name}(cJSON *node, {name} *value);") - for name in sorted(unions): - nsid, schema = unions[name] - out.append(f"static WF_LEX_UNUSED void wf_lex_clear_{name}({name} *value);") - out.append(f"static WF_LEX_UNUSED wf_status wf_lex_decode_{name}(cJSON *node, {name} *value);") - if catalog or unions: - out.append("") - for name, (object_nsid, object_schema) in encoders.items(): - out.extend(self.emit_object_encoder(object_nsid, name, object_schema)) - for name in sorted(catalog): - object_nsid, object_schema = catalog[name] - out.extend(self.emit_object_decoder(object_nsid, name, object_schema)) - for name in sorted(unions): - nsid, schema = unions[name] - out.extend(self.emit_union_decoder(name, nsid, schema)) - out.extend(self.emit_union_clear(name, nsid, schema)) - for doc in self.docs: - nsid = doc["id"] - for def_name, definition in doc.get("defs", {}).items(): - if definition.get("type") not in ("query", "procedure"): - continue - base = type_name(nsid, def_name) - output_schema = definition.get("output", {}).get("schema") - if output_schema and output_schema.get("type") == "object": - out += [f"wf_status {base}_output_decode_json(", - f" const char *json, size_t length, {base}_output **out_value) {{", - " if (!json || !out_value) return WF_ERR_INVALID_ARG;", - " *out_value = NULL; cJSON *root = cJSON_ParseWithLength(json, length);", - " if (!root) return WF_ERR_INVALID_ARG;", - f" {base}_output *value = calloc(1, sizeof(*value));", - " if (!value) { cJSON_Delete(root); return WF_ERR_ALLOC; }", - f" wf_status status = wf_lex_decode_{base}_output(root, value);", - " cJSON_Delete(root);", - " if (status != WF_OK) { free(value); return status; }", - " *out_value = value; return WF_OK;", "}", "", - f"void {base}_output_free({base}_output *value) {{", - f" wf_lex_clear_{base}_output(value); free(value);", "}", ""] - schema = definition.get("input", {}).get("schema") - input_type = (self.json_input_type(nsid, base, schema) - if schema else None) - if input_type: - if schema.get("type") == "object": - encoder = base + "_input" - else: - resolved = (self.resolve_ref(nsid, schema["ref"]) - if schema.get("type") == "ref" else None) - encoder = (ref_type(nsid, schema["ref"]) - if resolved and resolved[1].get("type") == "object" - else None) - function = [f"wf_status {base}_input_encode_json(", - f" const {input_type} *value, char **out_json) {{", - " if (!value || !out_json) return WF_ERR_INVALID_ARG;", - " *out_json = NULL; cJSON *root = NULL;", - " wf_status status = WF_OK;", " (void)status;"] - if encoder: - function += [f" status = wf_lex_encode_{encoder}(value, &root);", - " if (status != WF_OK) return status;"] - else: - self.add_encoded_value(function, nsid, base, schema, "*value", - "root", "input", " ") - function += [" *out_json = cJSON_PrintUnformatted(root);", " cJSON_Delete(root);", - " return *out_json ? WF_OK : WF_ERR_ALLOC;"] - if any("goto invalid;" in line for line in function): - function += ["invalid:", " cJSON_Delete(root); return WF_ERR_INVALID_ARG;"] - if any("goto fail;" in line for line in function): - function += ["fail:", " cJSON_Delete(root); return WF_ERR_ALLOC;"] - if any("goto status_fail;" in line for line in function): - function += ["status_fail:", " cJSON_Delete(root); return status;"] - out += function + ["}", "", f"void {base}_json_free(char *json) {{ cJSON_free(json); }}", ""] - if definition.get("type") == "procedure" and not definition.get("parameters"): - if input_type: - out += [f"wf_status {base}_call(wf_xrpc_client *client,", - f" const {input_type} *input, wf_response *response) {{", - " if (!client || !input || !response) return WF_ERR_INVALID_ARG;", - " char *json = NULL;", - f" wf_status status = {base}_input_encode_json(input, &json);", - " if (status != WF_OK) return status;", - f' status = wf_xrpc_procedure(client, "{nsid}", json, response);', - " cJSON_free(json);", " return status;", "}", ""] - out += [f"wf_status {base}_call_auth(wf_auth_client *client,", - f" const {input_type} *input, wf_response *response) {{", - " if (!client || !input || !response) return WF_ERR_INVALID_ARG;", - " char *json = NULL;", - f" wf_status status = {base}_input_encode_json(input, &json);", - " if (status != WF_OK) return status;", - f' status = wf_auth_client_procedure(client, "{nsid}", json, response);', - " cJSON_free(json);", " return status;", "}", ""] - else: - out += [f"wf_status {base}_call(wf_xrpc_client *client, wf_response *response) {{", - " if (!client || !response) return WF_ERR_INVALID_ARG;", - f' return wf_xrpc_procedure(client, "{nsid}", NULL, response);', "}", ""] - out += [f"wf_status {base}_call_auth(wf_auth_client *client, wf_response *response) {{", - " if (!client || !response) return WF_ERR_INVALID_ARG;", - f' return wf_auth_client_procedure(client, "{nsid}", NULL, response);', "}", ""] - main = doc.get("defs", {}).get("main", {}) - if main.get("type") == "query": - base = type_name(nsid, "main") - params = main.get("parameters") - if not params: - out += [f"wf_status {base}_call(wf_xrpc_client *client, wf_response *response) {{", - " if (!client || !response) return WF_ERR_INVALID_ARG;", - f' return wf_xrpc_query(client, "{nsid}", NULL, response);', "}", ""] - out += [f"wf_status {base}_call_auth(wf_auth_client *client, wf_response *response) {{", - " if (!client || !response) return WF_ERR_INVALID_ARG;", - f' return wf_auth_client_query(client, "{nsid}", NULL, response);', "}", ""] - elif params.get("type") == "params": - props = params.get("properties", {}) - required = set(params.get("required", [])) - for wire_name, prop in props.items(): - kind = prop.get("type") - if kind == "array": - item_kind = prop.get("items", {}).get("type") - if item_kind not in ("string", "integer", "boolean"): - raise ValueError( - f"query parameter {wire_name} has unsupported array item type {item_kind}" - ) - elif kind not in ("string", "integer", "boolean"): - raise ValueError(f"query parameter {wire_name} has unsupported type {kind}") - # Build shared param-encoding body - call_body: list[str] = [ - " if (!params || !response) return WF_ERR_INVALID_ARG;", - " size_t encoded_capacity = 0, number_capacity = 0;"] - for wire_name, prop in props.items(): - field = self.field_name(params, wire_name) - condition = None if wire_name in required else f"params->has_{field}" - indent = " " - if condition: - call_body.append(f" if ({condition}) {{") - indent = " " - kind = prop.get("type") - if kind == "array": - item_kind = prop["items"]["type"] - call_body.append(f"{indent}if (params->{field}.count && !params->{field}.items) return WF_ERR_INVALID_ARG;") - if item_kind == "string": - call_body.append(f"{indent}for (size_t i = 0; i < params->{field}.count; ++i)") - call_body.append(f"{indent} if (!params->{field}.items[i]) return WF_ERR_INVALID_ARG;") - call_body.append(f"{indent}if (params->{field}.count > SIZE_MAX - encoded_capacity) return WF_ERR_INVALID_ARG;") - call_body.append(f"{indent}encoded_capacity += params->{field}.count;") - if item_kind == "integer": - call_body.append(f"{indent}if (params->{field}.count > SIZE_MAX - number_capacity) return WF_ERR_INVALID_ARG;") - call_body.append(f"{indent}number_capacity += params->{field}.count;") - else: - if kind == "string": - call_body.append(f"{indent}if (!params->{field}) return WF_ERR_INVALID_ARG;") - call_body.append(f"{indent}if (encoded_capacity == SIZE_MAX) return WF_ERR_INVALID_ARG;") - call_body.append(f"{indent}++encoded_capacity;") - if kind == "integer": - call_body.append(f"{indent}if (number_capacity == SIZE_MAX) return WF_ERR_INVALID_ARG;") - call_body.append(f"{indent}++number_capacity;") - if condition: - call_body.append(" }") - call_body += [" if (encoded_capacity > SIZE_MAX / sizeof(wf_xrpc_param) ||", - " number_capacity > SIZE_MAX / sizeof(char[32])) return WF_ERR_INVALID_ARG;", - " wf_xrpc_param *encoded = encoded_capacity ? calloc(encoded_capacity, sizeof(*encoded)) : NULL;", - " char (*number_values)[32] = number_capacity ? malloc(number_capacity * sizeof(*number_values)) : NULL;", - " if ((encoded_capacity && !encoded) || (number_capacity && !number_values)) {", - " free(encoded); free(number_values); return WF_ERR_ALLOC;", - " }", " size_t count = 0, number_count = 0;", - " (void)number_count;"] - for index, (wire_name, prop) in enumerate(props.items()): - field = self.field_name(params, wire_name) - condition = None if wire_name in required else f"params->has_{field}" - indent = " " - if condition: - call_body.append(f" if ({condition}) {{") - indent = " " - kind = prop.get("type") - if kind == "string": - value = f"params->{field}" - elif kind == "boolean": - value = f'(params->{field} ? "true" : "false")' - elif kind == "integer": - call_body.append(f'{indent}snprintf(number_values[number_count], sizeof(number_values[number_count]), "%' + '" PRId64, params->' + field + ");") - value = "number_values[number_count++]" - else: - item_kind = prop["items"]["type"] - call_body.append(f"{indent}for (size_t i = 0; i < params->{field}.count; ++i) {{") - if item_kind == "string": - value = f"params->{field}.items[i]" - elif item_kind == "boolean": - value = f'(params->{field}.items[i] ? "true" : "false")' - else: - call_body.append(f'{indent} snprintf(number_values[number_count], sizeof(number_values[number_count]), "%' + '" PRId64, params->' + field + ".items[i]);") - value = "number_values[number_count++]" - call_body.append(f'{indent} encoded[count++] = (wf_xrpc_param){{"{wire_name}", {value}}};') - call_body.append(f"{indent}}}") - if condition: - call_body.append(" }") - continue - call_body.append(f'{indent}encoded[count++] = (wf_xrpc_param){{"{wire_name}", {value}}};') - if condition: - call_body.append(" }") - # _call variant (wf_xrpc) - out += [f"wf_status {base}_call(wf_xrpc_client *client,", - f" const {base}_params *params, wf_response *response) {{", - " if (!client) return WF_ERR_INVALID_ARG;"] - out += call_body - out += [f' wf_status status = wf_xrpc_query_params(client, "{nsid}", encoded, count, response);', - " free(encoded); free(number_values); return status;", - "}", ""] - # _call_auth variant (wf_auth_client) - out += [f"wf_status {base}_call_auth(wf_auth_client *client,", - f" const {base}_params *params, wf_response *response) {{", - " if (!client) return WF_ERR_INVALID_ARG;"] - out += call_body - out += [f' wf_status status = wf_auth_client_query_params(client, "{nsid}", encoded, count, response);', - " free(encoded); free(number_values); return status;", - "}", ""] - return "\n".join(out) - - -def load(path: Path) -> dict[str, Any]: - with path.open(encoding="utf-8") as stream: - doc = json.load(stream) - if doc.get("lexicon") != 1 or not isinstance(doc.get("id"), str): - raise ValueError(f"{path}: expected a Lexicon 1 document with an id") - if not isinstance(doc.get("defs"), dict): - raise ValueError(f"{path}: expected a defs object") - return doc - - -def main(argv: list[str] | None = None) -> int: - parser = argparse.ArgumentParser(description=__doc__) - parser.add_argument("lexicons", nargs="+", type=Path) - parser.add_argument("-o", "--output", type=Path, - help="output header (stdout when omitted)") - parser.add_argument("--source-output", type=Path, - help="also write JSON codecs and XRPC wrappers") - parser.add_argument("--guard", default="WOLFRAM_GENERATED_LEXICONS_H") - parser.add_argument("--header-rel", default=None, - help="include path for the header (e.g. wolfram/foo.h)") - args = parser.parse_args(argv) - try: - generator = Generator([load(path) for path in args.lexicons], args.guard) - result = generator.generate() - if args.output: - args.output.parent.mkdir(parents=True, exist_ok=True) - args.output.write_text(result, encoding="utf-8") - else: - sys.stdout.write(result) - if args.source_output: - if not args.output: - raise ValueError("--source-output requires --output") - args.source_output.parent.mkdir(parents=True, exist_ok=True) - header_name = args.header_rel or args.output.name - args.source_output.write_text( - generator.generate_source(header_name), encoding="utf-8") - except (OSError, ValueError, json.JSONDecodeError) as error: - parser.error(str(error)) - return 0 - - -if __name__ == "__main__": - raise SystemExit(main())