diff --git a/docs/design/local-provider.md b/docs/design/local-provider.md deleted file mode 100644 index f7885dc0e..000000000 --- a/docs/design/local-provider.md +++ /dev/null @@ -1,128 +0,0 @@ -# Local provider - -Historical design note, July 2026: this document records the local-provider -implementation lode. Current provider routing is documented in -`docs/PROVIDERS.md`; the shipped system uses one active brain and does not use -automatic backup-provider fallback or tier-based provider routing. Body -references to `PROVIDER_DEFAULTS`, `get_backup_provider`, backup rewrites, or -tier keys are historical, not live implementation instructions. - -## D1 Provider identity and registry - -Decision: use the literal provider key `local`, registered as `PROVIDER_REGISTRY["local"] = "solstone.think.providers.local"`. Replace the current `ollama` provider entry; do not keep an alias. Use `PROVIDER_METADATA["local"] = {"label": "Local (on-device)", "env_key": ""}` with no `cogitate_runtime`, `cogitate_cli`, or install-command metadata. - -Justification: `local` is the owner-facing backend identity. The old `ollama` implementation depended on an external daemon and an OpenCode CLI; the new provider is a bundled llama-server loopback provider and should be selected through the same provider registry without compatibility shims. - -Implementation note: update `solstone/think/providers/__init__.py`. `build_provider_status` must remove the Ollama `/api/version` probe and add a `local` branch whose readiness is true only when the pinned llama-server binary is installed, the selected GGUF is present and sha256-verified, and the loopback daemon is healthy. Report distinct issue codes for `binary_missing`, `model_missing`, `server_unhealthy`, and `ram_insufficient`. - -## BYO OpenAI-compatible endpoint override - -Decision: `providers.local` may also carry a power-user endpoint override with non-secret `endpoint_url` and `served_model_id` fields plus secret `credential`. The override is active only when both non-secret fields are non-empty; otherwise `local` uses the bundled llama-server path described below. Non-confidential BYO overrides are governed by `providers.local.parallel_slots`, an integer at least 1 that defaults to 2 when absent or invalid. Bundled local ignores this key because bundled capacity is live server state. Confidential BYO overrides, identified by a present `services.confidential` block, also ignore this key and remain ungoverned. - -Justification: this keeps the owner-facing provider identity as `local` while allowing an explicitly configured OpenAI-compatible endpoint to receive local generate and cogitate requests. The top-level `env` subtree remains reserved for cloud provider API keys keyed by environment variable name. - -Implementation note: the settings surface owns writes through `POST /app/settings/api/local/endpoint` and `DELETE /app/settings/api/local/endpoint`. Public config and provider payloads expose only `{enabled, endpoint_url, served_model_id, credential_configured}` for the override and never return `providers.local.credential`. BYO readiness uses a shallow endpoint-root reachability probe; vision and JSON-schema support are classified as runtime contract failures when the endpoint rejects actual requests. - -## D2 OpenHands facade branch - -Decision: add a `local` branch to `solstone/think/providers/openhands.py::_build_llm`. The bundled kwargs are `model=f"openai/{local_model_id}"`, `base_url=f"http://127.0.0.1:{port}/v1"`, `api_key="EMPTY"`, `native_tool_calling=False`, `input_cost_per_token=0`, and `litellm_extra_body={"chat_template_kwargs": {"enable_thinking": False}}`. When the BYO endpoint override is active, the same branch uses `model=f"openai/{served_model_id}"`, `base_url=f"{endpoint_root}/v1"`, and the configured credential or `"EMPTY"`. - -Justification: llama-server exposes an OpenAI-compatible loopback API, but it does not provide native OpenHands tool calling or cloud API-key semantics. LiteLLM should treat it as an OpenAI-compatible custom endpoint with zero cost. - -Implementation note: call `solstone.think.providers.local_server.connect()` before constructing the bundled `LLM`; skip that call entirely when `resolve_local_endpoint().is_bundled` is false. Do not add `local` to `_GENERATE_MODULES` or `_API_KEY_ENV`; `local` generate is owned by `solstone.think.providers.local`, while `local` cogitate delegates through `openhands.run_cogitate`. Add `"local"` to `_KNOWN_MODEL_PREFIXES` so `_prefixed_model` can strip local-prefixed ids when a caller hands one to cloud code, but make the `local` branch return before `_MODEL_PREFIXES` lookup. `local.py::run_cogitate(config, on_event=None)` should import and call `openhands.run_cogitate(config, on_event, slot_lease=...)` after bundled readiness is established or the BYO endpoint is resolved. - -## D3 Model-id scheme - -Decision: replace `ollama-local/` with `local/`. Ship `local/qwen3.5-4b`. - -Justification: the provider prefix must match the new provider identity and should not preserve an Ollama-shaped namespace. Model ids stay stable across binary releases and map to pinned GGUF artifacts. - -Implementation note: in `solstone/think/models.py`, add `get_model_provider()` handling for `local/`, add `local` to the zero-cost provider set in `calc_token_cost`, and replace `PROVIDER_DEFAULTS["ollama"]` with `PROVIDER_DEFAULTS["local"]`. Define model specs with fields `model_id`, `repo`, `filename`, `revision`, `sha256`, `size_bytes`, `min_ram_bytes`, `mmproj_filename`, and `mmproj_sha256`. The current spec is `local/qwen3.5-4b`: repo `unsloth/Qwen3.5-4B-GGUF`, filename `Qwen3.5-4B-Q4_K_M.gguf`, revision `main`, sha256 `00fe7986ff5f6b463e62455821146049db6f9313603938a70800d1fb69ef11a4`, size `2740937888` bytes (~2.74 GB), min RAM `8 * 1024**3`, mmproj filename `mmproj-F16.gguf`, and mmproj sha256 `cd88edcf8d031894960bb0c9c5b9b7e1fea6ebee02b9f7ce925a00d12891f864`. - -## D4 First-slice GGUF default - -Decision: ship `Qwen3.5-4B Q4_K_M` as the practical default unified vision VLM. Set all local tiers to `LOCAL_MODEL = "local/qwen3.5-4b"`. - -Justification: the 4B GGUF is about 2.74 GB before runtime overhead, supports text and vision through the F16 mmproj, and fits the macOS arm64 and Linux CPU first slice with an 8 GiB RAM gate. - -Implementation note: set `min_ram_bytes` to `8 * 1024**3` for the Qwen3.5-4B spec. Default provider selection for all local tiers resolves to this model. - -## D5 Installer placement and state - -Decision: implement local installation with a local-specific core installer plus a settings bootstrap module: `solstone/think/providers/local_install.py` for artifact paths, pins, downloads, extraction, chmod, and verification; `solstone/apps/settings/local_bootstrap.py` for availability/progress state and HTTP route helpers. Do not extend `solstone/think/providers/bundled.py`. - -Justification: `bundled.py` delegates Codex binary installation to an external SDK and does not contain reusable tarball download/extract logic. Local needs two sha256-verified artifact classes: GitHub release tarballs for llama-server and Hugging Face LFS GGUF files, which matches the explicit verification pattern in `mlx_bootstrap.py`. - -Implementation note: define `LLAMA_SERVER_PINS` as `{artifact_key: {"release_tag": str, "filename": str, "sha256": str, "binary_name": "llama-server"}}`. Pin v1 to `aarch64-apple-darwin` -> `b9291`, `llama-b9291-bin-macos-arm64.tar.gz`, sha256 `0e985f87dd71f96a9cb9ebc3ad26f8388030342d000e7e82d4a38d14913373ff`; and `x86_64-unknown-linux-gnu` -> `b9291`, `llama-b9291-bin-ubuntu-vulkan-x64.tar.gz`, sha256 `7e3bf4202bedc71c2c9fbfbe02d10075b8d596bb963e7ab006663582dc2e92c2`. Store binaries under `/cache/providers/local/bin///` and models under `/cache/providers/local/models//`. Extract tarballs with path traversal checks, chmod the `llama-server` binary executable, and on macOS run `xattr -dr com.apple.quarantine ` best-effort after sha256 verification. Persist install state in `providers.bundled.local` because that key already represents bundled provider install state; include fields for `binary_artifact`, `binary_sha256`, `binary_path`, `model_id`, `model_path`, `model_sha256`, `state`, `last_transition_at`, and `install_error`. - -## D6 Linux Vulkan slice - -Decision: v1 ships macOS arm64 Metal and Linux x86_64 Vulkan. Linux requires a hardware Vulkan GPU; CPU/software Vulkan devices are rejected. - -Justification: recent llama.cpp releases expose Ubuntu CPU, Vulkan, SYCL, OpenVINO, and ROCm tarballs plus Windows CUDA zips, but no `ubuntu` or `linux` CUDA tarball equivalent was present in the checked release asset set. Building from source would violate the prebuilt zero-install-size constraint for this lode. - -Implementation note: do not add a CUDA artifact key in `LLAMA_SERVER_PINS` until an upstream Linux CUDA prebuilt tarball with a published sha256 is available. Add a completion note flagging this as a founder-facing decision point: v1 accepts no Linux CUDA acceleration path; the next decision is whether to wait for upstream, use a container, or own a build pipeline. - -## D7 Daemon ownership - -Decision: use lazy-start on first `local` generate or cogitate call, not a supervisor-registered always-on service. Implement `solstone/think/providers/local_server.py` as the daemon manager. - -Justification: there is no existing lazy subprocess-daemon pattern, and running llama-server all the time would impose memory cost on users who only occasionally choose the local backend. The provider can own a single loopback daemon and reuse existing process and port primitives without adding a new top-level service. - -Implementation note: the local server launch must use `[binary_path, "-m", model_path, "--alias", model_id, "--host", "127.0.0.1", "--port", str(port), "--jinja"]`, with `["--mmproj", mmproj_path]` appended only when the spec has an mmproj. Allocate the port with `find_available_port()`, persist it with `write_service_port("local", port)`, and reattach by reading `read_service_port("local")` plus probing `/health`. Poll `/health` until HTTP 200; treat HTTP 503 with "Loading model" as `loading`; timeout becomes `model_load_timeout`. Enforce one instance with a process-local lock plus a file lock at `/health/local-server.lock` so concurrent first calls do not double-spawn. Provide `stop()` using `ManagedProcess.terminate()`. New state names are `idle`, `starting`, `loading`, `ready`, `failed`, and `stopped`; install/bootstrap keeps the MLX-style `idle`, `downloading`, `verifying`, `installed`, `failed`. - -## D8 Fallback opt-out - -Decision: `local` and the no-thinking-engine state must never silently fall back to a cloud provider for either `generate` or `cogitate`. - -Justification: selecting `local` is an explicit privacy and locality choice, and choosing no thinking engine is an explicit deferred state. Silent cloud fallback would violate that intent and hide fixable local installation, runtime, or setup failures. - -Implementation note: in `solstone/think/models.py::get_backup_provider`, return `None` when the resolved primary provider is `local` or `none` for all agent types. This makes the preflight swap in `talents.py` and the on-failure cogitate fallback no-ops because both paths already require a non-empty backup. On local failure, emit or surface a recovery reason instead of setting `config["fallback_from"]`: `binary_missing`, `model_missing`, `server_crashed`, `model_load_failed`, `port_conflict`, or `ram_insufficient`. When no thinking engine is chosen, surface `thinking_engine_not_chosen` instead of consulting a backup. - -## D9 Migration command - -Decision: implement the migration as a manual settings maintenance module at `solstone/apps/settings/maint/_migrate_ollama_to_local.py`, with a CLI wrapper `sol call settings providers migrate-ollama-to-local [--commit] [--json]`. - -Justification: app `maint/` scripts are auto-discovered and run without flags, but this data-format migration must dry-run by default and require `--commit` per L5. The leading underscore keeps the shared migration implementation under settings maint while preventing accidental automatic execution. - -Implementation note: the dry-run prints a JSON-compatible report of every planned rewrite and exits without writing. `--commit` rewrites `providers.generate.provider`, `providers.generate.backup`, `providers.cogitate.provider`, and `providers.cogitate.backup` values from `ollama` to `local`; moves `providers.models.ollama` to `providers.models.local`, filling missing local tier keys and reporting conflicts; moves `providers.auth.ollama` to `providers.auth.local`; moves `providers.key_validation.ollama` to `providers.key_validation.local`; updates every `providers.contexts.*.provider == "ollama"` to `local`; and rewrites known Ollama-local model strings to `local/qwen3.5-4b`, with other `ollama-local/` values becoming `local/` plus an `unsupported_model` warning. Do not touch `providers.api_keys`. Do not rewrite `providers.contexts` map keys because those are owner-defined context patterns; only rewrite their values. Running twice must produce an empty rewrite report. - -## D10 Failure taxonomy and copy - -Decision: use one local failure taxonomy shared by provider status, bootstrap routes, provider CLI, and UI. - -Justification: local has more failure modes than a cloud API key provider, but the UI should still present the same recovery vocabulary as MLX bootstrap and bundled providers: install, verify, retry, and choose another configured provider manually. - -Implementation note: map failures as follows. `binary_missing`: action `install_local_runtime`, copy "Local runtime is not installed." `gguf_missing`: action `install_local_model`, copy "Local model files are not installed." `server_crashed`: action `restart_local_runtime`, copy "Local runtime stopped unexpectedly." `model_load_failed`: action `retry_model_load`, copy "Local model could not be loaded." `model_load_timeout`: action `retry_model_load`, copy "Local model is still loading." `port_conflict`: action `restart_local_runtime`, copy "Local runtime port is unavailable." `ram_insufficient`: action `choose_smaller_model`, copy "This computer does not have enough memory for the selected local model." No message should offer automatic cloud fallback. - -## D11 Settings, CLI, and UI surface - -Decision: rename all Ollama settings surfaces to Local and add a Local bootstrap region modeled on MLX. - -Justification: the old UI was an external Ollama/OpenCode readiness check. Local needs install, model availability, and daemon readiness controls that match the new bundled runtime. - -Implementation note: replace `/api/providers/ollama/status` and `get_ollama_provider_status` with `/api/providers/local/status` and `get_local_provider_status`. Add `/api/local/models`, `/api/local/availability`, `/api/local/bootstrap`, and `/api/local/bootstrap/status` mirroring `/api/mlx/*`. In `workspace.html`, rename `#ollamaCogitateStatus*` to `#localCogitateStatus*`, `.ollama-command-row` to `.local-command-row`, `ollamaStatusLoading` to `localStatusLoading`, `OLLAMA_OPENCODE_*` to `LOCAL_RUNTIME_*`, `renderOllamaCogitateStatus` to `renderLocalCogitateStatus`, and `recheckOllamaStatus` to `recheckLocalStatus`. Add `.local-bootstrap-region` and `.local-progress-shell` modeled on `.mlx-bootstrap-region` and `.mlx-progress-shell`. Update `chat_reasons.py` and `chat_reasons.js` display names from `"ollama": "Ollama"` to `"local": "Local"`. In `providers_cli.py`, include `local` in the OpenHands-backed cogitate check set, remove Ollama-specific skip copy, and emit Local readiness copy from the D10 taxonomy. - -## D12 Importability and zero-fetch rule - -Decision: `solstone.think.providers.local`, `solstone.think.providers.local_server`, `solstone.think.providers.local_install`, and `solstone.think.providers.openhands` must import without llama-server, GGUF files, llama.cpp, OpenHands, LiteLLM, or Hugging Face network access present. - -Justification: provider registration and settings pages must work before the owner enables the local backend. Imports must not trigger downloads, subprocess starts, or optional SDK imports. - -Implementation note: keep OpenHands/LiteLLM imports inside `openhands` functions, keep Hugging Face and HTTP download imports inside bootstrap/install functions, and keep daemon startup inside `start_local_server()` / `local_server.connect()` paths. Nothing fetches or installs until the owner runs the Local bootstrap/install action. - -## Deferred decisions and completion notes - -Decision: record these completion notes with the implementation PR. - -Justification: they capture resolved scope boundaries and the one founder-facing risk that should not be rediscovered during implementation. - -Implementation note: B2 is resolved by shipping all local tiers on `local/qwen3.5-4b` with an 8 GiB RAM gate. B3 is resolved by supervisor-owned local startup when the local provider is selected. B4 is deferred to v1.1: add a non-gating warning constant named `LOCAL_BOOTSTRAP_DISK_WARNING_THRESHOLD_BYTES` before downloading large models. B5 is resolved by the Qwen3.5-4B unified VLM plus `mmproj-F16.gguf`; image contents passed to local generate use the OpenAI-compatible image-url path. D6 remains the explicit CUDA decision point: v1 ships no Linux CUDA slice. - -## Implementation sequence - -Decision: implement in this order: provider identity, facade, models, fallback; then install, daemon, bootstrap; then settings UI and CLI; then migration; then tests, baselines, and docs. - -Justification: provider identity and model resolution are prerequisites for every route and test. Daemon/bootstrap depend on model specs and install paths. UI and migration should target stable provider/status APIs. - -Implementation note: first update `providers/__init__.py`, `models.py`, `openhands.py`, `local.py`, and fallback tests. Next add `local_install.py`, `local_server.py`, and `apps/settings/local_bootstrap.py`. Then update settings routes, workspace IDs/classes/JS, `providers_cli.py`, and chat reason display names. Then add the manual migration wrapper and update fixtures/baselines. Finish with provider, fallback, migration, settings route, workspace, and baseline tests. diff --git a/solstone/apps/thinking/tests/test_providers_payload_extended.py b/solstone/apps/thinking/tests/test_providers_payload_extended.py index 9ab891234..0e55470c2 100644 --- a/solstone/apps/thinking/tests/test_providers_payload_extended.py +++ b/solstone/apps/thinking/tests/test_providers_payload_extended.py @@ -462,7 +462,13 @@ def test_providers_payload_omits_bundled_block(settings_client): "cogitate_ready", "issues", } - assert provider_status["local"]["cogitate_cli"] == "llama-server" + assert set(provider_status["local"]) == { + "configured", + "selected", + "generate_ready", + "cogitate_ready", + "issues", + } assert REMOVED_PROVIDER not in payload _assert_install_status(payload["local"]) @@ -1083,8 +1089,6 @@ def test_providers_payload_local_status_uses_endpoint_readiness_under_byo( "selected": False, "generate_ready": True, "cogitate_ready": True, - "cogitate_cli": None, - "cogitate_cli_found": False, "issues": [], } @@ -1125,8 +1129,6 @@ def test_get_providers_uses_state_local_status(settings_client, monkeypatch): "selected": True, "generate_ready": True, "cogitate_ready": True, - "cogitate_cli": "llama-server", - "cogitate_cli_found": True, "issues": ["sentinel"], } monkeypatch.setattr( @@ -1236,8 +1238,6 @@ def test_get_providers_ai_readiness_degrades_without_changing_status_payload( "selected": True, "generate_ready": True, "cogitate_ready": True, - "cogitate_cli": "llama-server", - "cogitate_cli_found": True, "issues": ["sentinel"], } monkeypatch.setattr( diff --git a/solstone/think/providers/__init__.py b/solstone/think/providers/__init__.py index d41b23388..9353d592a 100644 --- a/solstone/think/providers/__init__.py +++ b/solstone/think/providers/__init__.py @@ -148,7 +148,6 @@ def build_provider_status( ------- Dict[str, Dict[str, Any]] Keyed by provider name. Each entry has readiness fields and issues. - Local readiness also includes cogitate_cli and cogitate_cli_found. """ if providers_list is None: providers_list = get_provider_list() diff --git a/solstone/think/providers/state.py b/solstone/think/providers/state.py index b47caa918..0798b9af3 100644 --- a/solstone/think/providers/state.py +++ b/solstone/think/providers/state.py @@ -285,7 +285,7 @@ def local_runtime_ready(model_id: str | None = None) -> bool: def local_status_dict() -> dict: - """Build the legacy local provider status dict.""" + """Build the local provider status dict.""" from solstone.think.models import is_local_provider_needed from solstone.think.providers.local_endpoint import ( probe_local_endpoint, @@ -301,8 +301,6 @@ def local_status_dict() -> dict: "selected": selected, "generate_ready": reachable, "cogitate_ready": reachable, - "cogitate_cli": None, - "cogitate_cli_found": False, "issues": [] if reachable else ["local_endpoint_unreachable"], } @@ -320,8 +318,6 @@ def local_status_dict() -> dict: "selected": False, "generate_ready": False, "cogitate_ready": False, - "cogitate_cli": "mlx-vlm", - "cogitate_cli_found": runtime_available, "issues": [], } @@ -340,8 +336,6 @@ def local_status_dict() -> dict: "selected": True, "generate_ready": ready, "cogitate_ready": ready, - "cogitate_cli": "mlx-vlm", - "cogitate_cli_found": runtime_available, "issues": issues, } @@ -358,8 +352,6 @@ def local_status_dict() -> dict: "selected": False, "generate_ready": False, "cogitate_ready": False, - "cogitate_cli": "llama-server", - "cogitate_cli_found": binary_installed, "issues": [], } @@ -387,8 +379,6 @@ def local_status_dict() -> dict: "selected": True, "generate_ready": ready, "cogitate_ready": ready, - "cogitate_cli": "llama-server", - "cogitate_cli_found": binary_installed, "issues": issues, } diff --git a/solstone/think/providers_cli.py b/solstone/think/providers_cli.py index 2918e5723..437f9da84 100644 --- a/solstone/think/providers_cli.py +++ b/solstone/think/providers_cli.py @@ -122,13 +122,13 @@ async def _check_cogitate( state.readiness_for_provider("local", "cogitate").reason_code or "unknown" ) - if not status.get("cogitate_cli_found"): - from solstone.think.providers.local_endpoint import ( - resolve_local_endpoint, - ) + from solstone.think.providers.local_endpoint import resolve_local_endpoint - if not resolve_local_endpoint().is_bundled: - return "skip", _local_readiness_message(status), reason_code + if not resolve_local_endpoint().is_bundled: + return "skip", _local_readiness_message(status), reason_code + if {"runtime_missing", "binary_missing", "model_missing"}.intersection( + status.get("issues", []) + ): from solstone.think.providers import local_install return ( diff --git a/tests/baselines/api/thinking/providers.json b/tests/baselines/api/thinking/providers.json index ead61d3fb..60a3bc034 100644 --- a/tests/baselines/api/thinking/providers.json +++ b/tests/baselines/api/thinking/providers.json @@ -198,8 +198,6 @@ "provider": "google" }, "local": { - "cogitate_cli": "llama-server", - "cogitate_cli_found": false, "cogitate_ready": false, "configured": false, "generate_ready": false, diff --git a/tests/test_google_sdk.py b/tests/test_google_sdk.py index debead6a1..d866274c1 100644 --- a/tests/test_google_sdk.py +++ b/tests/test_google_sdk.py @@ -8,8 +8,7 @@ import pytest from solstone.think.providers import PROVIDER_METADATA, build_provider_status -def test_google_provider_metadata_has_no_cogitate_cli() -> None: - assert "cogitate_cli" not in PROVIDER_METADATA["google"] +def test_google_provider_metadata_has_no_runtime_adapter() -> None: assert "cogitate_runtime" not in PROVIDER_METADATA["google"] @@ -51,6 +50,4 @@ def test_google_provider_status_uses_managed_key_only( [{"name": "google", "env_key": "GOOGLE_API_KEY"}], )["google"] - assert "cogitate_cli" not in status - assert "cogitate_cli_found" not in status assert status == expected_status diff --git a/tests/test_local.py b/tests/test_local.py index 9d5491437..55c1ed754 100644 --- a/tests/test_local.py +++ b/tests/test_local.py @@ -2397,7 +2397,6 @@ def test_build_provider_status_local_readiness(monkeypatch): assert status["configured"] is True assert status["generate_ready"] is True assert status["cogitate_ready"] is True - assert status["cogitate_cli"] == "llama-server" assert status["issues"] == [] @@ -2527,8 +2526,6 @@ def test_local_provider_status_carries_install_hint_substring(monkeypatch): assert status["configured"] is False assert status["generate_ready"] is False assert status["cogitate_ready"] is False - assert status["cogitate_cli"] == "llama-server" - assert status["cogitate_cli_found"] is False assert status["issues"] == [ "binary_missing", "model_missing", diff --git a/tests/test_provider_state.py b/tests/test_provider_state.py index d77646f72..3ae2765a4 100644 --- a/tests/test_provider_state.py +++ b/tests/test_provider_state.py @@ -862,8 +862,6 @@ def test_local_status_dict_byo( "selected": selected_config["providers"]["active"]["provider"] == "local", "generate_ready": reachable, "cogitate_ready": reachable, - "cogitate_cli": None, - "cogitate_cli_found": False, "issues": expected_issues, } @@ -887,14 +885,10 @@ def test_local_status_dict_darwin(monkeypatch): "selected", "generate_ready", "cogitate_ready", - "cogitate_cli", - "cogitate_cli_found", "issues", } assert status["generate_ready"] is True assert status["cogitate_ready"] is True - assert status["cogitate_cli"] == "mlx-vlm" - assert status["cogitate_cli_found"] is True assert status["issues"] == [] diff --git a/tests/test_providers_check.py b/tests/test_providers_check.py index c74d09113..66db3940e 100644 --- a/tests/test_providers_check.py +++ b/tests/test_providers_check.py @@ -320,7 +320,8 @@ def test_check_cogitate_local_missing_runtime_names_local_install_hint(monkeypat "_provider_status", lambda _name: { "configured": True, - "cogitate_cli_found": False, + "cogitate_ready": False, + "issues": ["model_missing"], }, ) monkeypatch.setattr( @@ -349,7 +350,6 @@ def test_check_cogitate_local_endpoint_unreachable_uses_endpoint_reason(monkeypa "_provider_status", lambda _name: { "configured": True, - "cogitate_cli_found": False, "cogitate_ready": False, "issues": ["local_endpoint_unreachable"], }, diff --git a/tests/test_talents_process.py b/tests/test_talents_process.py index 4fddb6efd..b210f82e6 100644 --- a/tests/test_talents_process.py +++ b/tests/test_talents_process.py @@ -60,7 +60,6 @@ def test_talent_main_sigterm_exits_without_cancelled_traceback(tmp_path): providers.PROVIDER_METADATA["test"] = {{ "label": "Test", "env_key": "", - "cogitate_cli": "", }} sys.modules["solstone_test_provider"] = fake_provider """ diff --git a/tests/test_thinking_call_parity.py b/tests/test_thinking_call_parity.py index a876efc0c..e5571e006 100644 --- a/tests/test_thinking_call_parity.py +++ b/tests/test_thinking_call_parity.py @@ -326,7 +326,6 @@ def test_providers_show_human_and_set_errors( "local": { "generate_ready": False, "cogitate_ready": False, - "cogitate_cli": "llama-server", "issues": ["binary_missing"], }, "openai": {"generate_ready": True, "cogitate_ready": True, "issues": []}, diff --git a/tests/verify_api.py b/tests/verify_api.py index cb20e4a3b..5f8756a21 100644 --- a/tests/verify_api.py +++ b/tests/verify_api.py @@ -438,12 +438,10 @@ def normalize(data: Any, journal_path: str) -> Any: status["cogitate_ready"] = False status["issues"] = [f"{env_keys[_name]} not set"] continue - if isinstance(status, dict) and "cogitate_cli" in status: - status["cogitate_cli_found"] = False + if _name == "local" and isinstance(status, dict): status["cogitate_ready"] = False status["configured"] = False status["generate_ready"] = False - cli = status.get("cogitate_cli", "") issues = [ i for i in status.get("issues", []) @@ -451,28 +449,24 @@ def normalize(data: Any, journal_path: str) -> Any: and "not set" not in i and "not reachable" not in i ] - if _name == "local": - local_issues = [ - i - for i in issues - if i - in { - "binary_missing", - "gpu_unavailable", - "model_missing", - "ram_insufficient", - "server_unhealthy", - } - ] - for local_issue in ("binary_missing", "model_missing"): - if local_issue not in local_issues: - local_issues.append(local_issue) - local_issues.append("run `journal install-provider local`") - status["issues"] = sorted(local_issues) - continue - if cli and _name not in {"anthropic", "openai", "google"}: - issues.append(f"{cli} CLI not found on PATH") - status["issues"] = sorted(issues) + local_issues = [ + i + for i in issues + if i + in { + "binary_missing", + "gpu_unavailable", + "model_missing", + "ram_insufficient", + "server_unhealthy", + } + ] + for local_issue in ("binary_missing", "model_missing"): + if local_issue not in local_issues: + local_issues.append(local_issue) + local_issues.append("run `journal install-provider local`") + status["issues"] = sorted(local_issues) + continue # Normalize env-dependent API key presence if key in ("api_keys", "runtime_env"): for k in result: