From 89df12f6bf53b4e56cb6c3f53279361de9055e1e Mon Sep 17 00:00:00 2001 From: aboba Date: Thu, 6 Aug 2026 03:39:17 +0300 Subject: [PATCH] close two rollback paths: firmware-owned dmem.bin, and remove.sh hanging Booter payload override moved out of the firmware package's namespace. cmpunlock.c read it from /lib/firmware/nvidia/ga100/gsp/dmem.bin, which belongs to nvidia-gpu-firmware (linux-firmware on Fedora). That directory already accumulates per-driver blobs on every firmware release, so an update landing a dmem.bin there would have been preferred over the built-in payload and quietly broken the unlock. Pinning that package is not an option - it is a linux-firmware subpackage and holding it back wedges system upgrades on a dependency conflict. Reading from /var/lib/cmpunlocker/dmem.bin instead makes the question moot. The GSP firmware the driver actually loads lives in the versioned directory from the driver package, which is pinned. remove.sh no longer swaps the running driver by default. Loading the stock nvidia-drm against a CMP wedges the machine: traced it to that exact modprobe, after which the kernel still answers pings while userspace stops making progress. Nothing needed the swap - the modules are already off disk, so the next boot picks up the stock driver on its own, and an uninstall ends in a reboot anyway. Behind --reload for anyone who wants it, now guarded with timeouts. Dropped the rmmod -f fallback: forcing out a module that is still referenced can take the kernel with it, which is a bad trade when a reboot finishes the job. remove.sh also syncs after depmod, matching build.sh. Without it a power cut between depmod and the next flush leaves a zero-length modules.dep, and then nothing resolves on the next boot - not the NIC driver, not storage. Hit this three times in testing via hard power-off. --- README.md | 8 ++++- driver/src/cmpunlock.c | 16 ++++++++- remove.sh | 77 +++++++++++++++++++++++++++++++----------- 3 files changed, 79 insertions(+), 22 deletions(-) diff --git a/README.md b/README.md index bd85b35..cdf1572 100644 --- a/README.md +++ b/README.md @@ -87,7 +87,9 @@ Two more things keep the stock driver from winning: - `/etc/depmod.d/cmpunlocker.conf` makes the patched modules outrank the stock ones. The distro driver is rebuilt on every kernel update too, into `extra/` (akmod) or `updates/dkms/` (dkms), right next to ours. - `nvidia-fallback.service` is masked and nouveau is blacklisted, so a driver that fails to load does not hand the card to nouveau. -**The NVIDIA packages are pinned to their installed version.** A driver upgrade past the versions in `driver/VERSION` makes every later rebuild fail, which is the rollback this is meant to prevent. GPU *firmware* packages are deliberately left unpinned — they belong to `linux-firmware` and holding them back can wedge system upgrades. +**The NVIDIA packages are pinned to their installed version.** A driver upgrade past the versions in `driver/VERSION` makes every later rebuild fail, which is the rollback this is meant to prevent. This covers the GSP firmware the driver actually loads, from `/lib/firmware/nvidia//`, which ships in the driver package itself. + +GPU *firmware* packages (`nvidia-gpu-firmware` and friends) are deliberately left unpinned — they belong to `linux-firmware`, and holding them back can wedge system upgrades on a dependency conflict. They are also no longer able to affect the unlock: the optional Booter payload override is read from `/var/lib/cmpunlocker/dmem.bin` rather than from `/lib/firmware/nvidia/ga100/gsp/`, a directory that firmware updates add files to and that no amount of pinning could safely protect. ```bash sudo ./install.sh --no-pin # allow driver upgrades, accept the risk @@ -177,6 +179,10 @@ sudo ./remove.sh --yes Then perform a cold reboot (full power off, then boot). +This removes the patched modules from disk, undoes the kernel-update hooks, releases the package pin, and rebuilds the initramfs. The driver already running in memory is left alone — the card comes up on the stock driver at the next boot, which is the safe order. + +`--reload` swaps the running driver for the stock one immediately instead of waiting for the reboot. It is off by default because loading the stock `nvidia-drm` against a CMP 170HX can wedge the machine: the card has no usable display engine, and the kernel keeps answering pings while userspace stops making progress. There is no reason to take that risk during an uninstall you are going to reboot from anyway. + ## Community Join our [Discord community](https://discord.gg/CdHSakKSFv) to discuss with other users. diff --git a/driver/src/cmpunlock.c b/driver/src/cmpunlock.c index b2b9c2f..1a70222 100644 --- a/driver/src/cmpunlock.c +++ b/driver/src/cmpunlock.c @@ -80,7 +80,21 @@ /* Booter payload geometry. */ #define CMP_SIGNATURE_SIZE 0x0000f800ULL #define CMP_PAYLOAD_FILL_DWORD 0x000004a7U -#define CMP_DMEM_PATH "/lib/firmware/nvidia/ga100/gsp/dmem.bin" +/* + * Optional override for the Booter payload. + * + * This used to read /lib/firmware/nvidia/ga100/gsp/dmem.bin, which belongs to + * the distro's GPU firmware package (nvidia-gpu-firmware on Fedora, part of + * linux-firmware). That directory already collects per-driver blobs on every + * firmware release, so a future update dropping a dmem.bin there would have + * been picked up in preference to the built-in payload and quietly broken the + * unlock - with no way to prevent it short of holding back linux-firmware, + * which is not something worth doing to a system. + * + * The path now lives in cmpunlocker's own directory, where nothing else + * writes. Absent - which is the normal case - the payload is generated below. + */ +#define CMP_DMEM_PATH "/var/lib/cmpunlocker/dmem.bin" /* Unlocked framebuffer sizes. */ #define CMP_FB_BYTES_8GB 0x0000001000000000ULL /* 64GB */ diff --git a/remove.sh b/remove.sh index f9d7855..fc7dca6 100755 --- a/remove.sh +++ b/remove.sh @@ -20,19 +20,32 @@ echo -e "${CYAN}║ cmpunlocker ║${NC}" echo -e "${CYAN}╚════════════════════════════════════════╝${NC}" echo "" -if [[ "${1:-}" != "--yes" && "${1:-}" != "-y" ]]; then +CONFIRMED=0 +RELOAD_DRIVER=0 +for arg in "$@"; do + case "${arg}" in + --yes|-y) CONFIRMED=1 ;; + --reload) RELOAD_DRIVER=1 ;; + esac +done + +if (( CONFIRMED == 0 )); then warn "This removes cmpunlocker patched kernel modules:" echo " - Removes /lib/modules/*/updates/cmpunlocker/" echo " - Rebuilds initramfs" - echo " - Reloads stock NVIDIA modules (brief display interruption)" echo " - Restores the pre-install kernel command line (reverts IOMMU changes)" echo " - Removes the kernel-update hooks and the boot-time rebuild service" echo " - Releases the NVIDIA package version pin" echo " - Unmasks nvidia-fallback.service and unblocks nouveau" echo "" + echo " The driver already in memory is left alone — reboot to finish." echo " Logs under /var/log/cmpunlocker/ are kept." echo "" echo "Run: sudo ./remove.sh --yes" + echo "" + echo " --reload also swap the running driver for the stock one now," + echo " instead of at the next boot. On a CMP this can hang the" + echo " machine, so it is off by default — see README." exit 1 fi @@ -114,6 +127,13 @@ for mod_dir in /lib/modules/*/updates/cmpunlocker; do kernel="$(basename "$(dirname "$(dirname "${mod_dir}")")")" rm -rf "${mod_dir}" depmod -a "${kernel}" 2>/dev/null || true + # + # depmod's output sits in the page cache until something flushes it. + # A power cut or hard reset before that leaves a zero-length + # modules.dep, and then nothing resolves on the next boot - not the + # NIC driver, not storage. Cheap insurance against an unbootable box. + # + sync ok "Removed patched modules for kernel ${kernel}" mod_removed=$((mod_removed + 1)) kernels_touched+=("${kernel}") @@ -136,9 +156,26 @@ if [[ ${#kernels_touched[@]} -gt 0 ]]; then ok "initramfs rebuilt" fi -info "Reloading stock NVIDIA driver..." -if lsmod | grep -q '^nvidia'; then - warn "Unloading NVIDIA modules (display may flicker)" +# +# Swapping the running driver is opt-in, because loading the stock one on a +# CMP 170HX can wedge the machine: the card has no usable display engine, and +# nvidia-drm binding to it hangs in a way that leaves the kernel answering +# pings while userspace stops making progress. Nothing needs the swap either - +# the files are already gone, so the next boot comes up on the stock driver by +# itself, and a reboot is the documented last step of an uninstall anyway. +# +if (( RELOAD_DRIVER == 0 )); then + if lsmod | grep -q '^nvidia'; then + info "Leaving the running driver in place" + echo " The patched modules are removed from disk, but the copy already in" + echo " memory keeps running until you reboot. That is the safe order." + echo " ./remove.sh --yes --reload swaps it now instead, at the risk of" + echo " hanging the machine on a CMP." + fi +else + info "Swapping the running driver for the stock one (--reload)..." + warn "This can hang on a CMP 170HX. Reboot instead if it does not return." + for svc in gdm3 sddm lightdm display-manager; do systemctl stop "${svc}" 2>/dev/null || true done @@ -146,24 +183,26 @@ if lsmod | grep -q '^nvidia'; then killall -9 Xorg Xwayland nvidia-persistenced 2>/dev/null || true sleep 1 + # + # No rmmod -f fallback: forcing a module out from under a driver that is + # still referenced is documented as able to take the kernel down with it, + # which is a poor trade when a reboot finishes the job cleanly. + # for mod in nvidia_drm nvidia_uvm nvidia_modeset nvidia; do - modprobe -r "${mod}" 2>/dev/null || true + timeout 30 modprobe -r "${mod}" 2>/dev/null || true done sleep 1 if lsmod | grep -q '^nvidia'; then - for mod in nvidia_uvm nvidia_drm nvidia_modeset nvidia; do - rmmod -f "${mod}" 2>/dev/null || true - done - fi - - if modprobe nvidia 2>/dev/null; then - modprobe nvidia-modeset 2>/dev/null || true - modprobe nvidia-uvm 2>/dev/null || true - modprobe nvidia-drm 2>/dev/null || true - ok "Stock NVIDIA driver reloaded" + warn "Some NVIDIA modules are still in use — reboot to finish cleanup" + elif timeout 60 modprobe nvidia 2>/dev/null; then + timeout 30 modprobe nvidia-modeset 2>/dev/null || true + timeout 30 modprobe nvidia-uvm 2>/dev/null || true + timeout 30 modprobe nvidia-drm 2>/dev/null || \ + warn "nvidia-drm did not load — harmless on a headless card" + ok "Stock NVIDIA driver loaded" else - warn "Could not reload NVIDIA driver — reboot to finish cleanup" + warn "Stock driver did not load — reboot to finish cleanup" fi for svc in gdm3 sddm lightdm display-manager; do @@ -172,14 +211,12 @@ if lsmod | grep -q '^nvidia'; then break fi done -else - warn "NVIDIA modules not loaded — skipping driver reload" fi echo "" ok "cmpunlocker removed" echo "Log saved to: ${LOG_FILE}" echo "" -echo "If the GPU or display is not working, reboot:" +echo "Reboot to finish — the card comes back up on the stock driver:" echo -e " ${CYAN}sudo reboot${NC}" echo "" -- 2.51.2