How to run WinRE Manager on a single machine or across a managed fleet.
There are two ways to deploy WinRE Manager. Pick the one that matches your situation.
For a single broken machine, run the production script once from an elevated PowerShell prompt. It inspects the machine, repairs the recovery environment, writes a state file, and exits. There is no installer, no service, and no configuration file to maintain.
As of v48 patch 2, the script refuses to run unelevated in milliseconds. The refusal is a single FATAL log line (FATAL: WinRE Manager requires an elevated (Administrator) PowerShell session.) and exit code 3, before the program lock, before hardware probes, and before any network fetch. The guard creates the log directory itself so its FATAL message has somewhere to be written. Before v48 patch 2, an unelevated launch proceeded through hardware detection, manifest fetch, VMD detection, and the multi-minute GitHub base-WIM download before failing at Mount-WindowsImage with The requested operation requires elevation. — a failure mode that wasted roughly four minutes per run in the 2026-10-05 field log. Both the interactive Administrator session and the SYSTEM context used by the scheduled task pass the guard.
# Elevated PowerShell (Run as Administrator)
powershell -ExecutionPolicy Bypass -File .\scripts\WinRE.ps1 -DryRun # walk the flow, log every decision, change nothing
powershell -ExecutionPolicy Bypass -File .\scripts\WinRE.ps1 # actually deploy
The -ExecutionPolicy Bypass -File form applies the bypass only to the child process; it does not change the machine’s execution policy.
If you want to inspect the machine first, run the read-only harness (Test-WinRE.ps1). It requires no elevation and modifies nothing. See the README for the harness menu options.
For a single-machine deployment, the rest of this document is reference material — the exit-code table, the Audit Mode precondition, and the Device Encryption policy all apply to a one-off run exactly as they apply to a fleet run. You can skip the scheduled-task sections.
For a managed fleet, deploy the script as a scheduled task running as NT AUTHORITY\SYSTEM, triggered at boot and weekly. This is the recommended model. Every machine keeps its recovery environment healthy on its own; a healthy machine’s fast path is dominated by Windows’ own CIM and PnP enumeration (typically 10–20 seconds), and a full repair runs only when something has changed.
The sections below this point cover the fleet deployment in detail: how to create the task, where to put the script, how to handle the Audit Mode precondition, how to interpret the exit codes, and how to roll out in rings.
Run WinRE.ps1 as NT AUTHORITY\SYSTEM from a scheduled task, triggered at boot and weekly. The script is fully idempotent; running it on a healthy machine costs a few seconds of CIM and PnP enumeration and a read-only scan.
Do not run it from a user logon script. It modifies partition tables and the BitLocker state of the target recovery partition; those operations require SYSTEM.
The script’s design is organized around four rules, in this order — see docs/architecture.md for the full hierarchy.
reagentc /disable → reagentc /enable window.Rules 1–3 are invariants: no deployment scenario should weaken them. Rule 4 is the working rule: the discipline that makes the fast path and the enable-only path free of side effects, and that makes the destructive path fail safe.
Each rule has a direct operational translation for a fleet operator.
The script’s destructive sequence is protected by a read-only plan, a fail-closed layout assertion, and a reversible shrink window. It does not need to be bullet-proof to be deployed; it needs to be canaried so that any regression is caught on one machine, not a thousand.
Before pushing to a fleet, run on one machine end-to-end:
powershell -ExecutionPolicy Bypass -File .\scripts\Test-WinRE.ps1 # read-only, no elevation
powershell -ExecutionPolicy Bypass -File .\scripts\WinRE.ps1 -DryRun # walk the flow, log every decision, change nothing
powershell -ExecutionPolicy Bypass -File .\scripts\WinRE.ps1 # actually deploy
Confirm exit 0 and Operating mode: DEDICATED in the log. Only then expand the ring.
The one case in which the script cannot guarantee a working recovery route on a single run is the post-deletion residual corner: a failure after the old recovery partition has been deleted, on a machine whose C: is encrypted, with no successful retry. The trigger set is wider than the commonly cited New-Partition or Format-Volume pair: the whole-layout assertion (Assert-RecoveryPartitionLayout), drive-letter availability, drive-letter assignment, in-place decryption of the new partition, the post-delete extension fallback, the post-delete geometry verification, the planned-extent-available re-check, and the recovery-partition deletion itself when the previous route cannot be restored, all reach the same corner. The correct operational response is not to interrupt the script while it is inside the destructive window — every interruption between the delete and the enable is an opportunity for the machine to land in that corner.
Concretely:
WinRE.ps1 process if it appears to be slow. A full-update pass on an NVMe machine takes 3–5 minutes of I/O; the destructive window is a small fraction of that. If you kill the process inside the window, the recovery route is not restored by the script and the next run must recover from whatever state the interruption left.The scheduled task’s own ExecutionTimeLimit = PT6H and the script’s internal timeouts (300-second BitLocker polling, 5-second sleeps, 15-second network timeout) are the effective limits inside that. Nothing outside the script needs to enforce a shorter one.
v48 expands the destructive-surface scope. As of v48 patch 1, the intervening-anchor path extends the destructive window’s entry to include a resize of the intervening data partition (typically D:) in addition to or instead of the OS partition. The window’s scope is unchanged — the same partition-delete / create / format / deploy / register sequence runs — but the pre-shrink target may be the anchor rather than C:. The operational guidance above applies to the intervening-anchor case without change. The narrow scope means the participating partitions are predictable: one non-recovery data partition immediately preceding the type-coded recovery cluster, on the OS disk, validated before any mutation.
reagentc /disable → reagentc /enable window: do not schedule around itThe window is the interval during which WinRE is not registered and the recovery partition is in the process of being replaced. On a modern SSD the window is typically 15–45 seconds for the delete / create / format / deploy / reagentc /setreimage / reagentc /enable sequence. On a spinning disk or a very large WIM it can be longer.
Rule 3 is enforced by the script’s own ordering — every preparatory step that does not require a disabled WinRE runs before the disable. Your scheduling decisions should not undo that.
ExecutionTimeLimit is the correct upper bound; a second, shorter policy that runs in parallel can kill the run at the worst possible moment.The single-run recovery story is: Restore-PreviousWinRERoute on delete failure, Remove-OrphanPartition plus Restore-OSPartitionSize on create failure, and the enable-only path on a clean-but-disabled state. These handle every ordinary failure. They do not handle a process kill inside the window.
Rule 4 has four concrete operational consequences.
ScriptVersion forces one full-update pass per machine. Do it only when there is a reason — the migration sections below describe the five version boundaries that do this deliberately.LocalInputsId mismatch — none is improved by running the script again before the underlying condition changes. The correct response is to resolve the condition, then re-run. One nuance: the v45 pre-shrink deferral has two classes. The four RetrySuppressible = $true reasons (post-shrink free space below minimum, pre-shrink failed, pre-shrink verification failed, resize rounding reduced planned extent) write a deferral marker and will not re-attempt until the marker is cleared. The other reasons — including the C: volume could not be read for free-space check failure and the pre-deletion resolver guard — return RetrySuppressible = $false and will retry on the next scheduled run without operator intervention. Check the specific reason in the log.EXIT_FATAL with a “manual intervention required” message. Retrying automatically produces the same result. Resolve the underlying enable failure, delete the state file, then re-run.Before the first run on any machine, confirm:
EXIT_WARNING before any state mutation. The refusal is a stable, named deferral: a machine in the fleet whose architecture is not x64 will continue to exit 2 on every run until the manager adds support or the machine is replaced.C:\Program Files\7-Zip\7z.exe, or winget available so the script can install it.gist.github.com, api.github.com, downloads.dell.com, ftp.ext.hp.com, and the vendor-specific Lenovo endpoints (download.lenovo.com, support.lenovo.com — the latter is only reached by the map builder, not by production).EXIT_WARNING when no eligible volume has 3 GiB free. See the “Workspace selection” section below.There is no BitLocker precondition on the OS volume. The v43 patch 5 (further revision 5) policy is target-volume-based: the enable-only and dedicated-partition paths do not depend on C:’s BitLocker state at all. If the target recovery partition is encrypted, the script decrypts it in place via Set-RecoveryPartitionReadyForWinRE before calling reagentc. The only route that depends on C:’s BitLocker state is the OS-fallback route, because on that route the target volume is C:. The destructive partition path does not depend on C:’s state either — as of v44 patch 6 it does not consult C: at all. Neither the OS-fallback gate nor the destructive path modifies C:’s BitLocker state. See the “Device Encryption” section below.
The harness’s System diagnostic (scripts\Test-WinRE.ps1 Option 1) reports ImageState, C:’s BitLocker state, the target recovery partition’s BitLocker state, and the workspace-selection eligibility. Run it on a representative machine before pushing the task to a fleet.
Create the task with schtasks.exe, or as an XML payload via Group Policy / Intune.
schtasks.exeschtasks /Create /TN "WinRE Manager" ^
/TR "powershell.exe -NoProfile -ExecutionPolicy Bypass -File C:\ProgramData\OEM\WinRE.ps1" ^
/SC ONSTART /RU SYSTEM /RL HIGHEST /F
schtasks /Create /TN "WinRE Manager Weekly" ^
/TR "powershell.exe -NoProfile -ExecutionPolicy Bypass -File C:\ProgramData\OEM\WinRE.ps1" ^
/SC WEEKLY /D SUN /ST 03:00 /RU SYSTEM /RL HIGHEST /F
Save as WinRE-Manager.xml and register with schtasks /Create /XML WinRE-Manager.xml /TN "WinRE Manager":
<?xml version="1.0" encoding="UTF-16"?>
<Task version="1.4" xmlns="http://schemas.microsoft.com/windows/2004/02/mit/task">
<RegistrationInfo>
<Description>Self-healing Windows Recovery Environment manager.</Description>
<URI>\WinRE Manager</URI>
</RegistrationInfo>
<Triggers>
<BootTrigger>
<Enabled>true</Enabled>
<Delay>PT2M</Delay>
</BootTrigger>
<CalendarTrigger>
<StartBoundary>2026-01-04T03:00:00</StartBoundary>
<Enabled>true</Enabled>
<ScheduleByWeek>
<DaysOfWeek><Sunday /></DaysOfWeek>
<WeeksInterval>1</WeeksInterval>
</ScheduleByWeek>
</CalendarTrigger>
</Triggers>
<Principals>
<Principal id="Author">
<UserId>S-1-5-18</UserId>
<RunLevel>HighestAvailable</RunLevel>
</Principal>
</Principals>
<Settings>
<MultipleInstancesPolicy>IgnoreNew</MultipleInstancesPolicy>
<DisallowStartIfOnBatteries>false</DisallowStartIfOnBatteries>
<StopIfGoingOnBatteries>false</StopIfGoingOnBatteries>
<AllowHardTerminate>false</AllowHardTerminate>
<StartWhenAvailable>true</StartWhenAvailable>
<RunOnlyIfNetworkAvailable>false</RunOnlyIfNetworkAvailable>
<IdleSettings>
<StopOnIdleEnd>false</StopOnIdleEnd>
<RestartOnIdle>false</RestartOnIdle>
</IdleSettings>
<AllowStartOnDemand>true</AllowStartOnDemand>
<Enabled>true</Enabled>
<Hidden>false</Hidden>
<RunOnlyIfIdle>false</RunOnlyIfIdle>
<WakeToRun>false</WakeToRun>
<ExecutionTimeLimit>PT6H</ExecutionTimeLimit>
<Priority>4</Priority>
</Settings>
<Actions Context="Author">
<Exec>
<Command>powershell.exe</Command>
<Arguments>-NoProfile -ExecutionPolicy Bypass -File C:\ProgramData\OEM\WinRE.ps1</Arguments>
</Exec>
</Actions>
</Task>
Notes on the settings:
MultipleInstancesPolicy = IgnoreNew prevents two scheduled runs from overlapping. A long defrag /x on a nearly-full volume can take 30+ minutes. It has no effect on processes launched outside the scheduled task — see “One instance per machine” below for how the script defends itself against manual invocations.ExecutionTimeLimit = PT6H gives the run enough headroom; the script’s own internal timeouts (300-second BitLocker polling, 5-second sleeps, 15-second network timeout) are the effective limit inside that. Do not add a shorter limit at the orchestrator level — see Rule 3 above.RunOnlyIfNetworkAvailable = false is deliberate as of v44 patch 5. The script now has an offline fallback: when the manifest fetch fails, it trusts the state file’s stored DesiredStateId and takes the fast path if the local safety checks pass. With true, the task would not fire when the network is down, so the fallback would only help manual invocations — not the scheduled path where the machine actually needs it. The offline fallback is self-limiting (EXIT_WARNING when a full update is needed, never a silent partial deployment) and safe on a machine whose state file is valid. A first deployment on a machine with no state file will still exit EXIT_FATAL when offline, which is correct: the live manifest is required to compute the initial DesiredStateId. See the “Offline behavior” section below.StartWhenAvailable = true handles missed triggers after a cold boot.WinRE.ps1 assumes exclusive access to its selected image workspace and to the recovery partition it is operating on. Full updates select an eligible fixed NTFS volume and store the selection in the checkpoint for resume. As of v44 patch 4, the script enforces single-instance execution with a startup file lock.
Before any state-modifying action, and immediately after the startup banner, WinRE.ps1 opens C:\ProgramData\OEM\Logs\WinREManager.lock with an exclusive handle:
$Script:ProgramLockStream = [System.IO.File]::Open(
$lockPath,
[System.IO.FileMode]::OpenOrCreate,
[System.IO.FileAccess]::ReadWrite,
[System.IO.FileShare]::None
)
FileShare.None is enforced by the Windows kernel. Cross-session and cross-privilege exclusion are guaranteed by the OS, not by a security descriptor the script would have to configure. The handle is released automatically when the process exits, cleanly or after a crash — there is no stale-lock recovery logic to get wrong.
The lock is not acquired under -DryRun. A dry run is read-only and safe to run concurrently with a live deployment, so an operator can inspect a machine with WinRE.ps1 -DryRun or scripts\Test-WinRE.ps1 while a scheduled run is in progress.
The lock is released last in the finally block, after the WIM-mount discard and after every temporary drive letter has been cleaned up. This ensures the lock is held until all cleanup is complete, so a waiting second instance cannot begin work while the first is still releasing resources.
A second WinRE.ps1 process launched while another instance holds the lock fails fast. The log records:
[WARN] Another WinRE Manager instance is already running (program lock file is exclusively held). Exiting without making any changes. This is not a deployment failure - the other instance is doing the work and will complete on its own. …
The second instance exits with EXIT_WARNING (code 2) within seconds, before the Audit Mode guard, before the hardware check, and before any state-modifying action. It has consumed no state, written no state file, and left no checkpoint.
This is a change from the pre-patch-4 behavior. Before the lock, a concurrent instance would reach Step 2 and fail with FATAL ERROR: Cannot rename because item at '<workspace>\winre.wim' does not exist., which the orchestration layer saw as EXIT_FATAL (code 3). As of patch 4, the same scenario produces a clean EXIT_WARNING from the second instance.
C:\ProgramData\OEM\Logs\WinREManager.lock remains on disk between runs. It is not deleted on release, and it is empty by design — nothing is ever written to it. Its existence does not indicate a running instance. Only an active exclusive handle on the file blocks a second instance. Operators inspecting the log directory should not mistake the file’s presence for a stuck process.
What it protects against. The MultipleInstancesPolicy = IgnoreNew setting on the scheduled task prevents two scheduled runs from overlapping. Windows will refuse to start a second instance of the task while the first is still running.
What it does not protect against. A manual invocation — from an interactive PowerShell session, an RMM tool’s “run now” button, an Intune remediation script, or anything else that launches WinRE.ps1 outside the scheduled task — can overlap with a scheduled run. IgnoreNew is a property of the task, not of the script, so it has no effect on processes launched any other way. The v44 patch 4 file lock closes this gap for every launch path: the manual invocation and the scheduled run cannot both hold the lock.
WinRE.ps1 manually on a machine that has the scheduled task installedThe lock makes manual invocations safe, but a manual run that coincides with a scheduled run will still find itself on one side or the other of the lock:
EXIT_WARNING and the “Another WinRE Manager instance is already running” message. This is not a failure. Wait for the scheduled run to complete and try again.EXIT_WARNING on its own side.schtasks /Change /TN "WinRE Manager" /DISABLE (and the same for the Weekly task) before, /ENABLE afterwards.scripts\Test-WinRE.ps1 does not acquire the lock, does not touch the workspace, does not modify partitions, and is safe to run at any time, alongside any number of other processes.If the lock cannot be acquired for a reason other than contention — a permission error on C:\ProgramData\OEM\Logs\, a missing Logs directory that could not be created, or a transient filesystem issue — the script logs:
[WARN] Could not set up program lock at C:\ProgramData\OEM\Logs\WinREManager.lock : <error> - proceeding without single-instance protection; concurrent runs may collide
and continues without the lock. This is deliberate: the lock is defensive, and a broken lock must not prevent a legitimate deployment. The run is then vulnerable to the pre-patch-4 rename error for its entire duration. If you see this line, investigate why the lock could not be acquired and fix the underlying condition — otherwise the machine is one concurrent invocation away from the pre-patch-4 failure mode.
Place WinRE.ps1 in a directory that:
SYSTEM can read. C:\ProgramData\OEM\ is the standard choice.C:\ProgramData\OEM\ with default ACLs is safe; do not use C:\Temp\.The script writes to C:\ProgramData\OEM\Logs\ for the log, checkpoint file, and lock file. SYSTEM can write there by default.
The image workspace is where the script mounts, injects into, and exports the WinRE image during a full update. As of v45 patch 1, the workspace is chosen at run time from a set of eligible volumes rather than assumed to be C:\Temp\WinREWork.
The selector accepts only volumes that are:
ATA, SATA, NVMe, RAID, SAS, Spaces, Virtual, File Backed Virtual, SCM.Recovery or WINRE.<volume>:\Temp or <volume>:\Temp\WinREWork.USB, SD/MMC, network, FireWire, Fibre Channel, unknown-bus disks, and reparse-point workspace paths are excluded even when Windows reports them as fixed.
$MinFreeSpaceGB.WIM_READY checkpoint exists in the recorded workspace. The WIM_READY flag is set at Step 4 after a verified dism /Export-Image. base.wim is deleted at that point to reclaim workspace capacity before partition planning.A run that cannot find any eligible volume above the floor defers with EXIT_WARNING before touching the machine. Free space on any eligible volume and re-run.
X:\Temp\WinREWork directories on eligible internal volumes. It never touches user files or arbitrary temporary directories elsewhere.The harness Option 1 diagnostic reports the eligible workspace candidates and their free space.
The Audit Mode guard is the highest-priority startup check in the script, after the program lock. It runs before the hardware detection, before the manifest fetch, before the OEM pack resolution, before the DesiredStateId computation, before the pending-reboot block, and before the classifier. It reads a single registry value:
HKLM\SOFTWARE\Microsoft\Windows\CurrentVersion\Setup\State
→ ImageState (string)
If the value is absent, or if it is exactly IMAGE_STATE_COMPLETE, the guard passes and the script proceeds normally. Any other value causes a deferral.
During Audit Mode, OOBE, the sysprep generalize phase, and the sysprep specialize phase, Windows blocks reagentc /enable with ERROR_CANCELLED (0x4c7, 1223). This is not a bug in the script, and it is not related to the correctness of the deployed WIM or the state of the recovery partition. The OS refuses the call by design until the machine has reached a normal desktop.
The destructive partition work the script would otherwise perform on such a machine accomplishes nothing — the partition will be re-created on the next run anyway — and leaves a state file recording the deployment as complete. On the next run, the enable-only path fires because the state file matches and WinRE is still disabled, and it fails the same way. Without the guard, the machine loops forever, and the operator sees a fleet of machines that “deployed successfully” but never enabled WinRE.
Two additional defenses exist for the enable failure itself, in case the Audit Mode guard is somehow bypassed or the failure comes from a different source:
"failed" result or a "bitlocker" result — or when the target partition could not be prepared — it increments EnableFailureAttempts in the state file and exits EXIT_WARNING. It does not fall through to full update; the deployment is current and a rebuild would not change the outcome.EXIT_FATAL with an actionable message. The state file is left in place; the operator resolves the underlying cause and deletes C:\Recovery\OEM\winre_state.json to reset the counter.See exit-codes.md and troubleshooting.md for the full description.
The scheduled task does not need to be aware of the machine’s ImageState. The script handles it. But if you are deploying to a fleet of freshly imaged machines, be aware of the timing:
BootTrigger with a PT2M delay, a machine that reboots out of image but does not sign in for a day will exit EXIT_WARNING on every boot until a user reaches the desktop.EXIT_WARNING exits are cosmetic in this case; they will not loop, and they will not consume the enable-failure counter (the Audit Mode gate runs before the state file is even read).If you want to confirm the machine is ready to run before pushing the task, check ImageState on a representative image:
(Get-ItemProperty "HKLM:\SOFTWARE\Microsoft\Windows\CurrentVersion\Setup\State" -Name ImageState -ErrorAction SilentlyContinue).ImageState
A machine ready for deployment returns IMAGE_STATE_COMPLETE, or the command returns nothing because the SKU omits the key. Both are safe.
If the command returns any other value, the machine has not finished OOBE. Wait for a user to sign in, then let the scheduled task fire on the next weekly or boot trigger (or run the script manually).
Windows 11 24H2+ enables Device Encryption by default on hardware meeting TPM 2.0 and Secure Boot requirements. The relevant fact about Device Encryption for this project is that it can encrypt a newly created partition on the same disk, including a freshly created recovery partition, before the recovery type GUID and attributes can be applied to it. The partition then carries the default Basic Data type for a short window and can be claimed by the encryption service during it.
The v43 patch 5 (further revision 5) policy addresses this by acting on the volume reagentc will enable, not on the OS volume. Concretely:
reagentc /enable is called. Set-RecoveryPartitionReadyForWinRE prepares it: if the partition is already clean it returns immediately, otherwise it runs manage-bde -off against the partition and polls for completion at 5-second intervals up to a 300-second timeout. This is called on the enable-only path, the full-update path with an existing recovery partition, the full-update path with a newly created partition, and the pending-reboot repair path.VolumeStatus before deploying and defers unless it is FullyDecrypted. The script never modifies C:’s state. As of v44 patch 6, the destructive partition path does not consult C:’s state either — the dedicated-partition path’s success or failure depends on the partition geometry, not on C:’s encryption state.New-Partition time. This closes the window between partition creation and Set-RecoveryPartitionAttributes during which a plain Basic Data partition could be claimed by the Device Encryption service.Suspend-BitLockerForWinRE was deleted. Suspension does not prevent Device Encryption from claiming new partitions — proven in the field — and it is useless for OS-fallback, where reagentc refuses regardless.The v43 patch 5 (further revision) design gated on C:’s BitLocker state at startup. It deferred any run where C: was not FullyDecrypted with Protection On. That design was correct in the sense that it prevented the Dell Latitude 3550 / HP ProBook 450 G10 failure mode, but it was over-broad: it deferred on the FullyEncrypted + Protection Off state, which is the normal state of every fresh Win11 local-account machine before a Microsoft account sign-in. The startup gate was deferring on the majority of the fleet.
The change was driven by a same-machine test on 2026-09-29. With C: in the FullyEncrypted + Protection Off state, reagentc /enable against a dedicated recovery partition succeeded while reagentc /enable against the OS volume failed on the same machine in the same session. That established that reagentc’s check is on the target volume. The policy was inverted accordingly: prepare the target; do not gate on C: except on the OS-fallback route, where the target is C:.
The scheduled task does not need to be aware of the machine’s encryption state. The script handles it. But three deployment-time facts are worth knowing:
FullyDecrypted. The log names the C: state and tells the operator what to do: complete decryption of C: (manage-bde -off C:) or wait for an in-progress decryption to finish. Actions that complete encryption, add a protector, or enable protection do not resolve the state — they make C: more protected, not less. This deferral does not consume the enable-failure counter.New-Partition or Format-Volume failure — the run falls through to OS-fallback. The OS-fallback gate then consults C: and defers cleanly if C: is encrypted. As of v45 patch 1 the shrink is no longer inside the destructive window: a failed pre-shrink returns Deferred with the old route preserved and never reaches OS-fallback, so the shrink case no longer produces this fall-through.If you are deploying to a fresh fleet and want the first run to succeed rather than defer on the OS-fallback route, wait until manage-bde -status C: shows Fully Decrypted before pushing the task. On a modern SSD, decrypting C: from a mid-encryption state typically takes 30–90 minutes.
The VMD hardware presence check determines which driver set the script selects. As of v44 patch 6 the check is fail-closed: if the PnP enumeration reports an error during the query, the run treats VMD presence as indeterminate rather than as absent, and defers with EXIT_WARNING before committing any state.
The reasoning: guessing “absent” on a machine that genuinely has VMD hardware would select a driver set that omits the VMD package, and the deployed WinRE would not be able to see the OS disk. An empty device list from an errored enumeration is not evidence of absence. The cost of a deferral is one pipeline run; the cost of guessing wrong is a machine with a non-functional recovery environment.
A run that defers this way leaves the machine unchanged: no WIM deployed, no partition touched, no reagentc call made. The next scheduled run retries the enumeration; a transient PnP service issue is the most likely cause. If the deferral fires repeatedly, resolve the underlying PnP service issue — the deferral is a protective stop, not a degraded-success.
LocalInputsId added in v48 patch 1)The script’s network calls — the driver manifest fetch, the OEM map fetches, and the base WIM download — each default to a 15-second per-call timeout. The timeout applies to each individual request, not to the aggregate; the retry loops (2 attempts for the manifest, 3 for the OEM pack download) can multiply it. Before v44 patch 5, an offline machine would hang for over 10 minutes across the aggregate of the default 100-second timeouts before failing. That is now capped at roughly 90 seconds in the fully-offline case, dominated by the fixed sleeps in the retry loops rather than by the network timeouts.
More importantly, v44 patch 5 adds an offline fallback for the manifest fetch. The behavior depends on the state file’s contents.
The script reads the state file at C:\Recovery\OEM\winre_state.json and trusts its stored DesiredStateId directly. It does not recompute the ID from cached inputs — recomputing would require the OEM pack version, which is resolved from the OEM map, another gist on the same unavailable network.
The local safety checks remain fully enforced and do not depend on the manifest. Under v47, the fast-path conditions the offline path validates are the same as the online path:
Enabled.DesiredStateId matches the stored value (used directly; the DSI is not recomputed offline).Version + SPBuild) matches the state file’s DeployedWinREMetadata anchor, and no force-upgrade, driver-version change, or missing-anchor condition applies.LocalInputsId matches the current locally-computable inputs (v48 patch 1). A stored value of $null (a state file predating v48 patch 1) skips the check with a WARN.UsedOSFallback = $true in the state file and the active location on the OS partition.If all pass, the script takes the fast path, logs Offline fallback: using state file's stored DesiredStateId <id>, verifies the machine is healthy, and exits EXIT_WARNING (code 2) because $Script:offlineFallback = $true is set. The exit code is a degraded-success signal, not a failure. The machine is unchanged. Total runtime is under 90 seconds.
The v46 → v47 transition on offline machines. A v46-or-earlier state file lacks DeployedWinREMetadata. Its absence forces $needInject = $true regardless of network state, and the offline guard then fires: the machine needs a full update, but the live manifest is required for that, so the run exits EXIT_WARNING without further work. From the second v47 run onward, the state file has the anchor written by the first successful online run, and the offline fast path can fire normally. The one-time cost of the v47 migration on offline machines is therefore: the first scheduled v47 run must happen while the machine can reach the driver manifest.
On the next scheduled run with network, the script performs a full manifest fetch, detects any hardware drift that occurred while offline (a CPU swap, a BIOS update that flipped VMD, a motherboard replacement), and either takes the fast path (if no drift) or forces a rebuild (if drift).
The v48 patch 1 LocalInputsId field closes the offline hardware-drift residual. The state file now records a hash over the three deployment inputs observable on the local machine without a network fetch: hardware identity (Manufacturer|Model|MachineType), OS build, and CPU vendor/generation. VMD presence is deliberately excluded because its detection depends on the manifest’s requiredDevices patterns, which is precisely the input the offline fallback does not have. On the offline fallback path, the run recomputes the hash from the current machine and compares it against the stored value; a mismatch defers with EXIT_WARNING (Offline fallback: stored LocalInputsId <hash> does not match the current locally-computable inputs <hash>) rather than trusting a stored DesiredStateId that describes a machine whose hardware or OS has since changed. A VMD flip is not caught by LocalInputsId: VMD is a BIOS-configurable setting that can be toggled independently of the hardware, so HW and CPU do not capture it, and the hash deliberately excludes VMD because its detection needs the manifest. A machine whose VMD state changed while offline could take the fast path with a stale DesiredStateId — a documented residual of the offline fallback.
If the state file indicates a full update is needed (metadata drift, force-upgrade, driver version drift, or a locally-detected input change), the script cannot proceed offline — a full update requires the live manifest to resolve the driver set. It logs:
Offline fallback: the machine requires a full update (state file is stale or unhealthy), but the driver manifest is unavailable. Cannot proceed without a live manifest. Will retry on the next scheduled run when the network is available.
and exits EXIT_WARNING. No WIM is deployed, no partition is touched, WinRE is not disabled, no state file is written or modified. The next scheduled run with network completes the work.
A first deployment on a fresh machine requires the live manifest to compute the initial DesiredStateId. The script logs:
Driver manifest unavailable and no state file exists - a live manifest is required for the first deployment on this machine.
and throws. The outer catch logs FATAL ERROR: … and exits EXIT_FATAL (code 3). The machine retries on the next scheduled run when the network is available. No state-modifying action is attempted before the throw.
Because the offline fallback makes an offline scheduled run useful, the recommended task XML sets RunOnlyIfNetworkAvailable = false (see the XML above). With true, the task would not fire when the network is down, so the fallback would only help manual invocations — not the scheduled path where the machine actually needs it.
The scheduled task’s StartWhenAvailable = true continues to handle missed triggers after a cold boot. Combined with the offline fallback, a machine that reboots offline and misses its weekly window will still take a fast-path validation run on the next opportunity and exit EXIT_WARNING, then complete normally on the next run with network.
See exit-codes.md and troubleshooting.md for the full offline exit-code discussion.
The v44 patch 1 revision changes the DesiredStateId inputs (adding CPU vendor/generation and VMD presence) and therefore bumps ScriptVersion from 43 to 44. The effect on a managed fleet is a one-time full-update pass per machine on the next scheduled run.
Concretely, on the first run after the update, every machine will:
DesiredStateId.needInject = $true, and take the full-update path.On a healthy NVMe laptop this is roughly 3–5 minutes of I/O and CPU. Machines with a suitable existing recovery partition re-use it; no partition work occurs on healthy machines. This is the intended behaviour and the reason the Migration Note in CHANGELOG.md is present.
Two cases that are handled cleanly without operator attention:
PendingReboot = true. The DSI mismatch is detected before the pending-reboot block runs, so the full-update path executes rather than the pending-reboot repair path. No machine is lost.EnableFailureAttempts >= 3). The DSI mismatch makes the state file stale before the loop-breaker check runs, so those machines get one fresh attempt under the new ID. If the underlying cause is resolved, they recover; if not, they re-enter the loop-breaker on the fourth fresh attempt.Plan for the v44 patch 1 rollout the same way you would plan for a manifest-version bump: brief the operator community, expect the first run after the update to be slower than usual, and check the log on a canary machine to confirm the full-update path executed and the state file was rewritten under the new ID.
Change $ScriptVersion back to 43, revert the Get-DesiredStateId $parts array, and restore the previous Get-HardwareObject if the manufacturer normalisation differed. No data is lost; another fleet-wide rebuild occurs on the next run.
The v45 patch 1 revision is the shrink-first redesign of the destructive partition path. It bumps ScriptVersion from 44 to 45, which changes the SCRIPT component of the DesiredStateId and forces one full-update pass per managed machine on the next scheduled run — exactly the same shape as the v44 patch 1 migration, and for the same reason: the DSI is the deployment-identity fingerprint, and any change to the inputs that determine the deployed artifact is a version boundary.
The v45 pass is not slower on a healthy machine than any other full-update pass. The reorder moves the shrink into the reversible window; the total work is the same. Machines with a suitable existing recovery partition reuse it and perform no destructive work.
Concretely, on the first run after the update, every machine will:
DesiredStateId (with SCRIPT=45).needInject = $true, and take the full-update path.Two cases handled without operator attention:
C:\Recovery\OEM\winre_partition_deferred.json. The marker was introduced in v45 patch 1 and is keyed by DesiredStateId; it is honored only when the current run computes the same DSI. No operator action is needed on the v44→v45 transition, because no v44-era marker can exist.PendingReboot = true or EnableFailureAttempts >= 3. Handled the same way as in the v44 patch 1 migration: the DSI mismatch is detected before the pending-reboot and loop-breaker blocks.Change $ScriptVersion back to 44. The state file written under the v45 DSI becomes stale, the next run takes the full-update path under the older behavior, and the pipeline falls back to the pre-v45 destructive sequence. No data is lost; another fleet-wide rebuild occurs on the next run.
The v46 patch 1 revision clamps tailEnd to diskSize − 1 MiB before aligning in Get-PartitionPlan. It bumps ScriptVersion from 45 to 46, which changes the SCRIPT component of the DesiredStateId and forces one full-update pass per managed machine on the next scheduled run — the same shape as the v44 patch 1 and v45 patch 1 migrations.
The v46 pass is not slower on a healthy machine than any other full-update pass. The clamp changes the plan computation only when the last partition on the disk ends at diskSize (equivalently, at the disk-end reserve boundary); on every other layout, the plan is identical to v45’s. Machines with a suitable existing recovery partition reuse it and perform no destructive work.
The distinguishing behaviour of this migration is what happens to machines that v45 patch 1 left in OS-fallback because of the plan-clamp bug. Under v45, the plan was rejected after the old recovery partition had already been deleted, and the machine fell back to OS-fallback. The state file recorded the failing DesiredStateId and was accepted on every subsequent run, so the machine stayed in OS-fallback and did not retry. The v46 patch 1 ScriptVersion bump clears that state automatically: the state file is stale under the new DSI, the next run takes the full-update path, and the plan-clamp fix makes the plan valid on the same layout v45 rejected. The machine reaches DEDICATED on the first full-update pass.
Concretely, on the first run after the update, every machine will:
DesiredStateId (with SCRIPT=46).needInject = $true, and take the full-update path.Two cases handled without operator attention:
C:\Recovery\OEM\winre_partition_deferred.json. The marker is keyed by DesiredStateId; it is stale under the v46 DSI and is cleared on read by Read-PartitionDeferral. No operator action is needed.Change $ScriptVersion back to 45. The state file written under the v46 DSI becomes stale, the next run takes the full-update path under v45’s behavior, and the pipeline is once again vulnerable to the plan-clamp bug on machines whose last partition ends at diskSize. No data is lost; another fleet-wide rebuild occurs on the next run.
The v47 patch 1 revision introduces the third-party driver strip stage and rewrites several other rebuild-path elements. It bumps ScriptVersion from 46 to 47, which changes the SCRIPT component of the DesiredStateId and forces one full-update pass per managed machine on the next scheduled run — the same shape as the v44 patch 1, v45 patch 1, and v46 patch 1 migrations.
The v47 release is architecturally different from the earlier ScriptVersion bumps. v44 added DSI fields; v45 and v46 patch 1 each corrected a specific bug that had left machines stuck. v47 patch 1 changes what a rebuild produces, not just how a rebuild is triggered: the mounted image is stripped to zero third-party drivers, proven by re-enumeration, and only then is the current driver recipe injected. The deployed WIM bytes therefore differ from v46’s on the same inputs. The version bump is the mechanism by which the fleet converges on the strip-normalized driver set.
The v47 pass is not materially slower on a healthy machine than any other full-update pass. The strip stage on an already-clean image is a no-op — the ASUS field run logged Strip: image already has zero third-party drivers in under a second. On a machine whose registered image carries the OEM/VMD drivers from an earlier WinRE Manager cycle, the strip removes them and the current recipe is injected; the total time depends on the driver set size, not on the strip itself.
The distinguishing property of this migration is the metadata anchor. Every machine’s first v47 run writes a DeployedWinREMetadata field to the state file — the DISM servicing metadata (Version + SPBuild) of the WIM that was actually deployed. The field is the anchor the v47 drift detector compares against on every subsequent run. A machine whose state file lacks the anchor (all v46-or-earlier state files) forces a rebuild regardless of whether any other input changed; a machine whose state file has the anchor is evaluated against it.
Concretely, on the first run after the update, every machine will:
DesiredStateId (with SCRIPT=47).needInject = $true, and take the full-update path.DeployedWinREMetadata anchor.Three cases handled without operator attention:
PendingReboot = true or EnableFailureAttempts >= 3. As in every prior migration: the DSI mismatch is detected before the pending-reboot and loop-breaker blocks, so the full-update path executes and the counters are reset.C:\Recovery\OEM\winre_partition_deferred.json. The marker is keyed by DesiredStateId; it is stale under the v47 DSI and is cleared on read by Read-PartitionDeferral.Version or SPBuild. The metadata comparison would fire on the second v47 run, but is subsumed by the DSI mismatch on the first — the rebuild happens regardless.The offline-migration case is one step longer. On a machine whose first v47 run is offline, the state file lacks the anchor, $needInject is forced to $true, and the offline guard fires — the run exits EXIT_WARNING without a rebuild. The first successful online run writes the anchor; from the second v47 run onward, the offline fast path works normally. If a machine in your fleet runs offline for an extended period, expect its first online v47 run to be the one that writes the anchor.
Change $ScriptVersion back to 46 and revert the strip stage, the three-source selection logic, the race detector, and the metadata-drift changes. The state file written under the v47 DSI becomes stale, the next run takes the full-update path under v46’s behavior, and the strip-normalized image is replaced with a lineage-seeded one. The DeployedWinREMetadata field becomes inert on rollback — it is read only by the v47 drift detector. No data is lost; another fleet-wide rebuild occurs on the next run.
The v48 patch 1 revision is the largest single architectural change to the manager since v43. It bundles five distinct changes: intervening-partition handling, the architecture gate, the LocalInputsId state-file field, LKG-by-hash at any discovered recovery location, and transactional WIM replacement. It bumps ScriptVersion from 47 to 48, which changes the SCRIPT component of the DesiredStateId and forces one full-update pass per managed machine on the next scheduled run — the same shape as the v44 patch 1, v45 patch 1, v46 patch 1, and v47 patch 1 migrations.
The v48 pass is not materially slower on a healthy machine than any other full-update pass. On the ordinary non-intervening path, the pre-shrink target is C: and the pipeline is unchanged from v47; the transactional WIM replacement adds one hash of the existing target WIM (roughly 1-3 seconds on an NVMe) before the delete-then-copy sequence. On the intervening-anchor path, the pre-shrink target is the anchor instead of C:, and the anchor resize adds roughly one additional minute of I/O on a 720 GiB data partition.
Concretely, on the first run after the update, every machine will:
DesiredStateId (with SCRIPT=48).needInject = $true, and take the full-update path.EXIT_WARNING).LocalInputsId in the state file on the state write at the end of the run.C: | D: | Recovery), D: is the resize target and is validated before any mutation.LocalInputsId, and (as in v47) the DeployedWinREMetadata anchor.Three cases handled without operator attention:
PendingReboot = true or EnableFailureAttempts >= 3. The DSI mismatch is detected before the pending-reboot and loop-breaker blocks, so the full-update path executes and the counters are reset.C:\Recovery\OEM\winre_partition_deferred.json. The marker is keyed by DesiredStateId; it is stale under the v48 DSI and is cleared on read by Read-PartitionDeferral.Two cases where the v48 pipeline defers where v47 deferred differently:
intervening anchor is N MiB smaller than the plan requires ...). Under v47, the same layout would have produced the separation rejection because the intervening partition was in the way. Both reject; the v48 message names a different reason and offers the same deferral outcome. The operator path is unchanged: resolve the layout (extend the anchor, or reduce the desired recovery partition size), then delete the state file and deferral marker to force a retry. See troubleshooting.md.C: | D: | E: | Recovery layout still rejects, as it did under v47. The v48 message names the intervening partitions with Format-PartitionRef enrichment.Change $ScriptVersion back to 47 and revert the five bundled changes: the intervening-anchor plan logic, the architecture gate, the LocalInputsId computation and its state-file field, the LKG-by-hash discovery scan in Get-LKGWinREImagePath, and the transactional WIM replacement. The state file written under the v48 DSI becomes stale, the next run takes the full-update path under v47’s behavior, and the LocalInputsId field becomes inert. No data is lost; another fleet-wide rebuild occurs on the next run.
The v44 patches 2 through 8 and the v46 patch 2 additions do not bump ScriptVersion and do not change the DesiredStateId. They apply to every subsequent run without a state-file action. A machine that is already healthy continues to take the fast path. A machine that was mid-deployment when the patch rolled out continues from where it was — the checkpoint and state-file schemas are unchanged.
The v47 patch 2 revision is also not a fleet-wide rebuild: it does not bump ScriptVersion, and its only DSI-value change is a one-time event on a narrow machine class (machines with a padded Win32_ComputerSystemProduct.Version field). It is documented under ### v47 patch 2 specifically below. The v47 patch 1 revision is not in this list; it bumps ScriptVersion and is documented under “v47 patch 1 migration” above.
The v48 patch 2 revision is also not a fleet-wide rebuild: it does not bump ScriptVersion and does not change the DesiredStateId. It adds the fail-fast elevation guard, a source-WIM hash cache that closes a pre-existing enable-only fallthrough path, a resize-target resynchronisation fix on the extension-failure fallback path (the D1 fix), and several cosmetic log and comment corrections. On a machine that ran v48 patch 1, the patch 2 build takes the fast path unchanged; the only observable difference is the startup banner and, on an unelevated launch, the millisecond refusal instead of a delayed Mount-WindowsImage failure. Documented under ### v48 patch 2 specifically below.
dism /cleanup-image /StartComponentCleanup /ResetBase to the full-update pipeline; the size reduction materialises on the next natural rebuild, not immediately.Ensure-AdequateRecoveryPartition (removed in v44 patch 6) and a base.wim cleanup on the injection-failure abort branch (retained).C:\ProgramData\OEM\Logs\WinREManager.lock, so a second concurrent instance fails fast with EXIT_WARNING instead of colliding at Step 2.Find-SuitableRecoveryPartition. It also clears the VMD extraction directory before each extraction, and adds C:’s actual encryption state to the destructive-replacement WARN (diagnostic only; the decision to proceed is unchanged).Get-Partition -DiskNumber calls in Get-RecoveryPartitions against the empty-disk read that throws CmdletizationQuery_NotFound_DiskNumber on machines with an SD/MMC card reader. It also applies four post-review hardenings of Ensure-AdequateRecoveryPartition — the pre-deletion resolver guard, the extension-fallback bucket cap, the layout-assertion fail-closed branch, and the deletion-failure partial-rollback reporting. The build-logging values do not enter the DesiredStateId; the read guard silences a cosmetic error; the four hardenings change only the failure behaviour of the destructive sequence.Format-PartitionRef to the seven partition-identity plan-rejection reasons, so each reason now names the offending partition (disk number, partition number, size, label, type code); adds Remove-WindowsDriver and BusType to the harness parser self-test; makes -NonInteractive exit non-zero on any FAIL; guards the Lenovo-pack call on vendor; and adds the Compare-WimServicingMetadata harness mirror with an eight-case regression test. It does not change ScriptVersion or the DesiredStateId.ScriptVersion and does not change the DesiredStateId. The elevation guard refuses an unelevated launch in milliseconds; on an elevated launch the behavior is unchanged. The source-WIM hash cache closes a pre-existing enable-only fallthrough path where the state file would not be written because the source WIM became inaccessible after the temporary drive letter was removed.Because none of these patches changes ScriptVersion or the DesiredStateId, an already-completed machine will not rerun automatically; the deployment mechanism must invoke the script explicitly to pick the fixes up. The next natural rebuild (manifest bump, OEM pack version change, Windows build change, or CPU/VMD presence change) picks them up regardless.
dism /cleanup-image /StartComponentCleanup /ResetBase invocation from Step 3.base.wim cleanup on the abort branch is retained and does not need rollback.try block and the corresponding release in the finally block.$NetworkTimeoutSeconds to its absence on each call.Ensure-AdequateRecoveryPartition; revert Get-LenovoWinPEPack to a two-state return; revert the VMD presence check to treat enumeration errors as absent; remove the Step 2 stale-file cleanup; restore the earlier OS-fallback remediation wording.Find-SuitableRecoveryPartition and the main-flow $existingRecoveryParts back to accepting label-OR-type; the active-location classifier and the final-verification classifier back to promoting a label-only match to DEDICATED — remove the VMD extraction directory cleanup, and remove the C: encryption-state lookup from the destructive-replacement WARN.WinRE status: ..., Version: ..., Source WIM build: ..., Post-deploy WIM build: ...) from the startup, force-upgrade, and post-deploy paths; remove the Version extraction from Get-WinREState; remove the -ErrorAction SilentlyContinue guard from the two Get-Partition -DiskNumber calls in Get-RecoveryPartitions; revert the four Ensure-AdequateRecoveryPartition hardenings to their pre-v46-patch-2 form.Win32_ComputerSystemProduct.Version trim; revert the workspace candidate filter to IsSystem alone; revert Get-GitHubBaseWinRE to throw on extraction failure; revert the extension-failure fallback geometry and sizing to the pre-patch-2 form; remove the unrecognized-route guard.Remove-ItemIfExist to log unconditionally; revert the enable-only -not $targetPart branch to $needInject = $false; revert the seven plan-rejection reasons to the unenriched form; revert the race-detector comment to the “microseconds” phrasing.try block; remove the $cachedSourceHash / $cachedSourceBuild / $cachedSourceMetadata capture and revert the state-write section to re-reading $SourceWim; revert the two comment corrections. None of these affect the state-file schema or the DesiredStateId.None of these rollbacks affects the state-file schema or the DesiredStateId. (The v47 patch 1 rollback does affect the schema — it introduces the DeployedWinREMetadata field — and is documented separately under “v47 patch 1 migration” above.)
The script’s exit codes carry meaning. See exit-codes.md for the full matrix.
For MDM / orchestration:
| Exit code | Recommended action |
|---|---|
| 0 | None. Record success. |
| 1 | Reboot the machine at the next convenient window. The script will finish on the next boot. |
| 2 | Investigate. WinRE is functional but degraded, or the script deferred work, or the enable step failed and the counter incremented. Collect the log and check the state file’s LastUpdated timestamp, its LastEnableResult field, whether a deferral marker exists at C:\Recovery\OEM\winre_partition_deferred.json, and the state file’s LocalInputsId. Sixteen distinct cases are documented in exit-codes.md: OS-fallback, v47 strip-failure abort, v47 patch 2 base-WIM copy-integrity check, Step 3 → Step 4 pipeline gate, Audit Mode deferral, VMD-query-indeterminate deferral, OS-fallback BitLocker deferral, v45 pre-shrink deferral, enable-only failure, concurrent-instance deferral (v44 patch 4), offline-fallback deferral (v44 patch 5), v47 race-detector abort at the two /disable sites, v48 architecture-gate refusal, v48 intervening-anchor surplus rejection, v48 multi-intervening rejection, and v48 offline LocalInputsId mismatch. The concurrent-instance case is not a failure — the other instance is doing the work. The offline-fallback fast-path case is a degraded-success — the machine is healthy and unchanged. The VMD-query-indeterminate, v45 pre-shrink, v47 race-detector, v48 architecture-gate, v48 surplus, v48 multi-intervening, and v48 LocalInputsId cases are protective deferrals — no state was committed. |
| 3 | Investigate. The run failed and did not write state, or the enable-failure loop-breaker fired. Collect the log. Do not retry automatically. One exception: the v48 patch 2 elevation guard exits 3 with FATAL: WinRE Manager requires an elevated (Administrator) PowerShell session. — this is expected behavior on an unelevated launch, not a failure. Relaunch from an elevated prompt or via scripts\WinRE-Manager.cmd. |
Do not treat exit code 2 as success. A machine in OS-fallback is intentionally reported as a warning; it will be treated as a healthy machine by the fast path only if the state file records UsedOSFallback = true for the current DesiredStateId. A machine on which the Audit Mode guard or a v45 pre-shrink deferral fired will have an unchanged (or absent) state file and a single deferral line in the log. See exit-codes.md for how to distinguish the sixteen cases.
Do not configure retry loops that ignore the exit code and re-run unconditionally. Deferrals are not improved by retrying — the same gate will fire on the next run (Rule 4). Enable failures are handled by the counter; after three consecutive failures the loop-breaker fires and requires manual intervention. The concurrent-instance deferral is not a failure at all; retrying while the other instance is still running will simply produce another EXIT_WARNING. The VMD-query-indeterminate deferral is not improved by retrying — the underlying PnP service issue must be resolved. The v45 pre-shrink deferral is not improved by retrying until the constraint is resolved (free space on C:, disk layout, oversized recovery partition, or shrink failure). The one exception is the transient volume-read failure: it returns RetrySuppressible = $false and does not write a deferral marker, so the next scheduled run retries it automatically without operator intervention. The v47 race-detector abort is not improved by retrying immediately — the underlying Windows Update servicing of the registered WinRE is a transient condition, and the next scheduled run (or a manual re-run after the WU settles) is the correct response. The v47 strip-failure abort is not improved by retrying until the specific strip failure has been diagnosed — see troubleshooting.md for the specific failure modes. The v48 architecture-gate refusal is a stable deferral: any machine whose architecture is not x64 will continue to exit 2 on every run until the manager adds support or the machine is replaced. The v48 intervening-anchor surplus and multi-intervening rejections are stable layout deferrals; a retry produces the same rejection until the layout is changed. The v48 offline LocalInputsId mismatch resolves on the next run with network, which recomputes the DSI. Resolve the underlying condition, then re-run.
Deploy via a Win32 app or a PowerShell script platform script. The Intune “Scripts and remediations” feature is well-suited:
Test-Path C:\ProgramData\OEM\Logs\WinRE-Manager.log and reagentc /info reports Enabled.WinRE.ps1 invoked as SYSTEM.For a curated deployment, package the script as an Intune Win32 app with a detection rule on the log file and reagentc state. Run once, then allow the built-in weekly scheduled task to keep the state fresh.
When using Intune to push a remediation, be aware that the remediation script may run outside the scheduled task’s IgnoreNew gate. As of v44 patch 4, the program lock handles this: if the weekly task happens to be running at the same moment, the remediation exits with EXIT_WARNING (code 2) and the message Another WinRE Manager instance is already running (program lock file is exclusively held). This is not a failure. The remediation will succeed on the next cycle, or you can wait for the scheduled task to complete and re-run the remediation manually.
Do not disable the lock to “fix” the remediation. Before patch 4, the same collision caused an EXIT_FATAL (code 3) with a rename error. Some Intune remediation pipelines treated that as a hard failure and retried aggressively, which could produce a cascade of colliding runs. The lock converts the collision into a clean, self-limiting EXIT_WARNING. If your Intune pipeline retries on any non-zero exit, adjust it to treat code 2 as a soft failure (route to a queue) and code 3 as a hard failure (investigate). See exit-codes.md for the recommended orchestration policy.
By default, WinRE Manager fetches the driver manifest, the three OEM maps, and the base WIM repository from the project maintainer’s GitHub account. All five artifacts can be replaced with self-hosted equivalents so your fleet does not depend on those external endpoints.
See self-hosting.md for the trust model, the exact URLs and variables to change, the map-builder scripts, and the air-gapped deployment procedure.
The log at C:\ProgramData\OEM\Logs\WinRE-Manager.log is append-only and grows unbounded. Rotate it via your existing log-collection pipeline:
WinRE.ps1 does not write to the Event Log; add a wrapper script that reads the last N lines of the log and writes them as an event after each run.C:\ProgramData\OEM\Logs\WinRE-Manager.log.At ~100 log lines per healthy run, an unmanaged machine generates ~5 MB per year. Rotation is not urgent but is worth having.
The script is destructive on the recovery partition and non-destructive on the OS partition, but it does modify the OS partition geometry on the shrink path. Roll out in rings:
Test-WinRE.ps1 first, then WinRE.ps1 -DryRun, then WinRE.ps1. Verify exit code 0. A transient exit code 2 from a concurrent scheduled task that happened to be running is not a failure — see “One instance per machine” above.Add a step 0 before the canary ring: run the current script on one machine and confirm the full-update pass completes and the state file is rewritten under the new DesiredStateId. Then roll out normally. The Migration Note in CHANGELOG.md and the “v44 patch 1 migration” section above describe the expected behaviour.
No state-file action is needed on healthy machines — the surviving part of the fix (the base.wim cleanup) is exercised on the next full-update pass, and the guard portion is removed by v44 patch 6 anyway. See “Later patches (no ScriptVersion bump)” below.
Watch for the new EXIT_WARNING from concurrent invocations in your orchestration logs. A cluster of these on a machine usually means an RMM tool or a manual operation overlapped with the scheduled task; it is a signal about the operator side, not a script defect.
The first observable change is faster failure on machines with network issues. Machines that previously exited EXIT_FATAL after ~100-second timeouts now exit EXIT_WARNING after ~15-second timeouts, and healthy machines take the fast path offline instead of failing. Confirm on a canary that an offline run exits 2 with the Offline fallback: using state file's stored DesiredStateId … line rather than 3.
Watch for:
VMD hardware detection was indeterminate message.manage-bde -off C: instead of the earlier (wrong) “complete encryption” guidance.Watch for machines that exit 0 after one full-update pass where they previously cycled through rebuilds. Before patch 7, a machine with a Basic Data partition labelled “Recovery” on the OS disk never converged: the fast-path count included the label-only partition, so the “exactly one recovery partition on the OS disk” condition was never satisfied, while the destructive path correctly refused to delete it. After patch 7 the classifier, count, and destructive path all agree on the type-code rule, so the machine converges — either to a proper DEDICATED state (if the OS-shrink succeeds) or stably to OS-fallback (if it does not).
No state-file action and no observable orchestration change. The reorder is an internal cost reduction for healthy paths. Confirm on a canary that a fast-path run still exits 0 with the same log shape as before.
The v45 revision bumps ScriptVersion and forces one full-update pass per managed machine. The pass itself is not slower than a normal full-update pass. Watch for:
DEDICATED (via the reuse path, or via a successful destructive replacement). Returned to the fast path on the second run.Dedicated replacement deferred before partition deletion (<Reason>) in the log. The v45 pre-shrink deferral fired; the old recovery partition is preserved and the machine is unchanged. Investigate the <Reason>. Common causes: insufficient free space on C:, blocking layout, oversized recovery-typed partition, transient volume read, or shrink failure. Resolve the constraint and delete C:\Recovery\OEM\winre_partition_deferred.json to force a retry.Pre-deletion inventory: followed in the same run by OS-fallback deferred: C: could not be confirmed fully decrypted. This is the post-deletion residual corner the [v45 patch 1] CHANGELOG entry documents. The v45 reorder narrowed it: only New-Partition or Format-Volume failures after deletion, on an encrypted C:, now reach the corner (the shrink case is closed). Capture the full log and file a bug — the project wants field data on this corner.Extension-failure fallback: creating recovery partition, the run continued with a slightly different geometry than the plan and exited 2. This is expected. If the log shows Extension-failure fallback unavailable, the machine fell through to the OS-fallback decision. Check the machine’s end state with the harness.Exit-code observations. Watch for these exit-code patterns as well:
UsedOSFallback = true in the state file and a state file LastUpdated timestamp newer than the run’s start time — the post-delete partition creation could not complete and the run deliberately ended in OS-fallback. (As of v45 patch 1, a shrink failure no longer reaches this path — the pre-shrink runs in the reversible window and defers with the old route preserved.) The log will show the Dedicated recovery partition creation failed after all attempts. banner.Setup\State\ImageState=… in the log and the state file absent or unchanged — the Audit Mode / OOBE guard fired. The machine has not finished OOBE. This is expected on freshly imaged machines that are still pre-first-sign-in; wait for the machine to reach a normal desktop and re-run on the next scheduled trigger.VMD hardware detection was indeterminate; deferring because the driver set cannot be safely determined. in the log and the state file unchanged (or the state file absent) — the VMD hardware presence check could not complete because of a PnP enumeration error. No WIM was deployed, no partition was touched, no reagentc call was made. Resolve the PnP service issue and re-run. Do not retry immediately; the same check will fire.OS-fallback deferred: C: could not be confirmed fully decrypted (Test-VolumeEncrypted=…) in the log and the state file’s LastUpdated timestamp unchanged (or the state file absent) — the dedicated replacement could not complete and the run fell through to the OS-fallback route, whose target (C:) is not confirmed fully decrypted. No WIM was deployed and no reagentc call was made. Wait for C: to reach FullyDecrypted, or complete decryption of C: with manage-bde -off C:, then re-run on the next scheduled trigger. Do not retry immediately.Dedicated replacement deferred before partition deletion (<Reason>) in the log and the state file unchanged — the v45 pre-shrink deferral fired. The old recovery partition is still present and WinRE is still registered to it. Free space on C:, correct the disk layout, address the oversized partition, retry a transient volume read, or investigate the shrink or route-restore failure as appropriate, then delete C:\Recovery\OEM\winre_partition_deferred.json (and, if you also want to clear the deployment identity, C:\Recovery\OEM\winre_state.json) to force a retry.Image injection did not complete. Stopping before Step 4 and before any deployment. in the log — OEM or VMD injection failed and the pipeline gate stopped the run before deployment. The WIM was not deployed, the partition was not touched, and WinRE was not disabled. The checkpoint was set back to 2. The next run retries from step 2. Investigate the injection failure (OEM pack download, extraction, INF validation, or VMD driver download).Enable-only /enable failed (attempt N of 3) or Enable-only /enable refused with the BitLocker error after the target partition was confirmed unencrypted (attempt N of 3) in the log and a non-"ok" LastEnableResult in the state file — the enable step failed on a machine whose deployment is current. The counter incremented. Investigate the enable failure itself (ReAgent.xml corruption, missing registration, a Windows component problem, an antivirus product holding a file). The next run retries enable-only. If the counter reaches 3, the loop-breaker fires on the following run and the exit code becomes 3.Another WinRE Manager instance is already running (program lock file is exclusively held) in the log — a concurrent instance holds the lock. This is not a failure. It appears when a manual invocation overlaps with a scheduled run, or when an Intune remediation fires while the scheduled task is running. No action required; the other instance is completing the work.Offline fallback: the machine requires a full update (state file is stale or unhealthy), but the driver manifest is unavailable in the log — the machine is offline and the state file indicates a full update is needed. No operator action; the next scheduled run with network completes the work. If the log instead shows Offline fallback: using state file's stored DesiredStateId …, the fast path fired and the exit is a degraded-success, not a deferral — the machine is healthy and unchanged.Pre-deletion inventory: followed in the same run by OS-fallback deferred: C: could not be confirmed fully decrypted — the post-deletion residual corner that the [v45 patch 1] CHANGELOG entry documents. The existing type-coded recovery partition was deleted and a post-deletion New-Partition or Format-Volume failure occurred, and the OS-fallback gate deferred because C: is encrypted. The machine ends with neither a dedicated recovery partition nor OS-fallback. Capture the full log and file a bug — the project wants field data on this corner. See troubleshooting.md for the operator-facing procedure.Cannot rename because item at '<workspace>\winre.wim' does not exist — this is now only reachable if the lock could not be acquired for a non-contention reason. Check the log for Could not set up program lock at … earlier in the run. Investigate permissions on C:\ProgramData\OEM\Logs\, whether the directory exists, and whether a filesystem issue is affecting the log path.dism /Export-Image failed in the log — 7-Zip or DISM problem.cannot deploy a new WinRE image while WinRE is still Enabled — reagentc /disable returned nonzero. Investigate before retrying.FATAL: WinRE is not enabled at exit — the machine lost its recovery partition. This is the pre-patch-5 Device Encryption failure mode. Follow the recovery procedure in troubleshooting.md.Refusing to retry reagentc /enable — the enable-failure loop-breaker fired. Do not retry automatically; the same guard will fire because the state file still records the counter. Resolve the underlying enable failure (see troubleshooting.md), delete C:\Recovery\OEM\winre_state.json, then re-run.Driver manifest unavailable and no state file exists — the machine is offline and has no state file. A first deployment requires a live manifest. Retry when the network is available.The v46 patch 1 revision clamps tailEnd to diskSize − 1 MiB before aligning in Get-PartitionPlan, and bumps ScriptVersion from 45 to 46. It forces one full-update pass per managed machine on the next scheduled run — the same shape as the v44 patch 1 and v45 patch 1 migrations. The pass itself is not slower than a normal full-update pass.
The v46 patch 1 canary is primarily about confirming that machines the v45 plan-clamp bug left in OS-fallback recover. Watch for:
Operating mode: DEDICATED in the log is the confirmation.alignment reserve value and the post-delete check for overlapCount=0. See troubleshooting.md for the plan-rejection corner and its fall-through conditions.Dedicated replacement deferred before partition deletion (<Reason>). A pre-shrink deferral fired during the v46 migration pass. The old recovery partition is preserved and the machine is unchanged. Investigate the <Reason>, resolve the constraint, and delete C:\Recovery\OEM\winre_partition_deferred.json to force a retry.The v46 patch 2 revision does not bump ScriptVersion and does not change the DesiredStateId. It adds three build-drift log lines, a small read guard in Get-RecoveryPartitions, and four post-review hardenings of Ensure-AdequateRecoveryPartition. The canary is primarily observational:
WinRE status: ..., Location: ..., Version: <version>; the force-upgrade detection line reads Source WIM build: <build> (path: <path>); the post-deploy line reads Post-deploy WIM build: <build> (source: <path>). On a fast-path run, the first two lines appear and the third does not — that is expected, because the fast path does not deploy a WIM.CmdletizationQuery_NotFound_DiskNumber error is silent on machines with an SD/MMC card reader or an empty USB enclosure. Before v46 patch 2, such a machine produced three copies of the error in the harness diagnostic and one per Get-RecoveryPartitions invocation in the production log. After v46 patch 2, the same machine produces no such error. If it still appears, the log’s version banner will tell you whether the machine is running the patched build.Dedicated replacement deferred before partition deletion (active WinRE location could not be resolved). As of v48 patch 1, the pre-deletion resolver guard fires ahead of the pre-shrink and ahead of reagentc /disable; the machine is left with WinRE still Enabled and the old recovery partition intact. This is a code-review hardening that has not been exercised in the field; if you see it, please capture the full log and file a bug per the Reporting a bug section. See troubleshooting.md for the re-registration procedure.The v47 patch 1 revision introduces the third-party driver strip stage and bumps ScriptVersion from 46 to 47. It forces one full-update pass per managed machine on the next scheduled run — the same shape as the earlier ScriptVersion-bump migrations. The pass itself is not materially slower than a normal full-update pass; on an already-clean image, the strip stage is a no-op.
Watch for:
Strip: pre-strip third-party inventory: N package(s) and, if N was zero, Strip: image already has zero third-party drivers), a source-selection line naming the branch taken, ResetBase, the export, and Deployed WinRE metadata: <Version>|<SPBuild> immediately before the state write. Operating mode: DEDICATED confirms the end state.Strip: removed <OEM#.inf> lines. The strip stage actually ran against a non-empty third-party driver set — this is the case the field data does not yet cover. If any of these lines appear on the canary, the canary has exercised the strip-and-reinject loop against a non-empty set, which is exactly the case the project wants more field data on. Capture the full log and report it.Strip: initial Get-WindowsDriver enumeration failed or Strip: Remove-WindowsDriver failed for <OEM#.inf> — the strip stage rejected the candidate. The checkpoint was rolled back to Step 2, base.wim and any stale winre_optimized.wim were removed, and the run exited EXIT_WARNING. The old recovery route is intact. Capture the full log and file a bug — the strip-failure path has not been exercised in the field either.Registered-source recheck at Step 5 deployment : CHANGED or Registered-source recheck at Ensure-AdequateRecoveryPartition : CHANGED — the v47 race detector fired. The abort happened before the /disable call; the machine’s registered WinRE drifted during candidate preparation. On a machine whose WU is currently servicing WinRE, this is expected; the abort is protective. On a machine with no plausible WU activity, investigate. Both abort sites return EXIT_WARNING with the machine unchanged.State has no DeployedWinREMetadata anchor (previous deployment could not record it) - rebuilding — this is the expected message for the first v47 run on a machine whose state file was written by v46 or earlier. The machine will rebuild on this run and the anchor will be written by the end. Any subsequent run should log Registered WinRE metadata: <ver>|<sp>; last deployed: <ver>|<sp> and take the fast path.Registered WinRE metadata could not be read - forcing rebuild so the repair path can replace the image — the registered image’s DISM metadata is unreadable. The rebuild proceeds and the repair path replaces the image. If this fires repeatedly across runs, the registered image is persistently unreadable and the machine is not converging — investigate the registered recovery partition’s storage state.Pre-injection third-party driver count: N with N > 0 after the strip stage — the filterless and filtered DISM queries disagreed. The strip stage remains authoritative (the filterless form is the documented third-party inventory), but the discrepancy was flagged with a non-fatal warning. Capture the log; this is the observability cross-check working as designed, but the project wants field data on when it fires.Checkpoint step N has no recorded source identity (legacy or GitHub-sourced checkpoint); invalidating to rebuild from a known source, Checkpoint step N source identity <hash> cannot be re-validated against the current selection (no selected source, or GitHub cold-start has no on-disk hash); invalidating to rebuild, or Checkpoint step N source identity <hash> does not match the currently-selected source <hash> (<reason>); invalidating to rebuild — the v47 source-hash binding invalidated a stale checkpoint on resume. The run restarts from Step 0 and rebuilds from the current source. This is a protective invalidation, not a failure. If it fires on a machine whose checkpoint was from a same-source run, the source actually changed during the interval, which the invalidator correctly caught.Also check the ASUS PRIME H510M-D v47 field-verification transcript in CHANGELOG.md for the expected shape of a healthy migration run — the log lines the canary should produce are the same ones.
The v47 patch 2 revision does not bump ScriptVersion and does not change the DesiredStateId on any machine whose Win32_ComputerSystemProduct.Version field is not padded. It ships six hardening fixes on top of v47 patch 1: the OS-fallback missing-WIM guard without a status restriction; the base-WIM copy-integrity check; the machine-type fallback trim; the workspace candidate filter IsSystem -and -not IsBoot; the GitHub base-WIM extraction contract change (return $false instead of throw); and the -SuppressSuccessLog switch on the duplicate Verified partition log line. The canary is primarily observational:
v47 patch 2. The startup line reads ========== WinRE Manager Started (v47 patch 2) ==========.Copied base WIM from <path>; copy hash verified after the base-WIM copy. The mismatch branch (Base WIM copy hash mismatch: source=..., copied=... - rejecting candidate) is not expected on a healthy canary.Version field (observed: some ASUS), expect a one-time DesiredStateId change from the trim fix and one full-update pass. All other machines continue on the fast path.The v48 patch 1 revision bumps ScriptVersion from 47 to 48 and forces one full-update pass per managed machine. The pass is not materially slower than any other full-update pass on the ordinary non-intervening path; the intervening-anchor path adds roughly one additional minute of anchor resize on the anchored machine class.
Watch for:
Machines that exit 0 after the first v48 full-update pass. The normal case. The log should show the architecture gate line (Architecture gate passed: x64), the LocalInputsId: <hash> line, and — on a machine whose only candidate recovery partition is undersized and whose predecessor on the OS disk is a shrinkable data partition — the Intervening-anchor plan: anchor is disk N part M; anchor shrink ... line followed by the anchor shrink, delete, New-Partition, layout assertion, and reagentc /enable sequence. Operating mode: DEDICATED confirms the end state.
Machines whose log shows Intervening-anchor plan: ... and complete DEDICATED. This is the v48 feature working on the exact layout that motivated it (the AMD Ryzen 7 5825U machine with C: | D: | Recovery). The 2026-10-06 field run confirmed this path end-to-end; a canary that produces the same shape is expected.
Machines whose log shows Unsupported OS architecture '<arch>' and exit 2. The v48 architecture gate refused. This is expected on ARM64, x86, or machines whose architecture could not be resolved. No state was committed; the machine is unchanged. This is a stable deferral, not a failure: any machine in the fleet whose architecture is not x64 will continue to exit 2 on every run until the manager adds support or the machine is replaced. ARM64 support would be a separate, tested change to the WinRE image pipeline.
Machines whose log shows intervening anchor is N MiB smaller than the plan requires ... and exit 2. The v48 surplus-case rejection. The old recovery partition is intact; no partition was touched. Resolve the layout (extend the anchor to close the gap, or reduce the desired recovery partition size) and delete the state file and deferral marker to force a retry. See troubleshooting.md.
Machines whose log shows recovery-typed partitions exist beyond the intervening anchor's recovery cluster and exit 2. The v48 multi-intervening rejection. This is unchanged behaviour from v47 for the multi-intervening case; the message is now enriched with Format-PartitionRef entries naming the intervening partitions.
Machines whose log shows Preserving existing target ... as rollback copy at <path>. The v48 transactional WIM replacement took the rollback-copy branch, meaning the deployment target was an existing partition with an existing WIM. The subsequent log lines confirm the copy, the hash verification, and the release of the rollback copy after /setreimage and /enable. The branch is expected on the reuse path and on the OS-fallback path; it is skipped on the create-fresh-partition path, where there is no pre-existing WIM to preserve.
Machines whose log shows Deploy-WimTransactional ... copy failed - restoring rollback copy and exits code 3 (EXIT_FATAL). The transactional replacement caught a copy failure. The rollback copy was restored and hash-verified by Deploy-WimTransactional before it returned failure, so the previous WIM is in place on the active route — but both call sites (Deploy-WimToPartition and the OS-fallback deploy branch) treat any $false return as fatal and exit before /setreimage or /enable is attempted. The machine is left with WinRE disabled and the previous WIM in place; the next scheduled run re-attempts deployment. This branch is code-review-only in v48 and has not been exercised in the field. If a canary hits it, capture the full log per the Reporting a bug section.
$needInject = $true forced by the DSI mismatch and will exit 2 with Offline fallback: the machine requires a full update, since a rebuild needs the live manifest. The LocalInputsId comparison does not fire on that path; it fires when a full-update-needing machine has a v48 state file and its hardware has drifted. The next online run writes the v48 state file and the anchor.LocalInputsId mismatch (v48 patch 1). A machine whose state file was written by v48 (with a LocalInputsId), whose next run is offline, and whose locally-observable deployment inputs have changed (hardware identity, OS build, CPU vendor/generation) will exit 2 with Offline fallback: stored LocalInputsId <hash> does not match the current locally-computable inputs <hash>. This is a protective deferral: the stored DesiredStateId no longer describes this machine. The machine is unchanged; bring it online so a fresh manifest can be fetched and a new DSI computed. Note that this check fires before the offline fast path evaluates the drift detector, so a machine with both a LocalInputsId mismatch and a metadata anchor mismatch reports the LocalInputsId line first.The v48 patch 2 revision does not bump ScriptVersion and does not change the DesiredStateId. It adds the fail-fast elevation guard, a source-WIM hash cache that closes a pre-existing enable-only fallthrough path, a resize-target resynchronisation fix on the extension-failure fallback path (the D1 fix), and several cosmetic log and comment corrections. The canary is primarily observational:
v48 patch 2. The startup line reads ========== WinRE Manager Started (v48 patch 2) ==========..\scripts\WinRE.ps1 -DryRun should produce a single FATAL log line (FATAL: WinRE Manager requires an elevated (Administrator) PowerShell session.) and exit code 3 in milliseconds, before the log directory is created. From an elevated prompt, the same command should proceed as in v48 patch 1./setreimage fails. The change is code-review-only in v48 and has not been exercised in the field.The script does not have an uninstall path. To disable it:
schtasks /Delete /TN "WinRE Manager" /F
schtasks /Delete /TN "WinRE Manager Weekly" /F
The state file, deferral marker, log, lock file, and any deployed recovery partition remain. The machine is in a healthy end state and Windows Update will continue to service the recovery image normally. The lock file at C:\ProgramData\OEM\Logs\WinREManager.lock is inert once the scheduled task is removed; it can be deleted manually if desired, and it will be recreated if the script is ever run again. The deferral marker at C:\Recovery\OEM\winre_partition_deferred.json is likewise inert; it can be deleted manually.
If you need to revert a machine to its pre-WinRE-Manager state, restore the partition layout from a backup. The script does not create one.
To revert a DesiredStateId change specifically, see the “Rollback” subsections under “v44 patch 1 migration”, “v45 patch 1 migration”, “v46 patch 1 migration”, “v47 patch 1 migration”, and “v48 patch 1 migration” above. To revert code from a patch that does not change the DesiredStateId, see the “Rollback for later non-bumping patches” subsection above.