Deployment

How to run WinRE Manager on a single machine or across a managed fleet.

How to deploy

There are two ways to deploy WinRE Manager. Pick the one that matches your situation.

Single machine, one-off repair

For a single broken machine, run the production script once from an elevated PowerShell prompt. It inspects the machine, repairs the recovery environment, writes a state file, and exits. There is no installer, no service, and no configuration file to maintain.

As of v48 patch 2, the script refuses to run unelevated in milliseconds. The refusal is a single FATAL log line (FATAL: WinRE Manager requires an elevated (Administrator) PowerShell session.) and exit code 3, before the program lock, before hardware probes, and before any network fetch. The guard creates the log directory itself so its FATAL message has somewhere to be written. Before v48 patch 2, an unelevated launch proceeded through hardware detection, manifest fetch, VMD detection, and the multi-minute GitHub base-WIM download before failing at Mount-WindowsImage with The requested operation requires elevation. — a failure mode that wasted roughly four minutes per run in the 2026-10-05 field log. Both the interactive Administrator session and the SYSTEM context used by the scheduled task pass the guard.

# Elevated PowerShell (Run as Administrator)
powershell -ExecutionPolicy Bypass -File .\scripts\WinRE.ps1 -DryRun   # walk the flow, log every decision, change nothing
powershell -ExecutionPolicy Bypass -File .\scripts\WinRE.ps1           # actually deploy

The -ExecutionPolicy Bypass -File form applies the bypass only to the child process; it does not change the machine’s execution policy.

If you want to inspect the machine first, run the read-only harness (Test-WinRE.ps1). It requires no elevation and modifies nothing. See the README for the harness menu options.

For a single-machine deployment, the rest of this document is reference material — the exit-code table, the Audit Mode precondition, and the Device Encryption policy all apply to a one-off run exactly as they apply to a fleet run. You can skip the scheduled-task sections.

Managed fleet, continuous

For a managed fleet, deploy the script as a scheduled task running as NT AUTHORITY\SYSTEM, triggered at boot and weekly. This is the recommended model. Every machine keeps its recovery environment healthy on its own; a healthy machine’s fast path is dominated by Windows’ own CIM and PnP enumeration (typically 10–20 seconds), and a full repair runs only when something has changed.

The sections below this point cover the fleet deployment in detail: how to create the task, where to put the script, how to handle the Audit Mode precondition, how to interpret the exit codes, and how to roll out in rings.

Run WinRE.ps1 as NT AUTHORITY\SYSTEM from a scheduled task, triggered at boot and weekly. The script is fully idempotent; running it on a healthy machine costs a few seconds of CIM and PnP enumeration and a read-only scan.

Do not run it from a user logon script. It modifies partition tables and the BitLocker state of the target recovery partition; those operations require SYSTEM.

Design invariants at deployment time

The script’s design is organized around four rules, in this order — see docs/architecture.md for the full hierarchy.

  1. Never break Windows RE.
  2. Never leave a machine without a working recovery route — to the extent the machine, its OS, and its storage stack allow.
  3. Minimize the reagentc /disable → reagentc /enable window.
  4. Do no work unless needed. When work is needed, prepare everything before touching anything.

Rules 1–3 are invariants: no deployment scenario should weaken them. Rule 4 is the working rule: the discipline that makes the fast path and the enable-only path free of side effects, and that makes the destructive path fail safe.

Each rule has a direct operational translation for a fleet operator.

Rule 1 — never break Windows RE: canary, do not bulk-push

The script’s destructive sequence is protected by a read-only plan, a fail-closed layout assertion, and a reversible shrink window. It does not need to be bullet-proof to be deployed; it needs to be canaried so that any regression is caught on one machine, not a thousand.

Before pushing to a fleet, run on one machine end-to-end:

powershell -ExecutionPolicy Bypass -File .\scripts\Test-WinRE.ps1            # read-only, no elevation
powershell -ExecutionPolicy Bypass -File .\scripts\WinRE.ps1 -DryRun          # walk the flow, log every decision, change nothing
powershell -ExecutionPolicy Bypass -File .\scripts\WinRE.ps1                  # actually deploy

Confirm exit 0 and Operating mode: DEDICATED in the log. Only then expand the ring.

Rule 2 — never leave a machine without a working recovery route: do not interfere with the recovery window

The one case in which the script cannot guarantee a working recovery route on a single run is the post-deletion residual corner: a failure after the old recovery partition has been deleted, on a machine whose C: is encrypted, with no successful retry. The trigger set is wider than the commonly cited New-Partition or Format-Volume pair: the whole-layout assertion (Assert-RecoveryPartitionLayout), drive-letter availability, drive-letter assignment, in-place decryption of the new partition, the post-delete extension fallback, the post-delete geometry verification, the planned-extent-available re-check, and the recovery-partition deletion itself when the previous route cannot be restored, all reach the same corner. The correct operational response is not to interrupt the script while it is inside the destructive window — every interruption between the delete and the enable is an opportunity for the machine to land in that corner.

Concretely:

The scheduled task’s own ExecutionTimeLimit = PT6H and the script’s internal timeouts (300-second BitLocker polling, 5-second sleeps, 15-second network timeout) are the effective limits inside that. Nothing outside the script needs to enforce a shorter one.

v48 expands the destructive-surface scope. As of v48 patch 1, the intervening-anchor path extends the destructive window’s entry to include a resize of the intervening data partition (typically D:) in addition to or instead of the OS partition. The window’s scope is unchanged — the same partition-delete / create / format / deploy / register sequence runs — but the pre-shrink target may be the anchor rather than C:. The operational guidance above applies to the intervening-anchor case without change. The narrow scope means the participating partitions are predictable: one non-recovery data partition immediately preceding the type-coded recovery cluster, on the OS disk, validated before any mutation.

Rule 3 — minimize the reagentc /disable → reagentc /enable window: do not schedule around it

The window is the interval during which WinRE is not registered and the recovery partition is in the process of being replaced. On a modern SSD the window is typically 15–45 seconds for the delete / create / format / deploy / reagentc /setreimage / reagentc /enable sequence. On a spinning disk or a very large WIM it can be longer.

Rule 3 is enforced by the script’s own ordering — every preparatory step that does not require a disabled WinRE runs before the disable. Your scheduling decisions should not undo that.

The single-run recovery story is: Restore-PreviousWinRERoute on delete failure, Remove-OrphanPartition plus Restore-OSPartitionSize on create failure, and the enable-only path on a clean-but-disabled state. These handle every ordinary failure. They do not handle a process kill inside the window.

Rule 4 — do no work unless needed: trust the fast path; do not force reruns

Rule 4 has four concrete operational consequences.

Preconditions

Before the first run on any machine, confirm:

There is no BitLocker precondition on the OS volume. The v43 patch 5 (further revision 5) policy is target-volume-based: the enable-only and dedicated-partition paths do not depend on C:’s BitLocker state at all. If the target recovery partition is encrypted, the script decrypts it in place via Set-RecoveryPartitionReadyForWinRE before calling reagentc. The only route that depends on C:’s BitLocker state is the OS-fallback route, because on that route the target volume is C:. The destructive partition path does not depend on C:’s state either — as of v44 patch 6 it does not consult C: at all. Neither the OS-fallback gate nor the destructive path modifies C:’s BitLocker state. See the “Device Encryption” section below.

The harness’s System diagnostic (scripts\Test-WinRE.ps1 Option 1) reports ImageState, C:’s BitLocker state, the target recovery partition’s BitLocker state, and the workspace-selection eligibility. Run it on a representative machine before pushing the task to a fleet.

Scheduled task

Create the task with schtasks.exe, or as an XML payload via Group Policy / Intune.

Via schtasks.exe

schtasks /Create /TN "WinRE Manager" ^
  /TR "powershell.exe -NoProfile -ExecutionPolicy Bypass -File C:\ProgramData\OEM\WinRE.ps1" ^
  /SC ONSTART /RU SYSTEM /RL HIGHEST /F

schtasks /Create /TN "WinRE Manager Weekly" ^
  /TR "powershell.exe -NoProfile -ExecutionPolicy Bypass -File C:\ProgramData\OEM\WinRE.ps1" ^
  /SC WEEKLY /D SUN /ST 03:00 /RU SYSTEM /RL HIGHEST /F

Via XML (preferred for MDM)

Save as WinRE-Manager.xml and register with schtasks /Create /XML WinRE-Manager.xml /TN "WinRE Manager":

<?xml version="1.0" encoding="UTF-16"?>
<Task version="1.4" xmlns="http://schemas.microsoft.com/windows/2004/02/mit/task">
  <RegistrationInfo>
    <Description>Self-healing Windows Recovery Environment manager.</Description>
    <URI>\WinRE Manager</URI>
  </RegistrationInfo>
  <Triggers>
    <BootTrigger>
      <Enabled>true</Enabled>
      <Delay>PT2M</Delay>
    </BootTrigger>
    <CalendarTrigger>
      <StartBoundary>2026-01-04T03:00:00</StartBoundary>
      <Enabled>true</Enabled>
      <ScheduleByWeek>
        <DaysOfWeek><Sunday /></DaysOfWeek>
        <WeeksInterval>1</WeeksInterval>
      </ScheduleByWeek>
    </CalendarTrigger>
  </Triggers>
  <Principals>
    <Principal id="Author">
      <UserId>S-1-5-18</UserId>
      <RunLevel>HighestAvailable</RunLevel>
    </Principal>
  </Principals>
  <Settings>
    <MultipleInstancesPolicy>IgnoreNew</MultipleInstancesPolicy>
    <DisallowStartIfOnBatteries>false</DisallowStartIfOnBatteries>
    <StopIfGoingOnBatteries>false</StopIfGoingOnBatteries>
    <AllowHardTerminate>false</AllowHardTerminate>
    <StartWhenAvailable>true</StartWhenAvailable>
    <RunOnlyIfNetworkAvailable>false</RunOnlyIfNetworkAvailable>
    <IdleSettings>
      <StopOnIdleEnd>false</StopOnIdleEnd>
      <RestartOnIdle>false</RestartOnIdle>
    </IdleSettings>
    <AllowStartOnDemand>true</AllowStartOnDemand>
    <Enabled>true</Enabled>
    <Hidden>false</Hidden>
    <RunOnlyIfIdle>false</RunOnlyIfIdle>
    <WakeToRun>false</WakeToRun>
    <ExecutionTimeLimit>PT6H</ExecutionTimeLimit>
    <Priority>4</Priority>
  </Settings>
  <Actions Context="Author">
    <Exec>
      <Command>powershell.exe</Command>
      <Arguments>-NoProfile -ExecutionPolicy Bypass -File C:\ProgramData\OEM\WinRE.ps1</Arguments>
    </Exec>
  </Actions>
</Task>

Notes on the settings:

One instance per machine

WinRE.ps1 assumes exclusive access to its selected image workspace and to the recovery partition it is operating on. Full updates select an eligible fixed NTFS volume and store the selection in the checkpoint for resume. As of v44 patch 4, the script enforces single-instance execution with a startup file lock.

How the lock works

Before any state-modifying action, and immediately after the startup banner, WinRE.ps1 opens C:\ProgramData\OEM\Logs\WinREManager.lock with an exclusive handle:

$Script:ProgramLockStream = [System.IO.File]::Open(
    $lockPath,
    [System.IO.FileMode]::OpenOrCreate,
    [System.IO.FileAccess]::ReadWrite,
    [System.IO.FileShare]::None
)

FileShare.None is enforced by the Windows kernel. Cross-session and cross-privilege exclusion are guaranteed by the OS, not by a security descriptor the script would have to configure. The handle is released automatically when the process exits, cleanly or after a crash — there is no stale-lock recovery logic to get wrong.

The lock is not acquired under -DryRun. A dry run is read-only and safe to run concurrently with a live deployment, so an operator can inspect a machine with WinRE.ps1 -DryRun or scripts\Test-WinRE.ps1 while a scheduled run is in progress.

The lock is released last in the finally block, after the WIM-mount discard and after every temporary drive letter has been cleaned up. This ensures the lock is held until all cleanup is complete, so a waiting second instance cannot begin work while the first is still releasing resources.

What a second instance sees

A second WinRE.ps1 process launched while another instance holds the lock fails fast. The log records:

[WARN] Another WinRE Manager instance is already running (program lock file is exclusively held). Exiting without making any changes. This is not a deployment failure - the other instance is doing the work and will complete on its own. …

The second instance exits with EXIT_WARNING (code 2) within seconds, before the Audit Mode guard, before the hardware check, and before any state-modifying action. It has consumed no state, written no state file, and left no checkpoint.

This is a change from the pre-patch-4 behavior. Before the lock, a concurrent instance would reach Step 2 and fail with FATAL ERROR: Cannot rename because item at '<workspace>\winre.wim' does not exist., which the orchestration layer saw as EXIT_FATAL (code 3). As of patch 4, the same scenario produces a clean EXIT_WARNING from the second instance.

The lock file persists on disk

C:\ProgramData\OEM\Logs\WinREManager.lock remains on disk between runs. It is not deleted on release, and it is empty by design — nothing is ever written to it. Its existence does not indicate a running instance. Only an active exclusive handle on the file blocks a second instance. Operators inspecting the log directory should not mistake the file’s presence for a stuck process.

What the scheduled task protects against, and what it does not

What it protects against. The MultipleInstancesPolicy = IgnoreNew setting on the scheduled task prevents two scheduled runs from overlapping. Windows will refuse to start a second instance of the task while the first is still running.

What it does not protect against. A manual invocation — from an interactive PowerShell session, an RMM tool’s “run now” button, an Intune remediation script, or anything else that launches WinRE.ps1 outside the scheduled task — can overlap with a scheduled run. IgnoreNew is a property of the task, not of the script, so it has no effect on processes launched any other way. The v44 patch 4 file lock closes this gap for every launch path: the manual invocation and the scheduled run cannot both hold the lock.

If you need to run WinRE.ps1 manually on a machine that has the scheduled task installed

The lock makes manual invocations safe, but a manual run that coincides with a scheduled run will still find itself on one side or the other of the lock:

Non-contention lock failures

If the lock cannot be acquired for a reason other than contention — a permission error on C:\ProgramData\OEM\Logs\, a missing Logs directory that could not be created, or a transient filesystem issue — the script logs:

[WARN] Could not set up program lock at C:\ProgramData\OEM\Logs\WinREManager.lock : <error> - proceeding without single-instance protection; concurrent runs may collide

and continues without the lock. This is deliberate: the lock is defensive, and a broken lock must not prevent a legitimate deployment. The run is then vulnerable to the pre-patch-4 rename error for its entire duration. If you see this line, investigate why the lock could not be acquired and fix the underlying condition — otherwise the machine is one concurrent invocation away from the pre-patch-4 failure mode.

Where to put the script

Place WinRE.ps1 in a directory that:

  1. SYSTEM can read. C:\ProgramData\OEM\ is the standard choice.
  2. Cannot be modified by non-admin users. A scheduled task running as SYSTEM must not execute a script that a standard user can overwrite. C:\ProgramData\OEM\ with default ACLs is safe; do not use C:\Temp\.

The script writes to C:\ProgramData\OEM\Logs\ for the log, checkpoint file, and lock file. SYSTEM can write there by default.

Workspace selection (v45 patch 1)

The image workspace is where the script mounts, injects into, and exports the WinRE image during a full update. As of v45 patch 1, the workspace is chosen at run time from a set of eligible volumes rather than assumed to be C:\Temp\WinREWork.

Eligibility rules

The selector accepts only volumes that are:

USB, SD/MMC, network, FireWire, Fibre Channel, unknown-bus disks, and reparse-point workspace paths are excluded even when Windows reports them as fixed.

Selection order

  1. The recorded workspace from the checkpoint, if it is still eligible and above the free-space floor.
  2. The eligible non-OS volume with the most free space (fresh-run preference).
  3. The OS volume itself, only if no eligible non-OS volume is available.

Free-space floors

A run that cannot find any eligible volume above the floor defers with EXIT_WARNING before touching the machine. Free space on any eligible volume and re-run.

What this means for you

The harness Option 1 diagnostic reports the eligible workspace candidates and their free space.

Audit Mode and OOBE

The Audit Mode guard is the highest-priority startup check in the script, after the program lock. It runs before the hardware detection, before the manifest fetch, before the OEM pack resolution, before the DesiredStateId computation, before the pending-reboot block, and before the classifier. It reads a single registry value:

HKLM\SOFTWARE\Microsoft\Windows\CurrentVersion\Setup\State
  → ImageState (string)

If the value is absent, or if it is exactly IMAGE_STATE_COMPLETE, the guard passes and the script proceeds normally. Any other value causes a deferral.

Why the guard exists

During Audit Mode, OOBE, the sysprep generalize phase, and the sysprep specialize phase, Windows blocks reagentc /enable with ERROR_CANCELLED (0x4c7, 1223). This is not a bug in the script, and it is not related to the correctness of the deployed WIM or the state of the recovery partition. The OS refuses the call by design until the machine has reached a normal desktop.

The destructive partition work the script would otherwise perform on such a machine accomplishes nothing — the partition will be re-created on the next run anyway — and leaves a state file recording the deployment as complete. On the next run, the enable-only path fires because the state file matches and WinRE is still disabled, and it fails the same way. Without the guard, the machine loops forever, and the operator sees a fleet of machines that “deployed successfully” but never enabled WinRE.

Two additional defenses exist for the enable failure itself, in case the Audit Mode guard is somehow bypassed or the failure comes from a different source:

See exit-codes.md and troubleshooting.md for the full description.

What this means for you

The scheduled task does not need to be aware of the machine’s ImageState. The script handles it. But if you are deploying to a fleet of freshly imaged machines, be aware of the timing:

If you want to confirm the machine is ready to run before pushing the task, check ImageState on a representative image:

(Get-ItemProperty "HKLM:\SOFTWARE\Microsoft\Windows\CurrentVersion\Setup\State" -Name ImageState -ErrorAction SilentlyContinue).ImageState

A machine ready for deployment returns IMAGE_STATE_COMPLETE, or the command returns nothing because the SKU omits the key. Both are safe.

If the command returns any other value, the machine has not finished OOBE. Wait for a user to sign in, then let the scheduled task fire on the next weekly or boot trigger (or run the script manually).

Device Encryption (v43 patch 5, further revision 5)

Windows 11 24H2+ enables Device Encryption by default on hardware meeting TPM 2.0 and Secure Boot requirements. The relevant fact about Device Encryption for this project is that it can encrypt a newly created partition on the same disk, including a freshly created recovery partition, before the recovery type GUID and attributes can be applied to it. The partition then carries the default Basic Data type for a short window and can be claimed by the encryption service during it.

The v43 patch 5 (further revision 5) policy addresses this by acting on the volume reagentc will enable, not on the OS volume. Concretely:

Why the earlier further-revision policy was replaced

The v43 patch 5 (further revision) design gated on C:’s BitLocker state at startup. It deferred any run where C: was not FullyDecrypted with Protection On. That design was correct in the sense that it prevented the Dell Latitude 3550 / HP ProBook 450 G10 failure mode, but it was over-broad: it deferred on the FullyEncrypted + Protection Off state, which is the normal state of every fresh Win11 local-account machine before a Microsoft account sign-in. The startup gate was deferring on the majority of the fleet.

The change was driven by a same-machine test on 2026-09-29. With C: in the FullyEncrypted + Protection Off state, reagentc /enable against a dedicated recovery partition succeeded while reagentc /enable against the OS volume failed on the same machine in the same session. That established that reagentc’s check is on the target volume. The policy was inverted accordingly: prepare the target; do not gate on C: except on the OS-fallback route, where the target is C:.

What this means for you

The scheduled task does not need to be aware of the machine’s encryption state. The script handles it. But three deployment-time facts are worth knowing:

If you are deploying to a fresh fleet and want the first run to succeed rather than defer on the OS-fallback route, wait until manage-bde -status C: shows Fully Decrypted before pushing the task. On a modern SSD, decrypting C: from a mid-encryption state typically takes 30–90 minutes.

VMD fail-closed guard (v44 patch 6)

The VMD hardware presence check determines which driver set the script selects. As of v44 patch 6 the check is fail-closed: if the PnP enumeration reports an error during the query, the run treats VMD presence as indeterminate rather than as absent, and defers with EXIT_WARNING before committing any state.

The reasoning: guessing “absent” on a machine that genuinely has VMD hardware would select a driver set that omits the VMD package, and the deployed WinRE would not be able to see the OS disk. An empty device list from an errored enumeration is not evidence of absence. The cost of a deferral is one pipeline run; the cost of guessing wrong is a machine with a non-functional recovery environment.

A run that defers this way leaves the machine unchanged: no WIM deployed, no partition touched, no reagentc call made. The next scheduled run retries the enumeration; a transient PnP service issue is the most likely cause. If the deferral fires repeatedly, resolve the underlying PnP service issue — the deferral is a protective stop, not a degraded-success.

Offline behavior (v44 patch 5; fast-path gate updated in v47 patch 1; LocalInputsId added in v48 patch 1)

The script’s network calls — the driver manifest fetch, the OEM map fetches, and the base WIM download — each default to a 15-second per-call timeout. The timeout applies to each individual request, not to the aggregate; the retry loops (2 attempts for the manifest, 3 for the OEM pack download) can multiply it. Before v44 patch 5, an offline machine would hang for over 10 minutes across the aggregate of the default 100-second timeouts before failing. That is now capped at roughly 90 seconds in the fully-offline case, dominated by the fixed sleeps in the retry loops rather than by the network timeouts.

More importantly, v44 patch 5 adds an offline fallback for the manifest fetch. The behavior depends on the state file’s contents.

Case 1 — state file present, its fast-path checks pass

The script reads the state file at C:\Recovery\OEM\winre_state.json and trusts its stored DesiredStateId directly. It does not recompute the ID from cached inputs — recomputing would require the OEM pack version, which is resolved from the OEM map, another gist on the same unavailable network.

The local safety checks remain fully enforced and do not depend on the manifest. Under v47, the fast-path conditions the offline path validates are the same as the online path:

If all pass, the script takes the fast path, logs Offline fallback: using state file's stored DesiredStateId <id>, verifies the machine is healthy, and exits EXIT_WARNING (code 2) because $Script:offlineFallback = $true is set. The exit code is a degraded-success signal, not a failure. The machine is unchanged. Total runtime is under 90 seconds.

The v46 → v47 transition on offline machines. A v46-or-earlier state file lacks DeployedWinREMetadata. Its absence forces $needInject = $true regardless of network state, and the offline guard then fires: the machine needs a full update, but the live manifest is required for that, so the run exits EXIT_WARNING without further work. From the second v47 run onward, the state file has the anchor written by the first successful online run, and the offline fast path can fire normally. The one-time cost of the v47 migration on offline machines is therefore: the first scheduled v47 run must happen while the machine can reach the driver manifest.

On the next scheduled run with network, the script performs a full manifest fetch, detects any hardware drift that occurred while offline (a CPU swap, a BIOS update that flipped VMD, a motherboard replacement), and either takes the fast path (if no drift) or forces a rebuild (if drift).

The v48 patch 1 LocalInputsId field closes the offline hardware-drift residual. The state file now records a hash over the three deployment inputs observable on the local machine without a network fetch: hardware identity (Manufacturer|Model|MachineType), OS build, and CPU vendor/generation. VMD presence is deliberately excluded because its detection depends on the manifest’s requiredDevices patterns, which is precisely the input the offline fallback does not have. On the offline fallback path, the run recomputes the hash from the current machine and compares it against the stored value; a mismatch defers with EXIT_WARNING (Offline fallback: stored LocalInputsId <hash> does not match the current locally-computable inputs <hash>) rather than trusting a stored DesiredStateId that describes a machine whose hardware or OS has since changed. A VMD flip is not caught by LocalInputsId: VMD is a BIOS-configurable setting that can be toggled independently of the hardware, so HW and CPU do not capture it, and the hash deliberately excludes VMD because its detection needs the manifest. A machine whose VMD state changed while offline could take the fast path with a stale DesiredStateId — a documented residual of the offline fallback.

Case 2 — state file present, its fast-path checks fail

If the state file indicates a full update is needed (metadata drift, force-upgrade, driver version drift, or a locally-detected input change), the script cannot proceed offline — a full update requires the live manifest to resolve the driver set. It logs:

Offline fallback: the machine requires a full update (state file is stale or unhealthy), but the driver manifest is unavailable. Cannot proceed without a live manifest. Will retry on the next scheduled run when the network is available.

and exits EXIT_WARNING. No WIM is deployed, no partition is touched, WinRE is not disabled, no state file is written or modified. The next scheduled run with network completes the work.

Case 3 — no state file at all

A first deployment on a fresh machine requires the live manifest to compute the initial DesiredStateId. The script logs:

Driver manifest unavailable and no state file exists - a live manifest is required for the first deployment on this machine.

and throws. The outer catch logs FATAL ERROR: … and exits EXIT_FATAL (code 3). The machine retries on the next scheduled run when the network is available. No state-modifying action is attempted before the throw.

What this means for the scheduled task

Because the offline fallback makes an offline scheduled run useful, the recommended task XML sets RunOnlyIfNetworkAvailable = false (see the XML above). With true, the task would not fire when the network is down, so the fallback would only help manual invocations — not the scheduled path where the machine actually needs it.

The scheduled task’s StartWhenAvailable = true continues to handle missed triggers after a cold boot. Combined with the offline fallback, a machine that reboots offline and misses its weekly window will still take a fast-path validation run on the next opportunity and exit EXIT_WARNING, then complete normally on the next run with network.

See exit-codes.md and troubleshooting.md for the full offline exit-code discussion.

v44 patch 1 migration

The v44 patch 1 revision changes the DesiredStateId inputs (adding CPU vendor/generation and VMD presence) and therefore bumps ScriptVersion from 43 to 44. The effect on a managed fleet is a one-time full-update pass per machine on the next scheduled run.

Concretely, on the first run after the update, every machine will:

  1. Compute the new DesiredStateId.
  2. Compare it to the state file written by v43 patch 5 (further revision 5).
  3. Find a mismatch, set needInject = $true, and take the full-update path.
  4. Rebuild the WIM, deploy it, and write a state file under the new ID.
  5. Return to the fast path on the second run.

On a healthy NVMe laptop this is roughly 3–5 minutes of I/O and CPU. Machines with a suitable existing recovery partition re-use it; no partition work occurs on healthy machines. This is the intended behaviour and the reason the Migration Note in CHANGELOG.md is present.

Two cases that are handled cleanly without operator attention:

Plan for the v44 patch 1 rollout the same way you would plan for a manifest-version bump: brief the operator community, expect the first run after the update to be slower than usual, and check the log on a canary machine to confirm the full-update path executed and the state file was rewritten under the new ID.

Rollback

Change $ScriptVersion back to 43, revert the Get-DesiredStateId $parts array, and restore the previous Get-HardwareObject if the manufacturer normalisation differed. No data is lost; another fleet-wide rebuild occurs on the next run.

v45 patch 1 migration

The v45 patch 1 revision is the shrink-first redesign of the destructive partition path. It bumps ScriptVersion from 44 to 45, which changes the SCRIPT component of the DesiredStateId and forces one full-update pass per managed machine on the next scheduled run — exactly the same shape as the v44 patch 1 migration, and for the same reason: the DSI is the deployment-identity fingerprint, and any change to the inputs that determine the deployed artifact is a version boundary.

The v45 pass is not slower on a healthy machine than any other full-update pass. The reorder moves the shrink into the reversible window; the total work is the same. Machines with a suitable existing recovery partition reuse it and perform no destructive work.

Concretely, on the first run after the update, every machine will:

  1. Compute the new DesiredStateId (with SCRIPT=45).
  2. Compare it to the state file written by v44 (any patch).
  3. Find a mismatch, set needInject = $true, and take the full-update path.
  4. Rebuild the WIM, deploy it, and write a state file under the new ID.
  5. Return to the fast path on the second run.

Two cases handled without operator attention:

Rollback

Change $ScriptVersion back to 44. The state file written under the v45 DSI becomes stale, the next run takes the full-update path under the older behavior, and the pipeline falls back to the pre-v45 destructive sequence. No data is lost; another fleet-wide rebuild occurs on the next run.

v46 patch 1 migration

The v46 patch 1 revision clamps tailEnd to diskSize − 1 MiB before aligning in Get-PartitionPlan. It bumps ScriptVersion from 45 to 46, which changes the SCRIPT component of the DesiredStateId and forces one full-update pass per managed machine on the next scheduled run — the same shape as the v44 patch 1 and v45 patch 1 migrations.

The v46 pass is not slower on a healthy machine than any other full-update pass. The clamp changes the plan computation only when the last partition on the disk ends at diskSize (equivalently, at the disk-end reserve boundary); on every other layout, the plan is identical to v45’s. Machines with a suitable existing recovery partition reuse it and perform no destructive work.

The distinguishing behaviour of this migration is what happens to machines that v45 patch 1 left in OS-fallback because of the plan-clamp bug. Under v45, the plan was rejected after the old recovery partition had already been deleted, and the machine fell back to OS-fallback. The state file recorded the failing DesiredStateId and was accepted on every subsequent run, so the machine stayed in OS-fallback and did not retry. The v46 patch 1 ScriptVersion bump clears that state automatically: the state file is stale under the new DSI, the next run takes the full-update path, and the plan-clamp fix makes the plan valid on the same layout v45 rejected. The machine reaches DEDICATED on the first full-update pass.

Concretely, on the first run after the update, every machine will:

  1. Compute the new DesiredStateId (with SCRIPT=46).
  2. Compare it to the state file written by v45 patch 1.
  3. Find a mismatch, set needInject = $true, and take the full-update path.
  4. Rebuild the WIM, deploy it, and write a state file under the new ID.
  5. Return to the fast path on the second run.

Two cases handled without operator attention:

Rollback

Change $ScriptVersion back to 45. The state file written under the v46 DSI becomes stale, the next run takes the full-update path under v45’s behavior, and the pipeline is once again vulnerable to the plan-clamp bug on machines whose last partition ends at diskSize. No data is lost; another fleet-wide rebuild occurs on the next run.

v47 patch 1 migration

The v47 patch 1 revision introduces the third-party driver strip stage and rewrites several other rebuild-path elements. It bumps ScriptVersion from 46 to 47, which changes the SCRIPT component of the DesiredStateId and forces one full-update pass per managed machine on the next scheduled run — the same shape as the v44 patch 1, v45 patch 1, and v46 patch 1 migrations.

The v47 release is architecturally different from the earlier ScriptVersion bumps. v44 added DSI fields; v45 and v46 patch 1 each corrected a specific bug that had left machines stuck. v47 patch 1 changes what a rebuild produces, not just how a rebuild is triggered: the mounted image is stripped to zero third-party drivers, proven by re-enumeration, and only then is the current driver recipe injected. The deployed WIM bytes therefore differ from v46’s on the same inputs. The version bump is the mechanism by which the fleet converges on the strip-normalized driver set.

The v47 pass is not materially slower on a healthy machine than any other full-update pass. The strip stage on an already-clean image is a no-op — the ASUS field run logged Strip: image already has zero third-party drivers in under a second. On a machine whose registered image carries the OEM/VMD drivers from an earlier WinRE Manager cycle, the strip removes them and the current recipe is injected; the total time depends on the driver set size, not on the strip itself.

The distinguishing property of this migration is the metadata anchor. Every machine’s first v47 run writes a DeployedWinREMetadata field to the state file — the DISM servicing metadata (Version + SPBuild) of the WIM that was actually deployed. The field is the anchor the v47 drift detector compares against on every subsequent run. A machine whose state file lacks the anchor (all v46-or-earlier state files) forces a rebuild regardless of whether any other input changed; a machine whose state file has the anchor is evaluated against it.

Concretely, on the first run after the update, every machine will:

  1. Compute the new DesiredStateId (with SCRIPT=47).
  2. Compare it to the state file written by v46 (any patch).
  3. Find a mismatch, set needInject = $true, and take the full-update path.
  4. Obtain a base WIM via the v47 three-source preference order (registered image → hash-validated LKG → GitHub).
  5. Strip the mounted image to zero third-party drivers, proving zero by re-enumeration.
  6. Inject the current OEM and VMD recipe, run ResetBase, export the optimized WIM.
  7. Deploy, register, and write the state file with the new DSI and the DeployedWinREMetadata anchor.
  8. Return to the fast path on the second run.

Three cases handled without operator attention:

The offline-migration case is one step longer. On a machine whose first v47 run is offline, the state file lacks the anchor, $needInject is forced to $true, and the offline guard fires — the run exits EXIT_WARNING without a rebuild. The first successful online run writes the anchor; from the second v47 run onward, the offline fast path works normally. If a machine in your fleet runs offline for an extended period, expect its first online v47 run to be the one that writes the anchor.

Rollback

Change $ScriptVersion back to 46 and revert the strip stage, the three-source selection logic, the race detector, and the metadata-drift changes. The state file written under the v47 DSI becomes stale, the next run takes the full-update path under v46’s behavior, and the strip-normalized image is replaced with a lineage-seeded one. The DeployedWinREMetadata field becomes inert on rollback — it is read only by the v47 drift detector. No data is lost; another fleet-wide rebuild occurs on the next run.

v48 patch 1 migration

The v48 patch 1 revision is the largest single architectural change to the manager since v43. It bundles five distinct changes: intervening-partition handling, the architecture gate, the LocalInputsId state-file field, LKG-by-hash at any discovered recovery location, and transactional WIM replacement. It bumps ScriptVersion from 47 to 48, which changes the SCRIPT component of the DesiredStateId and forces one full-update pass per managed machine on the next scheduled run — the same shape as the v44 patch 1, v45 patch 1, v46 patch 1, and v47 patch 1 migrations.

The v48 pass is not materially slower on a healthy machine than any other full-update pass. On the ordinary non-intervening path, the pre-shrink target is C: and the pipeline is unchanged from v47; the transactional WIM replacement adds one hash of the existing target WIM (roughly 1-3 seconds on an NVMe) before the delete-then-copy sequence. On the intervening-anchor path, the pre-shrink target is the anchor instead of C:, and the anchor resize adds roughly one additional minute of I/O on a 720 GiB data partition.

Concretely, on the first run after the update, every machine will:

  1. Compute the new DesiredStateId (with SCRIPT=48).
  2. Compare it to the state file written by v47 (any patch).
  3. Find a mismatch, set needInject = $true, and take the full-update path.
  4. Run the architecture gate (must pass on x64; any other value exits EXIT_WARNING).
  5. Compute and store the LocalInputsId in the state file on the state write at the end of the run.
  6. Obtain a base WIM via the v47 three-source preference order (unchanged).
  7. Strip and inject as in v47 (unchanged).
  8. Plan geometry: on the ordinary layout, C: is the resize target; on an intervening-anchor layout (C: | D: | Recovery), D: is the resize target and is validated before any mutation.
  9. Deploy transactionally: the existing target WIM is preserved as a rollback copy in the workspace, the target is deleted, the new WIM is copied and hash-verified, and on any copy failure the rollback copy is restored.
  10. Write the state file with the new DSI, the LocalInputsId, and (as in v47) the DeployedWinREMetadata anchor.
  11. Return to the fast path on the second run.

Three cases handled without operator attention:

Two cases where the v48 pipeline defers where v47 deferred differently:

Rollback

Change $ScriptVersion back to 47 and revert the five bundled changes: the intervening-anchor plan logic, the architecture gate, the LocalInputsId computation and its state-file field, the LKG-by-hash discovery scan in Get-LKGWinREImagePath, and the transactional WIM replacement. The state file written under the v48 DSI becomes stale, the next run takes the full-update path under v47’s behavior, and the LocalInputsId field becomes inert. No data is lost; another fleet-wide rebuild occurs on the next run.

Later patches (no ScriptVersion bump)

The v44 patches 2 through 8 and the v46 patch 2 additions do not bump ScriptVersion and do not change the DesiredStateId. They apply to every subsequent run without a state-file action. A machine that is already healthy continues to take the fast path. A machine that was mid-deployment when the patch rolled out continues from where it was — the checkpoint and state-file schemas are unchanged.

The v47 patch 2 revision is also not a fleet-wide rebuild: it does not bump ScriptVersion, and its only DSI-value change is a one-time event on a narrow machine class (machines with a padded Win32_ComputerSystemProduct.Version field). It is documented under ### v47 patch 2 specifically below. The v47 patch 1 revision is not in this list; it bumps ScriptVersion and is documented under “v47 patch 1 migration” above.

The v48 patch 2 revision is also not a fleet-wide rebuild: it does not bump ScriptVersion and does not change the DesiredStateId. It adds the fail-fast elevation guard, a source-WIM hash cache that closes a pre-existing enable-only fallthrough path, a resize-target resynchronisation fix on the extension-failure fallback path (the D1 fix), and several cosmetic log and comment corrections. On a machine that ran v48 patch 1, the patch 2 build takes the fast path unchanged; the only observable difference is the startup banner and, on an unelevated launch, the millisecond refusal instead of a delayed Mount-WindowsImage failure. Documented under ### v48 patch 2 specifically below.

Because none of these patches changes ScriptVersion or the DesiredStateId, an already-completed machine will not rerun automatically; the deployment mechanism must invoke the script explicitly to pick the fixes up. The next natural rebuild (manifest bump, OEM pack version change, Windows build change, or CPU/VMD presence change) picks them up regardless.

Rollback for later non-bumping patches

None of these rollbacks affects the state-file schema or the DesiredStateId. (The v47 patch 1 rollback does affect the schema — it introduces the DeployedWinREMetadata field — and is documented separately under “v47 patch 1 migration” above.)

Exit code handling

The script’s exit codes carry meaning. See exit-codes.md for the full matrix.

For MDM / orchestration:

Exit code Recommended action
0 None. Record success.
1 Reboot the machine at the next convenient window. The script will finish on the next boot.
2 Investigate. WinRE is functional but degraded, or the script deferred work, or the enable step failed and the counter incremented. Collect the log and check the state file’s LastUpdated timestamp, its LastEnableResult field, whether a deferral marker exists at C:\Recovery\OEM\winre_partition_deferred.json, and the state file’s LocalInputsId. Sixteen distinct cases are documented in exit-codes.md: OS-fallback, v47 strip-failure abort, v47 patch 2 base-WIM copy-integrity check, Step 3 → Step 4 pipeline gate, Audit Mode deferral, VMD-query-indeterminate deferral, OS-fallback BitLocker deferral, v45 pre-shrink deferral, enable-only failure, concurrent-instance deferral (v44 patch 4), offline-fallback deferral (v44 patch 5), v47 race-detector abort at the two /disable sites, v48 architecture-gate refusal, v48 intervening-anchor surplus rejection, v48 multi-intervening rejection, and v48 offline LocalInputsId mismatch. The concurrent-instance case is not a failure — the other instance is doing the work. The offline-fallback fast-path case is a degraded-success — the machine is healthy and unchanged. The VMD-query-indeterminate, v45 pre-shrink, v47 race-detector, v48 architecture-gate, v48 surplus, v48 multi-intervening, and v48 LocalInputsId cases are protective deferrals — no state was committed.
3 Investigate. The run failed and did not write state, or the enable-failure loop-breaker fired. Collect the log. Do not retry automatically. One exception: the v48 patch 2 elevation guard exits 3 with FATAL: WinRE Manager requires an elevated (Administrator) PowerShell session. — this is expected behavior on an unelevated launch, not a failure. Relaunch from an elevated prompt or via scripts\WinRE-Manager.cmd.

Do not treat exit code 2 as success. A machine in OS-fallback is intentionally reported as a warning; it will be treated as a healthy machine by the fast path only if the state file records UsedOSFallback = true for the current DesiredStateId. A machine on which the Audit Mode guard or a v45 pre-shrink deferral fired will have an unchanged (or absent) state file and a single deferral line in the log. See exit-codes.md for how to distinguish the sixteen cases.

Do not configure retry loops that ignore the exit code and re-run unconditionally. Deferrals are not improved by retrying — the same gate will fire on the next run (Rule 4). Enable failures are handled by the counter; after three consecutive failures the loop-breaker fires and requires manual intervention. The concurrent-instance deferral is not a failure at all; retrying while the other instance is still running will simply produce another EXIT_WARNING. The VMD-query-indeterminate deferral is not improved by retrying — the underlying PnP service issue must be resolved. The v45 pre-shrink deferral is not improved by retrying until the constraint is resolved (free space on C:, disk layout, oversized recovery partition, or shrink failure). The one exception is the transient volume-read failure: it returns RetrySuppressible = $false and does not write a deferral marker, so the next scheduled run retries it automatically without operator intervention. The v47 race-detector abort is not improved by retrying immediately — the underlying Windows Update servicing of the registered WinRE is a transient condition, and the next scheduled run (or a manual re-run after the WU settles) is the correct response. The v47 strip-failure abort is not improved by retrying until the specific strip failure has been diagnosed — see troubleshooting.md for the specific failure modes. The v48 architecture-gate refusal is a stable deferral: any machine whose architecture is not x64 will continue to exit 2 on every run until the manager adds support or the machine is replaced. The v48 intervening-anchor surplus and multi-intervening rejections are stable layout deferrals; a retry produces the same rejection until the layout is changed. The v48 offline LocalInputsId mismatch resolves on the next run with network, which recomputes the DSI. Resolve the underlying condition, then re-run.

MDM / Intune

Deploy via a Win32 app or a PowerShell script platform script. The Intune “Scripts and remediations” feature is well-suited:

For a curated deployment, package the script as an Intune Win32 app with a detection rule on the log file and reagentc state. Run once, then allow the built-in weekly scheduled task to keep the state fresh.

When using Intune to push a remediation, be aware that the remediation script may run outside the scheduled task’s IgnoreNew gate. As of v44 patch 4, the program lock handles this: if the weekly task happens to be running at the same moment, the remediation exits with EXIT_WARNING (code 2) and the message Another WinRE Manager instance is already running (program lock file is exclusively held). This is not a failure. The remediation will succeed on the next cycle, or you can wait for the scheduled task to complete and re-run the remediation manually.

Do not disable the lock to “fix” the remediation. Before patch 4, the same collision caused an EXIT_FATAL (code 3) with a rename error. Some Intune remediation pipelines treated that as a hard failure and retried aggressively, which could produce a cascade of colliding runs. The lock converts the collision into a clean, self-limiting EXIT_WARNING. If your Intune pipeline retries on any non-zero exit, adjust it to treat code 2 as a soft failure (route to a queue) and code 3 as a hard failure (investigate). See exit-codes.md for the recommended orchestration policy.

Hosting your own maps, manifest, and base WIM repository

By default, WinRE Manager fetches the driver manifest, the three OEM maps, and the base WIM repository from the project maintainer’s GitHub account. All five artifacts can be replaced with self-hosted equivalents so your fleet does not depend on those external endpoints.

See self-hosting.md for the trust model, the exact URLs and variables to change, the map-builder scripts, and the air-gapped deployment procedure.

Log collection

The log at C:\ProgramData\OEM\Logs\WinRE-Manager.log is append-only and grows unbounded. Rotate it via your existing log-collection pipeline:

At ~100 log lines per healthy run, an unmanaged machine generates ~5 MB per year. Rotation is not urgent but is worth having.

Canary and ring deployment

The script is destructive on the recovery partition and non-destructive on the OS partition, but it does modify the OS partition geometry on the shrink path. Roll out in rings:

  1. Canary ring. One or two machines, GPT. Under the target-volume policy, C:’s BitLocker state may be any of fully decrypted, suspended, or mid-encryption; only the OS-fallback route depends on it, and a healthy canary machine will not take that route. Run Test-WinRE.ps1 first, then WinRE.ps1 -DryRun, then WinRE.ps1. Verify exit code 0. A transient exit code 2 from a concurrent scheduled task that happened to be running is not a failure — see “One instance per machine” above.
  2. Early ring. ~5% of the fleet. Mix of vendors and partition styles.
  3. Broad ring. The rest.

v44 patch 1 specifically

Add a step 0 before the canary ring: run the current script on one machine and confirm the full-update pass completes and the state file is rewritten under the new DesiredStateId. Then roll out normally. The Migration Note in CHANGELOG.md and the “v44 patch 1 migration” section above describe the expected behaviour.

v44 patch 3 specifically

No state-file action is needed on healthy machines — the surviving part of the fix (the base.wim cleanup) is exercised on the next full-update pass, and the guard portion is removed by v44 patch 6 anyway. See “Later patches (no ScriptVersion bump)” below.

v44 patch 4 specifically

Watch for the new EXIT_WARNING from concurrent invocations in your orchestration logs. A cluster of these on a machine usually means an RMM tool or a manual operation overlapped with the scheduled task; it is a signal about the operator side, not a script defect.

v44 patch 5 specifically

The first observable change is faster failure on machines with network issues. Machines that previously exited EXIT_FATAL after ~100-second timeouts now exit EXIT_WARNING after ~15-second timeouts, and healthy machines take the fast path offline instead of failing. Confirm on a canary that an offline run exits 2 with the Offline fallback: using state file's stored DesiredStateId … line rather than 3.

v44 patch 6 specifically

Watch for:

v44 patch 7 specifically

Watch for machines that exit 0 after one full-update pass where they previously cycled through rebuilds. Before patch 7, a machine with a Basic Data partition labelled “Recovery” on the OS disk never converged: the fast-path count included the label-only partition, so the “exactly one recovery partition on the OS disk” condition was never satisfied, while the destructive path correctly refused to delete it. After patch 7 the classifier, count, and destructive path all agree on the type-code rule, so the machine converges — either to a proper DEDICATED state (if the OS-shrink succeeds) or stably to OS-fallback (if it does not).

v44 patch 8 specifically

No state-file action and no observable orchestration change. The reorder is an internal cost reduction for healthy paths. Confirm on a canary that a fast-path run still exits 0 with the same log shape as before.

v45 patch 1 specifically

The v45 revision bumps ScriptVersion and forces one full-update pass per managed machine. The pass itself is not slower than a normal full-update pass. Watch for:

Exit-code observations. Watch for these exit-code patterns as well:

v46 patch 1 specifically

The v46 patch 1 revision clamps tailEnd to diskSize − 1 MiB before aligning in Get-PartitionPlan, and bumps ScriptVersion from 45 to 46. It forces one full-update pass per managed machine on the next scheduled run — the same shape as the v44 patch 1 and v45 patch 1 migrations. The pass itself is not slower than a normal full-update pass.

The v46 patch 1 canary is primarily about confirming that machines the v45 plan-clamp bug left in OS-fallback recover. Watch for:

v46 patch 2 specifically

The v46 patch 2 revision does not bump ScriptVersion and does not change the DesiredStateId. It adds three build-drift log lines, a small read guard in Get-RecoveryPartitions, and four post-review hardenings of Ensure-AdequateRecoveryPartition. The canary is primarily observational:

v47 patch 1 specifically

The v47 patch 1 revision introduces the third-party driver strip stage and bumps ScriptVersion from 46 to 47. It forces one full-update pass per managed machine on the next scheduled run — the same shape as the earlier ScriptVersion-bump migrations. The pass itself is not materially slower than a normal full-update pass; on an already-clean image, the strip stage is a no-op.

Watch for:

Also check the ASUS PRIME H510M-D v47 field-verification transcript in CHANGELOG.md for the expected shape of a healthy migration run — the log lines the canary should produce are the same ones.

v47 patch 2 specifically

The v47 patch 2 revision does not bump ScriptVersion and does not change the DesiredStateId on any machine whose Win32_ComputerSystemProduct.Version field is not padded. It ships six hardening fixes on top of v47 patch 1: the OS-fallback missing-WIM guard without a status restriction; the base-WIM copy-integrity check; the machine-type fallback trim; the workspace candidate filter IsSystem -and -not IsBoot; the GitHub base-WIM extraction contract change (return $false instead of throw); and the -SuppressSuccessLog switch on the duplicate Verified partition log line. The canary is primarily observational:

v48 patch 1 specifically

The v48 patch 1 revision bumps ScriptVersion from 47 to 48 and forces one full-update pass per managed machine. The pass is not materially slower than any other full-update pass on the ordinary non-intervening path; the intervening-anchor path adds roughly one additional minute of anchor resize on the anchored machine class.

Watch for:

v48 patch 2 specifically

The v48 patch 2 revision does not bump ScriptVersion and does not change the DesiredStateId. It adds the fail-fast elevation guard, a source-WIM hash cache that closes a pre-existing enable-only fallthrough path, a resize-target resynchronisation fix on the extension-failure fallback path (the D1 fix), and several cosmetic log and comment corrections. The canary is primarily observational:

Rolling back

The script does not have an uninstall path. To disable it:

schtasks /Delete /TN "WinRE Manager" /F
schtasks /Delete /TN "WinRE Manager Weekly" /F

The state file, deferral marker, log, lock file, and any deployed recovery partition remain. The machine is in a healthy end state and Windows Update will continue to service the recovery image normally. The lock file at C:\ProgramData\OEM\Logs\WinREManager.lock is inert once the scheduled task is removed; it can be deleted manually if desired, and it will be recreated if the script is ever run again. The deferral marker at C:\Recovery\OEM\winre_partition_deferred.json is likewise inert; it can be deleted manually.

If you need to revert a machine to its pre-WinRE-Manager state, restore the partition layout from a backup. The script does not create one.

To revert a DesiredStateId change specifically, see the “Rollback” subsections under “v44 patch 1 migration”, “v45 patch 1 migration”, “v46 patch 1 migration”, “v47 patch 1 migration”, and “v48 patch 1 migration” above. To revert code from a patch that does not change the DesiredStateId, see the “Rollback for later non-bumping patches” subsection above.