Skip to content

WSL: stale run\wsl-daemons record from pre-update instance blocks reconnect ("another live app instance already manages WSL distro") #4391

Description

@johnnyhaggs

Short summary

After a self-update on Windows, the relaunched app refuses every WSL environment with another live app instance already manages WSL distro. The cause is the record that the previous, now-dead instance left in %USERPROFILE%\.copilot\run\wsl-daemons\. The refusal persists across app restarts until that file is removed by hand. After removal, reconnect works immediately with no restart.

Affected version or release

v1.1.26 (commit 5ab79c8), updated from v1.1.25 (6ab388f). Also hit on the v1.1.24 → v1.1.25 update.

Installation context

Windows 11 (10.0.26200), WSL 2.9.4.0 (kernel 6.18.35.2-1). Two WSL environments, Ubuntu and <custom-distro>, both Ubuntu 24.04. copilotd 0.9.0 → 0.9.2.

What happened?

Timeline (UTC) for the 1.1.25 → 1.1.26 update:

Time Event
10-01 20:07:58 1.1.25 starts (pid 59952), connects both distros, writes run\wsl-daemons\59952-134353588756033422.json
10-01 22:08:13 Update staged, waiting for user to restart version=1.1.26
10-02 12:13:40 User confirmed staged update → Created pre-update backup. This is the last line in the 1.1.25 log. No shutdown or WSL teardown is logged, and the record is not removed.
10-02 12:14:34 1.1.26 starts (pid 40520). Both distros are refused about 100 ms later (below).
10-02 12:17:23 Manual app restart (pid 60836). Same refusal at startup and on every session resume through 13:24.
10-02 13:26:53 Stale JSON moved out of run\wsl-daemons\ (app left running)
10-02 13:27:15 Next session resume: copilotd 0.9.2 installed, wsl rehydrate complete

The stale record:

{"owner_pid":59952,"owner_executable":"...\\GitHub Copilot\\github.exe","owner_creation_time":134353588756033422,
 "daemons":[{"pid":64196,"distro":"<custom-distro>","executable":"C:\\Windows\\System32\\wsl.exe","creation_time":134353595937668693},
            {"pid":5628,"distro":"Ubuntu","executable":"C:\\Windows\\System32\\wsl.exe","creation_time":134353596493940384}]}

Errors:

WARN wsl::rehydrate: wsl rehydrate failed env_id=wsl:Ubuntu error=another live app instance already manages WSL distro `Ubuntu`
WARN wsl::rehydrate: wsl rehydrate failed env_id=wsl:<custom-distro> error=another live app instance already manages WSL distro `<custom-distro>`
WARN handlers::session: failed to resume session ... error=remote host `wsl:<custom-distro>` recovery failed: another live app instance already manages WSL distro `<custom-distro>`
WARN workspace::live_state::ahp: AHP producer: could not build host env ... error=host `wsl:<custom-distro>` is not registered or unavailable

State checked while the error was still occurring (around 13:25):

  • owner_pid 59952 was not running.
  • Daemon pids 64196 and 5628 were not running.
  • Exactly one github.exe (60836) was running.
  • No file under run\ was locked except single-instance.lock, which 60836 held.

So the "live instance" check treats a record whose owner is dead as live. The 1.1.25 instance also exited through the update-restart path without removing its record.

Second occurrence (1.1.24 → 1.1.25, 9/30): The 1.1.24 log (pid 62052) ends abruptly at the update at 00:09:52 UTC. The first 1.1.25 launch, 66 s later, logged another live app instance already manages WSL distro for the same distro. That record (62052-…json) stayed in place until it was removed by hand.

Steps to reproduce

  1. On Windows, configure and connect at least one WSL environment. %USERPROFILE%\.copilot\run\wsl-daemons\<pid>-<ctime>.json now exists.
  2. Let the app stage an update, then confirm the restart.
  3. After relaunch, open a session on the WSL environment.
  4. The connection fails, and the log shows another live app instance already manages WSL distro. Restarting the app doesn't help.
  5. Move the old wsl-daemons\*.json whose owner_pid isn't running. The next session resume reconnects.

Expected behavior

  • An instance that exits, including through the update-restart path, removes its wsl-daemons record after reaping its daemons.
  • On startup or rehydrate, a record whose owner_pid isn't running, or whose process creation time doesn't match owner_creation_time, is treated as stale. The new instance reaps any listed daemons that are still alive, then takes ownership instead of refusing.

Additional context

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions