Skip to content

platform/surface: Add s2idle/hibernate fix for Surface Laptop 5 - #162

Open
wowitsjack wants to merge 7 commits into
linux-surface:v6.18-surface-develfrom
wowitsjack:surface-laptop-5-s2idle-fix
Open

wowitsjack wants to merge 7 commits into
linux-surface:v6.18-surface-develfrom
wowitsjack:surface-laptop-5-s2idle-fix

Conversation

@wowitsjack

Copy link
Copy Markdown

Closes linux-surface/linux-surface#1782
Per request from @qzed in linux-surface/linux-surface#2011

Adds surface_s2idle_fix module to drivers/platform/surface/.

Problem

Intel INTC1055 pinctrl power-gating during s2idle corrupts pin 213's PADCFG0 register. The corruption fires a spurious SCI on GPE 0x52, which promotes to a full resume via acpi_any_gpe_status_set(). The system either never wakes (wakeup framework poisoned by pm_system_cancel_wakeup()) or wakes spuriously on every s2idle cycle.

Fix

The ACPI wakeup handler runs before acpi_ec_dispatch_gpe() in the s2idle wake decision path. The handler:

  1. Repairs PADCFG corruption (full register comparison, not just RXINV)
  2. Clears GPE 0x52 status so acpi_any_gpe_status_set() doesn't see it
  3. Returns false (no full wake)

Lid-open detection is deferred to an LPS0 check() callback which reads RXSTATE directly and only calls pm_system_wakeup() if the lid is genuinely open.

Key design decisions

  • GPE 0x52 masked during s2idle: VNN power-gating causes transient RXSTATE glitches that fire GPE 0x52 as SCI. Lid detection handled by 500ms poll timer reading RXSTATE directly.
  • GPE unmask delayed to PM_POST_SUSPEND: Stale GPE status from lid state changes fires immediately on unmask during resume_early, causing false lid events to logind.
  • Reference-counted GPE management: Uses acpi_enable_gpe/acpi_disable_gpe (not acpi_set_gpe) to maintain proper ACPI reference counting.
  • Full PADCFG0 corruption detection: Compares entire register (minus volatile GPIORXSTATE) instead of just RXINV, catching all VNN corruption.
  • 1ms GPIORXDIS toggle + 15s settling guard: Extended input buffer enable for reliable re-latching after PADCFG restore.

Additional layers

  • Failsafe re-suspend with exponential backoff (2s/4s/8s/15s)
  • RXSTATE polling for lid close detection
  • Suspend-to-lock mode (converts lid-open suspend to lock+blank)
  • Hibernate support with RTC time sync
  • SW_LID input events for userspace lid policy

Testing

Tested on Surface Laptop 5 (i5-1245U), Ubuntu 25.10, kernel 6.18.7-surface-1. s2idle cycles, hibernate cycles, extended sleep (11h+), repeated lid open/close, rapid-wake VNN cycles with failsafe re-suspend.

Intel INTC1055 pinctrl power-gating during s2idle corrupts pin 213's
PADCFG0 register. The corruption fires a spurious SCI on GPE 0x52,
which promotes to a full resume via acpi_any_gpe_status_set(). The
system either never wakes or wakes spuriously on every s2idle cycle.

Fix by intercepting in the ACPI wakeup handler (which runs before
acpi_ec_dispatch_gpe). The handler repairs PADCFG corruption, clears
GPE 0x52 status, and returns false. The actual lid-open decision is
deferred to an LPS0 check() callback which reads RXSTATE directly.

GPE 0x52 is kept masked during s2idle to prevent VNN RXSTATE glitches
from firing as SCI. Lid-open detection uses a 500ms poll timer reading
RXSTATE directly. GPE unmask is delayed to PM_POST_SUSPEND to avoid
stale GPE status causing false lid events during resume.

Additional features: failsafe re-suspend with exponential backoff,
RXSTATE polling for lid close detection, suspend-to-lock mode,
hibernate support with RTC time sync, and full PADCFG0 corruption
detection with 15s RXSTATE settling guard.

Link: linux-surface/linux-surface#1782
Signed-off-by: Jack <wowitsjack@users.noreply.github.com>
wowitsjack and others added 6 commits March 14, 2026 01:54
Removed old suspected root cause explanation for s2idle issue.
Between resume_noirq and resume_early, pinctrl-intel and ACPI methods
legitimately toggle RXINV. fix_padcfg_corruption() compared against our
stale pre-suspend saved state and 'corrected' RXINV back, triggering
restore_pin() which set last_padcfg_restore_time. This started the 15s
RXSTATE settling suppression in lid_poll_fn, making the lid appear stuck
closed for 15 seconds after every s2idle wake.

Real VNN corruption is already handled by lps0_check (during s2idle)
and resume_noirq (before ACPI touches RXINV).
- Call _L52 in PM_POST_SUSPEND to sync ACPI lid state immediately
  (RXSTATE is fresh and reliable before VNN settling starts)
- Reduce RXSTATE settling window from 15s to 3s
- GPE unmask after settling for natural lid close detection
- Poller calls _L52 manually on RXSTATE transitions as fallback
- suspend_to_lock: check RXINV to distinguish lid-triggered vs
  manual suspend (DE button, systemctl suspend). Only intercept
  lid-bounce suspends, allow manual suspends through.
- Debounce GPIO switch bounce in suspend_to_lock (5 polls x 200ms)
… glitches

Remove all _L52 calls (PM_POST_SUSPEND, poller open path) to eliminate
ghost suspends from logind's 30s lid-ignore buffer. Add RXSTATE guard
checks in lid_close_fn and failsafe_fn before pm_suspend to prevent
hardware input latch engaging with lid open. Fix suspend_to_lock to
trust lid_close_fn's guard for lid-triggered suspends (VNN glitches
caused alternating interceptions), and block external ghost suspends
when lid is physically open.
Send KEY_WAKEUP in PM_POST_SUSPEND when lid is open to wake display
(report_lid_state(0) was a no-op since SW_LID was already 0). Split
ghost suspend blocker to silently cancel without locking screen,
killing wifi, or blanking display.
Restore GPE 0x52 unmask after s2idle wake (matching original PR162) so
ACPI button driver handles lid events naturally. Restore report_lid_state
on both close and open for proper SW_LID 1->0 transition that GNOME uses
to wake the display. Add ghost suspend blocker in PM_SUSPEND_PREPARE to
silently cancel external open-lid suspends. Add RXSTATE guard checks in
lid_close_fn and failsafe_fn to prevent hardware latch death sleep. Fix
backlight restore to force FB_BLANK_UNBLANK before setting brightness.
Comment on lines +1726 to +1752
static void lock_restore_fn(struct work_struct *work)
{
cancel_delayed_work_sync(&lock_blank_work);

if (lock_bl_dev) {
if (lock_saved_brightness >= 0)
backlight_device_set_brightness(lock_bl_dev,
lock_saved_brightness);
put_device(&lock_bl_dev->dev);
lock_bl_dev = NULL;
lock_saved_brightness = -1;
}

{
static char *argv[] = {
"/usr/bin/nmcli", "networking", "on", NULL
};
static char *envp[] = {
"HOME=/root",
"PATH=/usr/bin:/bin",
NULL
};
call_usermodehelper(argv[0], argv, envp, UMH_NO_WAIT);
}

pr_info("suspend_to_lock: input detected, restoring display and networking\n");
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Kernel code should never be reliant on userland tooling. This can be a problem for a few reasons.

This code assumes Network Manager is installed, and that it exists at the path of /usr/bin/nmcli. For example, what if I was using systemd-networkd for networking instead? Kernel code should never rely on something in userspace being installed.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

True. This should be handled with something like a systemd service or some other trigger, entirely in userspace.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. Hrm.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think ideally the root cause should be found, preventing any interaction in user space that isn't required. But this could work as a temporary solution.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@qzed Up to you, ultimately

@wowitsjack

wowitsjack commented Sep 25, 2026 •

Copy link
Copy Markdown
Author

ALRIGHT updates time, take it away, Codex:

Follow-up: smaller PoC, measured S0ix residency, and screen wake

Further testing on a Surface Laptop 5, i7-1265U, Ubuntu 25.10, kernel 6.18.7-surface-1, BIOS 22.102.143.

I am currently testing a small, live-loaded SAM wake PoC (surface_sam_wake_poc v0.3) with surface_s2idle_fix unloaded. surface_gpe remains loaded. The PR's code has not been replaced with this PoC.

Confirmed sleep-power blocker on this installation

A local disable-deep-cstates.service, installed to avoid audio glitches, disabled C8 and C10 on all 12 logical CPUs. Those disabled states were also excluded during s2idle. Before correcting this, the PMC S0ix counter and the package C10 counter stayed at zero despite long suspend intervals.

A system-sleep hook now saves those settings, enables C8/C10 only for suspend while the PoC is loaded, and restores the original audio-workaround settings on resume. Battery drain during sleep is now substantially improved and behaving properly in actual use.

Hardware measurements:

  • Initial timed check: 29,759,948 us of S0ix, previously zero.
  • An approximately 18 h 38 min lid-closed cycle: 67,042,438,082 us of S0ix, about 18 h 37 min of actual low-power residency.
  • Latest cycle: 874.329 s between the userspace sleep/resume signals, with 868,000,720 us of S0ix.

This C-state restriction was local configuration. These results do not establish it as the cause of every SL5 sleep failure reported here.

Screen wake

The long cycle resumed, but still needed a keyboard press to light the screen. A GNOME session helper requests display power on and writes KEY_WAKEUP through a dedicated PoC input device after resume, gated on the lid being open and logind no longer preparing for sleep. It does not unlock the session.

I added bounded post-resume diagnostics because the earlier “wake applied” log only confirmed that a request was sent. The latest approximately 14½-minute test is now waking the screen correctly. At the first diagnostic sample, roughly 250 ms after the helper's wake request:

  • Mutter PowerSaveMode=0;
  • backlight power 0 and nonzero actual brightness;
  • logind and UPower both reported the lid open;
  • Mutter idle time was 3 ms.

Display power remained on through the 30-second diagnostic window. The latest change added observation, not a new wake mechanism. I am keeping this successful test separate from the earlier overnight result; overnight screen-wake reliability still needs repeat validation.

What remains unproven

The PoC binds as a SAM client and enables the SAM controller's wake capability. It also attempts a SAM 0x17 request during resume, but that request still returns -ETIMEDOUT (-110). There is no successful queued-event-drain result, so I cannot attribute the improvement to that command or claim the underlying firmware cause is established.

The PoC itself has no PADCFG writes, lid-poll timer, forced re-suspend, or clock correction. These are live-loaded tests on one machine, not a clean-boot or cross-device validation of a replacement kernel patch. No new hibernate validation is included in this update.

The useful result so far is that normal low-power sleep is working without the large surface_s2idle_fix module loaded, and screen wake is working in the latest observed test. I will keep the remaining kernel-level questions distinct from the confirmed local C-state fix.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Surface Laptop 5 DMI ID Missing From GPE Modules List Leading to Failed Lid Suspend Events

3 participants