platform/surface: Add s2idle/hibernate fix for Surface Laptop 5 - #162
wowitsjack wants to merge 7 commits into
Conversation
Intel INTC1055 pinctrl power-gating during s2idle corrupts pin 213's PADCFG0 register. The corruption fires a spurious SCI on GPE 0x52, which promotes to a full resume via acpi_any_gpe_status_set(). The system either never wakes or wakes spuriously on every s2idle cycle. Fix by intercepting in the ACPI wakeup handler (which runs before acpi_ec_dispatch_gpe). The handler repairs PADCFG corruption, clears GPE 0x52 status, and returns false. The actual lid-open decision is deferred to an LPS0 check() callback which reads RXSTATE directly. GPE 0x52 is kept masked during s2idle to prevent VNN RXSTATE glitches from firing as SCI. Lid-open detection uses a 500ms poll timer reading RXSTATE directly. GPE unmask is delayed to PM_POST_SUSPEND to avoid stale GPE status causing false lid events during resume. Additional features: failsafe re-suspend with exponential backoff, RXSTATE polling for lid close detection, suspend-to-lock mode, hibernate support with RTC time sync, and full PADCFG0 corruption detection with 15s RXSTATE settling guard. Link: linux-surface/linux-surface#1782 Signed-off-by: Jack <wowitsjack@users.noreply.github.com>
Removed old suspected root cause explanation for s2idle issue.
Between resume_noirq and resume_early, pinctrl-intel and ACPI methods legitimately toggle RXINV. fix_padcfg_corruption() compared against our stale pre-suspend saved state and 'corrected' RXINV back, triggering restore_pin() which set last_padcfg_restore_time. This started the 15s RXSTATE settling suppression in lid_poll_fn, making the lid appear stuck closed for 15 seconds after every s2idle wake. Real VNN corruption is already handled by lps0_check (during s2idle) and resume_noirq (before ACPI touches RXINV).
- Call _L52 in PM_POST_SUSPEND to sync ACPI lid state immediately (RXSTATE is fresh and reliable before VNN settling starts) - Reduce RXSTATE settling window from 15s to 3s - GPE unmask after settling for natural lid close detection - Poller calls _L52 manually on RXSTATE transitions as fallback - suspend_to_lock: check RXINV to distinguish lid-triggered vs manual suspend (DE button, systemctl suspend). Only intercept lid-bounce suspends, allow manual suspends through. - Debounce GPIO switch bounce in suspend_to_lock (5 polls x 200ms)
… glitches Remove all _L52 calls (PM_POST_SUSPEND, poller open path) to eliminate ghost suspends from logind's 30s lid-ignore buffer. Add RXSTATE guard checks in lid_close_fn and failsafe_fn before pm_suspend to prevent hardware input latch engaging with lid open. Fix suspend_to_lock to trust lid_close_fn's guard for lid-triggered suspends (VNN glitches caused alternating interceptions), and block external ghost suspends when lid is physically open.
Send KEY_WAKEUP in PM_POST_SUSPEND when lid is open to wake display (report_lid_state(0) was a no-op since SW_LID was already 0). Split ghost suspend blocker to silently cancel without locking screen, killing wifi, or blanking display.
Restore GPE 0x52 unmask after s2idle wake (matching original PR162) so ACPI button driver handles lid events naturally. Restore report_lid_state on both close and open for proper SW_LID 1->0 transition that GNOME uses to wake the display. Add ghost suspend blocker in PM_SUSPEND_PREPARE to silently cancel external open-lid suspends. Add RXSTATE guard checks in lid_close_fn and failsafe_fn to prevent hardware latch death sleep. Fix backlight restore to force FB_BLANK_UNBLANK before setting brightness.
| static void lock_restore_fn(struct work_struct *work) | ||
| { | ||
| cancel_delayed_work_sync(&lock_blank_work); | ||
|
|
||
| if (lock_bl_dev) { | ||
| if (lock_saved_brightness >= 0) | ||
| backlight_device_set_brightness(lock_bl_dev, | ||
| lock_saved_brightness); | ||
| put_device(&lock_bl_dev->dev); | ||
| lock_bl_dev = NULL; | ||
| lock_saved_brightness = -1; | ||
| } | ||
|
|
||
| { | ||
| static char *argv[] = { | ||
| "/usr/bin/nmcli", "networking", "on", NULL | ||
| }; | ||
| static char *envp[] = { | ||
| "HOME=/root", | ||
| "PATH=/usr/bin:/bin", | ||
| NULL | ||
| }; | ||
| call_usermodehelper(argv[0], argv, envp, UMH_NO_WAIT); | ||
| } | ||
|
|
||
| pr_info("suspend_to_lock: input detected, restoring display and networking\n"); | ||
| } |
There was a problem hiding this comment.
Kernel code should never be reliant on userland tooling. This can be a problem for a few reasons.
This code assumes Network Manager is installed, and that it exists at the path of /usr/bin/nmcli. For example, what if I was using systemd-networkd for networking instead? Kernel code should never rely on something in userspace being installed.
There was a problem hiding this comment.
True. This should be handled with something like a systemd service or some other trigger, entirely in userspace.
There was a problem hiding this comment.
I think ideally the root cause should be found, preventing any interaction in user space that isn't required. But this could work as a temporary solution.
|
ALRIGHT updates time, take it away, Codex: Follow-up: smaller PoC, measured S0ix residency, and screen wakeFurther testing on a Surface Laptop 5, i7-1265U, Ubuntu 25.10, kernel I am currently testing a small, live-loaded SAM wake PoC ( Confirmed sleep-power blocker on this installation A local A system-sleep hook now saves those settings, enables C8/C10 only for suspend while the PoC is loaded, and restores the original audio-workaround settings on resume. Battery drain during sleep is now substantially improved and behaving properly in actual use. Hardware measurements:
This C-state restriction was local configuration. These results do not establish it as the cause of every SL5 sleep failure reported here. Screen wake The long cycle resumed, but still needed a keyboard press to light the screen. A GNOME session helper requests display power on and writes I added bounded post-resume diagnostics because the earlier “wake applied” log only confirmed that a request was sent. The latest approximately 14½-minute test is now waking the screen correctly. At the first diagnostic sample, roughly 250 ms after the helper's wake request:
Display power remained on through the 30-second diagnostic window. The latest change added observation, not a new wake mechanism. I am keeping this successful test separate from the earlier overnight result; overnight screen-wake reliability still needs repeat validation. What remains unproven The PoC binds as a SAM client and enables the SAM controller's wake capability. It also attempts a SAM The PoC itself has no PADCFG writes, lid-poll timer, forced re-suspend, or clock correction. These are live-loaded tests on one machine, not a clean-boot or cross-device validation of a replacement kernel patch. No new hibernate validation is included in this update. The useful result so far is that normal low-power sleep is working without the large |
Closes linux-surface/linux-surface#1782
Per request from @qzed in linux-surface/linux-surface#2011
Adds
surface_s2idle_fixmodule todrivers/platform/surface/.Problem
Intel INTC1055 pinctrl power-gating during s2idle corrupts pin 213's PADCFG0 register. The corruption fires a spurious SCI on GPE 0x52, which promotes to a full resume via
acpi_any_gpe_status_set(). The system either never wakes (wakeup framework poisoned bypm_system_cancel_wakeup()) or wakes spuriously on every s2idle cycle.Fix
The ACPI wakeup handler runs before
acpi_ec_dispatch_gpe()in the s2idle wake decision path. The handler:acpi_any_gpe_status_set()doesn't see itLid-open detection is deferred to an LPS0
check()callback which reads RXSTATE directly and only callspm_system_wakeup()if the lid is genuinely open.Key design decisions
resume_early, causing false lid events to logind.acpi_enable_gpe/acpi_disable_gpe(notacpi_set_gpe) to maintain proper ACPI reference counting.Additional layers
Testing
Tested on Surface Laptop 5 (i5-1245U), Ubuntu 25.10, kernel 6.18.7-surface-1. s2idle cycles, hibernate cycles, extended sleep (11h+), repeated lid open/close, rapid-wake VNN cycles with failsafe re-suspend.