Skip to content

TextEncoder.encodeInto NOT work on BIG utf8 subarray ON 22.4.1 #62610

Description

@navegador5

Version

Welcome to Node.js v22.4.1

Platform

Linux dev 6.8.0-88-generic #89-Ubuntu SMP PREEMPT_DYNAMIC Sat Oct 11 01:02:46 UTC 2025 x86_64 x86_64 x86_64 GNU/Linux

Subsystem

No response

What steps will reproduce the bug?

`
const _TE = new TextEncoder();
var s = "aÿ我𝑒"
var u8a = new Uint8Array(2429682061);
var offset = 38928786 ;
_TE.encodeInto(s,u8a.subarray(offset)) // { read: 0, written: 0 } failed
_TE.encodeInto(s,u8a.subarray(offset,offset+10)) //{ read: 5, written: 10 } success

Welcome to Node.js v22.4.1.
Type ".help" for more information.

const _TE = new TextEncoder();
undefined
var s = "aÿ我𝑒"
undefined
var u8a = new Uint8Array(2429682061);
undefined
var offset = 38928786 ;
undefined
_TE.encodeInto(s,u8a.subarray(offset)) // { read: 0, written: 0 } failed
{ read: 0, written: 0 }
_TE.encodeInto(s,u8a.subarray(offset,offset+10)) //{ read: 5, written: 10 } success
{ read: 5, written: 10 }

Welcome to Node.js v24.14.0.
Type ".help" for more information.

const _TE = new TextEncoder();
undefined
var s = "aÿ我𝑒"
undefined
var u8a = new Uint8Array(2429682061);
undefined
var offset = 38928786 ;
undefined
_TE.encodeInto(s,u8a.subarray(offset)) // { read: 0, written: 0 } success
{ read: 5, written: 10 }
_TE.encodeInto(s,u8a.subarray(offset,offset+10)) //{ read: 5, written: 10 } success
{ read: 5, written: 10 }

`

How often does it reproduce? Is there a required condition?

alaways in 22.4.1

What is the expected behavior? Why is that the expected behavior?

same as 24.14.0

What do you see instead?

const _TE = new TextEncoder();
var s = "aÿ我𝑒"
var u8a = new Uint8Array(2429682061);
var offset = 38928786 ;
_TE.encodeInto(s,u8a.subarray(offset)) // 【{ read: 0, written: 0 } failed】
_TE.encodeInto(s,u8a.subarray(offset,offset+10)) //{ read: 5, written: 10 } success

Additional information

No response

Activity

  1. semimikoh commented on Apr 6, 2026

    @semimikoh
    Contributor

    Root cause

    This is a 32-bit integer overflow in EncodeInto in
    src/encoding_binding.cc:

    ```cpp
    size_t dest_length = dest->ByteLength(); // size_t — 2,390,753,275 OK

    int nchars;
    int written = source->WriteUtf8(
    isolate,
    write_result,
    dest_length, // implicit narrowing to int
    &nchars,
    String::NO_NULL_TERMINATION | String::REPLACE_INVALID_UTF8);
    ```

    v8::String::WriteUtf8 takes the capacity as an int. When
    dest_length exceeds INT32_MAX (2,147,483,647), the narrowing
    conversion underflows to a negative number, V8 treats it as "no
    capacity", and writes 0 bytes — hence { read: 0, written: 0 }.

    The subarray(offset, offset+10) case works because the view is
    only 10 bytes, well within int range.

    Already fixed on main / v24

    This was incidentally fixed by #58070 (src: use non-deprecated WriteUtf8V2() method), which migrated to WriteUtf8V2 whose
    capacity parameter is size_t. The intent of that PR was
    deprecation cleanup, not bug fixing, but it resolved this issue as
    a side effect.

    • v24.0.0+: fixed ✅
    • v22.x: still affected ❌ (int capacity still in place)

    v22 LTS backport?

    v22 is in maintenance LTS until 2027-04 and this causes silent data
    loss (not an error) for users with >2GB buffers. Would a minimal
    backport patch be acceptable — clamping the capacity to INT_MAX
    in EncodeInto specifically, rather than backporting the full
    WriteUtf8V2 migration?

    Happy to send a PR once the direction is confirmed.

  2. navegador5 commented on Apr 6, 2026

    @navegador5
    Author

    Root cause

    This is a 32-bit integer overflow in EncodeInto in src/encoding_binding.cc:

    int nchars; int written = source->WriteUtf8( isolate, write_result, dest_length, // implicit narrowing to int &nchars, String::NO_NULL_TERMINATION | String::REPLACE_INVALID_UTF8); ```
    
    `v8::String::WriteUtf8` takes the capacity as an `int`. When `dest_length` exceeds `INT32_MAX` (2,147,483,647), the narrowing conversion underflows to a negative number, V8 treats it as "no capacity", and writes 0 bytes — hence `{ read: 0, written: 0 }`.
    
    The `subarray(offset, offset+10)` case works because the view is only 10 bytes, well within int range.
    
    ## Already fixed on main / v24
    This was incidentally fixed by [#58070](https://github.com/nodejs/node/pull/58070) (`src: use non-deprecated WriteUtf8V2() method`), which migrated to `WriteUtf8V2` whose capacity parameter is `size_t`. The intent of that PR was deprecation cleanup, not bug fixing, but it resolved this issue as a side effect.
    
    * v24.0.0+: fixed ✅
    * v22.x: still affected ❌ (`int` capacity still in place)
    
    ## v22 LTS backport?
    v22 is in maintenance LTS until 2027-04 and this causes silent data loss (not an error) for users with >2GB buffers. Would a minimal backport patch be acceptable — clamping the capacity to `INT_MAX` in `EncodeInto` specifically, rather than backporting the full WriteUtf8V2 migration?
    
    Happy to send a PR once the direction is confirmed.

    v8 12.4.254(node 22.4.1) have NO WriteUtf8V2 , i think in 22.4 it should give a warning log

  3. added
    encodingIssues and PRs related to the TextEncoder and TextDecoder APIs.
    v22.xIssues that can be reproduced on v22.x or PRs targeting the v22.x-staging branch.
    on Apr 6, 2026
  4. github-actions commented on Jul 31, 2026

    @github-actions
    Contributor

    This issue has been marked as stale due to 90 days of inactivity.
    It will be automatically closed in 30 days if no further activity occurs. If this is still relevant, please leave a comment or update it to keep it open.

  5. added
    staleIssues and PRs marked stale due to inactivity and scheduled for automatic closure.
    on Jul 31, 2026
  6. richardlau commented on Aug 3, 2026

    @richardlau
    Member

    This was fixed by #62621 in Node.js 22.22.3.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    encodingIssues and PRs related to the TextEncoder and TextDecoder APIs.staleIssues and PRs marked stale due to inactivity and scheduled for automatic closure.v22.xIssues that can be reproduced on v22.x or PRs targeting the v22.x-staging branch.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions