Repository navigation
TextEncoder.encodeInto NOT work on BIG utf8 subarray ON 22.4.1 #62610
Description
Activity
Root cause
This is a 32-bit integer overflow in
EncodeIntoin
src/encoding_binding.cc:```cpp
size_t dest_length = dest->ByteLength(); // size_t — 2,390,753,275 OKint nchars;
int written = source->WriteUtf8(
isolate,
write_result,
dest_length, // implicit narrowing to int
&nchars,
String::NO_NULL_TERMINATION | String::REPLACE_INVALID_UTF8);
```v8::String::WriteUtf8takes the capacity as anint. When
dest_lengthexceedsINT32_MAX(2,147,483,647), the narrowing
conversion underflows to a negative number, V8 treats it as "no
capacity", and writes 0 bytes — hence{ read: 0, written: 0 }.The
subarray(offset, offset+10)case works because the view is
only 10 bytes, well within int range.Already fixed on main / v24
This was incidentally fixed by #58070 (
src: use non-deprecated WriteUtf8V2() method), which migrated toWriteUtf8V2whose
capacity parameter issize_t. The intent of that PR was
deprecation cleanup, not bug fixing, but it resolved this issue as
a side effect.- v24.0.0+: fixed ✅
- v22.x: still affected ❌ (
intcapacity still in place)
v22 LTS backport?
v22 is in maintenance LTS until 2027-04 and this causes silent data
loss (not an error) for users with >2GB buffers. Would a minimal
backport patch be acceptable — clamping the capacity toINT_MAX
inEncodeIntospecifically, rather than backporting the full
WriteUtf8V2 migration?Happy to send a PR once the direction is confirmed.
Reacted by sangwookRoot cause
This is a 32-bit integer overflow in
EncodeIntoinsrc/encoding_binding.cc:int nchars; int written = source->WriteUtf8( isolate, write_result, dest_length, // implicit narrowing to int &nchars, String::NO_NULL_TERMINATION | String::REPLACE_INVALID_UTF8); ``` `v8::String::WriteUtf8` takes the capacity as an `int`. When `dest_length` exceeds `INT32_MAX` (2,147,483,647), the narrowing conversion underflows to a negative number, V8 treats it as "no capacity", and writes 0 bytes — hence `{ read: 0, written: 0 }`. The `subarray(offset, offset+10)` case works because the view is only 10 bytes, well within int range. ## Already fixed on main / v24 This was incidentally fixed by [#58070](https://github.com/nodejs/node/pull/58070) (`src: use non-deprecated WriteUtf8V2() method`), which migrated to `WriteUtf8V2` whose capacity parameter is `size_t`. The intent of that PR was deprecation cleanup, not bug fixing, but it resolved this issue as a side effect. * v24.0.0+: fixed ✅ * v22.x: still affected ❌ (`int` capacity still in place) ## v22 LTS backport? v22 is in maintenance LTS until 2027-04 and this causes silent data loss (not an error) for users with >2GB buffers. Would a minimal backport patch be acceptable — clamping the capacity to `INT_MAX` in `EncodeInto` specifically, rather than backporting the full WriteUtf8V2 migration? Happy to send a PR once the direction is confirmed.
v8 12.4.254(node 22.4.1) have NO WriteUtf8V2 , i think in 22.4 it should give a warning log
- addedencodingIssues and PRs related to the TextEncoder and TextDecoder APIs.Issues and PRs related to the TextEncoder and TextDecoder APIs.v22.xIssues that can be reproduced on v22.x or PRs targeting the v22.x-staging branch.Issues that can be reproduced on v22.x or PRs targeting the v22.x-staging branch.
on Apr 6, 2026 - added a commit that references this issue
on Apr 30, 2026 - added 2 commits that reference this issue
on May 11, 2026 github-actions commented
on Jul 31, 2026 on Jul 31, 2026 – with GitHub ActionsContributorMore actionsThis issue has been marked as stale due to 90 days of inactivity.
It will be automatically closed in 30 days if no further activity occurs. If this is still relevant, please leave a comment or update it to keep it open.- addedstaleIssues and PRs marked stale due to inactivity and scheduled for automatic closure.Issues and PRs marked stale due to inactivity and scheduled for automatic closure.
on Jul 31, 2026 This was fixed by #62621 in Node.js 22.22.3.
Version
Welcome to Node.js v22.4.1
Platform
Subsystem
No response
What steps will reproduce the bug?
`
const _TE = new TextEncoder();
var s = "aÿ我𝑒"
var u8a = new Uint8Array(2429682061);
var offset = 38928786 ;
_TE.encodeInto(s,u8a.subarray(offset)) // { read: 0, written: 0 } failed
_TE.encodeInto(s,u8a.subarray(offset,offset+10)) //{ read: 5, written: 10 } success
Welcome to Node.js v22.4.1.
Type ".help" for more information.
Welcome to Node.js v24.14.0.
Type ".help" for more information.
`
How often does it reproduce? Is there a required condition?
alaways in 22.4.1
What is the expected behavior? Why is that the expected behavior?
same as 24.14.0
What do you see instead?
const _TE = new TextEncoder();
var s = "aÿ我𝑒"
var u8a = new Uint8Array(2429682061);
var offset = 38928786 ;
_TE.encodeInto(s,u8a.subarray(offset)) // 【{ read: 0, written: 0 } failed】
_TE.encodeInto(s,u8a.subarray(offset,offset+10)) //{ read: 5, written: 10 } success
Additional information
No response