ext/mbstring: Optimize valid UTF-8 in mb_substr() - #24080
kamil-tekiela wants to merge 1 commit into
Conversation
|
@kamil-tekiela Fuzzing is 100% necessary. Otherwise, it is very, very easy for subtle bugs to slip through. It's good that you optimized |
|
Do you have a fuzzer available? If not can you give me hints on how to build one? |
|
@kamil-tekiela When I have some time, I could fuzz this for you. Or, if you want to do it yourself, that might be a very good idea... fuzzing is an extremely powerful technique for finding bugs, and is an awesome tool for any programmer to know how to use. I once wrote up my workflow for fuzzing new code in mbstring here: #10828 (comment) |
|
I tried fuzzing as you said and it didn't find anything. Not sure I've done it properly, but when I intentionally introduced a bug it found it immediately. |
This PR adds a fast path for valid UTF-8 strings in
mb_substr. It skips bytes in 256 chunks. The performance improvement is noticeable with longer strings and only ones that are flagged as valid UTF-8.@alexdowad What do you think? Is this worth doing? I haven't got a fuzzer, but I don't anticipate any behavioural changes.