含Unicode字符的PHP编码字符串无法在JavaScript中解码的问题
Fixing Unicode Support for Your Rot-like Encoding/Decoding & Replacement Workflow
Got it, let's tackle why your Unicode characters are breaking in this PHP-to-JS rotation encoding and replacement setup. The core issues boil down to how both languages handle multi-byte/extended Unicode characters, plus some logic gaps in your rotation math and DOM replacement. Let's fix this step by step.
What's Going Wrong?
Let's break down the key problems:
- PHP is splitting Unicode characters into bytes: When you use
$str[$i], PHP treats strings as byte arrays. For characters likeÜ, emojis, or Chinese text (which take multiple bytes), you're processing each byte individually instead of the full character—this completely corrupts the encoding. - Your Unicode range is way too big: Setting
max = 99999999goes far beyond the actual maximum valid Unicode code point (0x10FFFF, or 1114111). Any code point beyond that becomes invalid, making decoding impossible. - JS isn't handling surrogate pairs: Characters like the dog emoji
live in Unicode's supplementary plane, which JS represents as two "surrogate" UTF-16 units. UsingcharCodeAt(i)only gets half the pair, andfromCharCodecan't rebuild the full character. - DOM replacement is risky: Using
html().replace()can break your HTML structure if your text contains characters like<or>, and you're not doing global replacements (so only the first match gets swapped).
Fixed Full Code Implementation
Here's the revised code with all fixes applied:
Updated PHP strRot Function
<?php function strRot($str, $n) { $len = mb_strlen($str); $min = 0; $max = 0x10FFFF; // Maximum valid Unicode code point $final = ''; for ($i = 0; $i < $len; $i++) { // Grab the full Unicode character (not just a byte) $char = mb_substr($str, $i, 1); $current = mb_ord($char); // Calculate rotated code point with proper wrapping $val = ($current + $n) % $max; // Handle negative values if n is negative if ($val < $min) { $val += $max; } $final .= mb_chr($val); } return $final; } ?>
Updated Full HTML/JavaScript Code
<!doctype html> <html lang="en"> <head> <meta charset="utf-8"> <meta name="viewport" content="width=device-width, initial-scale=1"> <link rel="stylesheet" href="https://stackpath.bootstrapcdn.com/bootstrap/4.2.1/css/bootstrap.min.css" integrity="sha384-GJzZqFGwb1QTTN6wy59ffF1BuGJpLSa9DkKMp0DgiMDm4iYMj70gZWKYbI706tWS" crossorigin="anonymous"> <title>Hello, world!</title> </head> <body> <h1>Hello, world!</h1> <h2>Ü and . 棕色的狐狸跳了起来.</h2> <p>The Hello, world! expression will be replaced.</p> <p>Ü and . 棕色的狐狸跳了起来. Should be replaced too.</p> <script src="https://code.jquery.com/jquery-3.3.1.slim.min.js" integrity="sha384-q8i/X+965DzO0rT7abK41JStQIAqVgRVzpbzo5smXKp4YfRvH+8abtTE1Pi6jizo" crossorigin="anonymous"></script> <script src="https://cdnjs.cloudflare.com/ajax/libs/popper.js/1.14.6/umd/popper.min.js" integrity="sha384-wHAiFfRlMFy6i5SRaxvfOCifBUQy1xHdJ/yoi7FRNXMRBu5WHdZYu1hA6ZOblgut" crossorigin="anonymous"></script> <script src="https://stackpath.bootstrapcdn.com/bootstrap/4.2.1/js/bootstrap.min.js" integrity="sha384-B0UglyR+jN6CkvvICOB2joaf5I4l3gm9GU6Hc1og6Ls7i6U/mkkaduKaBhlAXv9k" crossorigin="anonymous"></script> <script id="scriptId" type="text/javascript"> var data = [ ["Hello, world!", "<?php echo base64_encode(strRot('I got replaced.', 1000)); ?>"], ["Ü and . 棕色的狐狸跳了起来.", "<?php echo base64_encode(strRot(' before Ü and 棕色的.', 1000)); ?>"] ]; function b64DecodeUnicode(str) { return decodeURIComponent(atob(str).split('').map(function(c) { return '%' + ('00' + c.charCodeAt(0).toString(16)).slice(-2); }).join('')); } function strRot(str, n) { const min = 0; const max = 0x10FFFF; let final = ''; // Split string into full Unicode characters (handles emojis/supplementary plane) const chars = Array.from(str); for (const char of chars) { const current = char.codePointAt(0); let val = (current - n) % max; // Adjust negative values to stay within valid Unicode range if (val < min) { val += max; } final += String.fromCodePoint(val); } return final; } function replace() { data.forEach(([target, encodedReplacement]) => { const replacement = strRot(b64DecodeUnicode(encodedReplacement), -1000); // Escape regex special characters in the target string const escapedTarget = target.replace(/[.*+?^${}()|[\]\\]/g, '\\$&'); const regex = new RegExp(escapedTarget, 'g'); // Replace only text nodes to avoid breaking HTML jQuery("html *:not(script[id=scriptId])").contents().filter(function() { return this.nodeType === Node.TEXT_NODE; }).each(function() { this.textContent = this.textContent.replace(regex, replacement); }); }); } replace(); </script> </body> </html>
What Changed & Why?
PHP Unicode Handling:
- Used
mb_substrto get full Unicode characters instead of$str[$i] - Set
$maxto0x10FFFF(the actual maximum Unicode code point) - Swapped arbitrary subtraction/addition for modulo arithmetic to properly wrap code points
- Used
JS Unicode Handling:
- Used
Array.from(str)to split the string into full Unicode characters (fixes surrogate pairs) - Used
codePointAt(0)to get the full Unicode code point, andString.fromCodePointto reconstruct characters (works for emojis/Chinese) - Matched PHP's modulo logic for consistent encoding/decoding
- Used
DOM Replacement:
- Targeted only text nodes (instead of modifying HTML directly) to avoid breaking your page structure
- Escaped regex special characters in the target string to prevent regex syntax errors
- Used the
gflag in regex to replace all occurrences, not just the first
This setup will now correctly encode/decode all Unicode characters (including emojis, accented letters, and non-Latin scripts) and safely replace text in your DOM.
内容的提问来源于stack exchange,提问作者Augustin
相关产品推荐
相关产品推荐

