Notepad++中JSON文件Unicode转义序列解码部分字符失败求助
Hey there, let's break down exactly why your long dashes, Ç, and other characters aren't fixing, and walk through step-by-step solutions to get your JSON clean and ready for your program.
Why Some Characters Are Still Broken
First off, what you're dealing with isn't just "Unicode escape sequences"—it's mojibake (double encoding). Here's the simple breakdown:
- Characters like
éare stored as UTF-8 bytes (C3 A9), but if someone (or a tool) incorrectly decoded those bytes using ISO-8859-1 (ANSI), you geté. - A long dash (
—) has UTF-8 bytesE2 80 94; decoding that as ANSI givesâ. ÇbecomesÃfor the same reason: UTF-8 bytesC3 87→ ANSI-decoded toÃ.
Your earlier steps (Decode JS → ANSI → UTF-8) fixed some cases, but missed others because:
- The "Decode JS" feature targets
\uXXXXstyle escapes, not this byte-level encoding mess. - Some characters might have gone through multiple rounds of accidental encoding, making the single ANSI→UTF-8 flip insufficient.
Step-by-Step Fixes for Notepad++
Fix 1: Use Built-in Encoding Conversions (No Plugins Needed)
This is the fastest fix for most double-encoding cases:
- Open your messed-up JSON file in Notepad++.
- Go to the top menu: Encoding → Convert to ANSI. This takes the current "garbage" characters and saves them as the original UTF-8 bytes (since ANSI maps 1:1 to single-byte values).
- Immediately go back to Encoding → Convert to UTF-8. Now Notepad++ will decode those bytes correctly using UTF-8, and your
é,—, andÇshould pop into place.- If this doesn't work on the first try, close the file without saving, reopen it with Encoding → ANSI selected in the open dialog, then repeat steps 2-3.
Fix 2: Use Python Script for Stubborn Cases
If some characters still refuse to cooperate (maybe they were encoded three times?), use Notepad++'s Python Script plugin to automate the fix:
- Install the Python Script plugin via Plugins → Plugins Admin, search for it, and install.
- Go to Plugins → Python Script → New Script, name it
FixMojibake.py, and paste this code:
# Get the entire file content raw_text = editor.getText() # Reverse the double encoding: treat text as ISO-8859-1 bytes, decode to UTF-8 fixed_text = raw_text.encode('iso-8859-1').decode('utf-8') # Replace the file content with the fixed version editor.setText(fixed_text)
- Run the script via Plugins → Python Script → Scripts → FixMojibake. This will fix every case of UTF-8→ANSI mojibake in one go.
Fix 3: Target Specific Characters (Quick Manual Patch)
If you only have a few stubborn characters left, you can do a direct find-and-replace:
- Open the Replace dialog (Ctrl+H).
- In the "Find what" field, paste the messed-up character (e.g.,
âorÃ). - In the "Replace with" field, paste the correct character (e.g.,
—orÇ). - Click "Replace All" to clean up those last stragglers.
Final Check
After fixing, save the file as UTF-8 (Encoding → UTF-8) and verify that your program reads it correctly. This should eliminate all the encoding-related garbage from your JSON.
内容的提问来源于stack exchange,提问作者Don

