如何在Emacs中按Unicode编号范围搜索并替换字符?
If you need to target characters in the Unicode range U+1D000 to U+1DFFF (which includes mathematical alphanumeric symbols and other supplementary plane characters) and wrap each match in double quotes, here's a step-by-step guide using Emacs' regex tools:
Step 1: Launch the Regex Replace Tool
Open the buffer containing your text, then trigger the query-replace-regexp command—this lets you review each change before applying it (safer than a blind replace):
- Use the keyboard shortcut:
C-M-%(hold Ctrl + Alt, then press %) - Or type
M-x query-replace-regexp(press Alt+x, then enter the command name and hit Enter)
Step 2: Enter the Unicode Range Regex Pattern
When prompted for the "Regexp to replace:", input this pattern:
[\x{1d000}-\x{1dfff}]
\x{XXXXXX}is Emacs' syntax for specifying a Unicode code point (here, 5-digit values for the supplementary plane range).- The square brackets create a character class that matches any single character between U+1D000 and U+1DFFF inclusive.
Step 3: Define the Replacement String
Next, when asked for the "Replacement string:", enter:
"\\&"
\\&is Emacs' way of referencing the exact text that was matched by the regex.- Wrapping it in double quotes (
") ensures each matched character gets enclosed in".
Step 4: Navigate and Apply Changes
Once you've entered both values, you can control the replacement process:
- Press
yto replace the current match and move to the next one. - Press
nto skip the current match without changing it. - Press
!to replace all remaining matches automatically (no more prompts). - Press
Ctrl+gto cancel the operation at any time.
Optional: Blind Replace All
If you're confident you want to replace every occurrence without reviewing, use M-x replace-regexp instead. Follow the same pattern and replacement string steps—this will apply all changes immediately.
Note on Encoding
Ensure your buffer is using UTF-8 encoding (Emacs defaults to this for most cases). If not, run M-x set-buffer-file-coding-system and select utf-8 to properly handle Unicode characters.
内容的提问来源于stack exchange,提问作者Coeus Wang

