Python清洗带变音符号阿拉伯语文本时遇UnicodeEncodeError报错
Hey there, let's break down what's going wrong and how to fix it quickly:
The Root Cause
Your error happens because when you open the output file with open(localOutputPath, 'w'), Python uses the default system encoding (on Windows, that's usually cp1252). This old encoding doesn't support Arabic characters—especially not the diacritics (tashkeel) you want to keep. So when you try to write those characters, Python throws a UnicodeEncodeError because it can't map them to cp1252.
The Fixes
We need two key adjustments to make this work properly:
1. Force UTF-8 Encoding for Output Files
When opening your output file, explicitly specify encoding='utf-8'—this encoding supports every Unicode character, including all Arabic diacritics.
2. Confirm Your Regex is Correct
Your existing regex already targets the right Unicode ranges for Arabic text (including diacritics):
\u0600-\u06ff: Covers Arabic letters, diacritics, and punctuation- The other ranges (
\u0750-\u077f,\ufb50-\ufbc1, etc.): Cover extended Arabic script variants and ligatures
This regex will keep exactly what you want (Arabic characters with diacritics) and replace everything else with spaces—perfect for your use case.
Modified Working Code
Here's the updated code with these fixes, plus a small improvement using os.path.join to avoid path-related bugs:
import os import re generalPath = "C:/Users/Desktop/Code/dataset/" outputPath = "C:/Users/Desktop/Code/output/" # Make sure output directory exists (in case it doesn't) os.makedirs(outputPath, exist_ok=True) files = os.listdir(generalPath) for onefile in files: localPath = os.path.join(generalPath, onefile) localOutputPath = os.path.join(outputPath, onefile) print(f"Processing: {localPath}") print(f"Saving to: {localOutputPath}") with open(localPath, 'rb') as infile, open(localOutputPath, 'w', encoding='utf-8') as outfile: data = infile.read().decode('utf-8') # Your regex remains unchanged—it's correctly targeting Arabic text with diacritics new_data = re.sub(r'[^0-9\u0600-\u06ff\u0750-\u077f\ufb50-\ufbc1\ufbd3-\ufd3f\ufd50-\ufd8f\ufe70-\ufefc\uFDF0-\uFDFD]+', ' ', data) outfile.write(new_data)
Key Changes Explained
open(localOutputPath, 'w', encoding='utf-8'): Explicitly sets UTF-8 encoding for writing, which handles all Arabic characters.os.makedirs(outputPath, exist_ok=True): Ensures the output folder exists (prevents errors if it wasn't created already).os.path.join: Safely constructs file paths across different operating systems, avoiding issues with slashes.
This should resolve the encoding error while preserving all your Arabic diacritics as intended.
内容的提问来源于stack exchange,提问作者Moun

