You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python清洗带变音符号阿拉伯语文本时遇UnicodeEncodeError报错

Fix UnicodeEncodeError When Saving Arabic Text with Diacritics in Python

Hey there, let's break down what's going wrong and how to fix it quickly:

The Root Cause

Your error happens because when you open the output file with open(localOutputPath, 'w'), Python uses the default system encoding (on Windows, that's usually cp1252). This old encoding doesn't support Arabic characters—especially not the diacritics (tashkeel) you want to keep. So when you try to write those characters, Python throws a UnicodeEncodeError because it can't map them to cp1252.

The Fixes

We need two key adjustments to make this work properly:

1. Force UTF-8 Encoding for Output Files

When opening your output file, explicitly specify encoding='utf-8'—this encoding supports every Unicode character, including all Arabic diacritics.

2. Confirm Your Regex is Correct

Your existing regex already targets the right Unicode ranges for Arabic text (including diacritics):

  • \u0600-\u06ff: Covers Arabic letters, diacritics, and punctuation
  • The other ranges (\u0750-\u077f, \ufb50-\ufbc1, etc.): Cover extended Arabic script variants and ligatures

This regex will keep exactly what you want (Arabic characters with diacritics) and replace everything else with spaces—perfect for your use case.

Modified Working Code

Here's the updated code with these fixes, plus a small improvement using os.path.join to avoid path-related bugs:

import os
import re

generalPath = "C:/Users/Desktop/Code/dataset/"
outputPath = "C:/Users/Desktop/Code/output/"

# Make sure output directory exists (in case it doesn't)
os.makedirs(outputPath, exist_ok=True)

files = os.listdir(generalPath)
for onefile in files:
    localPath = os.path.join(generalPath, onefile)
    localOutputPath = os.path.join(outputPath, onefile)
    
    print(f"Processing: {localPath}")
    print(f"Saving to: {localOutputPath}")
    
    with open(localPath, 'rb') as infile, open(localOutputPath, 'w', encoding='utf-8') as outfile:
        data = infile.read().decode('utf-8')
        # Your regex remains unchanged—it's correctly targeting Arabic text with diacritics
        new_data = re.sub(r'[^0-9\u0600-\u06ff\u0750-\u077f\ufb50-\ufbc1\ufbd3-\ufd3f\ufd50-\ufd8f\ufe70-\ufefc\uFDF0-\uFDFD]+', ' ', data)
        outfile.write(new_data)

Key Changes Explained

  • open(localOutputPath, 'w', encoding='utf-8'): Explicitly sets UTF-8 encoding for writing, which handles all Arabic characters.
  • os.makedirs(outputPath, exist_ok=True): Ensures the output folder exists (prevents errors if it wasn't created already).
  • os.path.join: Safely constructs file paths across different operating systems, avoiding issues with slashes.

This should resolve the encoding error while preserving all your Arabic diacritics as intended.

内容的提问来源于stack exchange,提问作者Moun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 21:12:31