You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取TXT文件编码报错,求二进制打开实现方法

Fixing UnicodeDecodeError When Reading TXT Files in Python

Hey there, let's work through this encoding headache you're facing—it's super common when dealing with text files from different sources!

Why You're Getting These Errors

The problem is that your .txt files aren't strictly encoded in either cp1252 or utf-8. Characters like 0x9d (from your first error) and 0x92 (from the second) are often smart quotes or special symbols from Windows-based encodings, but some files might have mixed encoding or rare characters that fall outside the codec's map.

Binary Mode Solution: How to Implement It

Reading files in binary mode ('rb') lets you bypass automatic decoding, so you can handle the encoding step manually. Here are a few practical ways to modify your readingAndAddToList() function:

Option 1: Try Common Encodings with Error Handling

First, try decoding with cp1252 (a common Windows encoding), and fall back to utf-8 with error replacement if that fails:

import glob

# Make sure word_list is defined (e.g., initialize it outside the function)
word_list = []

def readingAndAddToList():
    list_of_files = glob.glob('./*.txt')
    for file_name in list_of_files:
        # Open in binary mode to read raw bytes
        with open(file_name, 'rb') as fi:
            file_content = fi.read()
        
        # Attempt to decode with cp1252 first
        try:
            stext = file_content.decode('cp1252')
        except UnicodeDecodeError:
            # Fall back to utf-8, replacing unreadable characters with �
            stext = file_content.decode('utf-8', errors='replace')
        
        # Your existing processing logic
        stext_processed = textProcessing(stext)
        secondaryWord_list = stext_processed.split()
        word_list.extend(secondaryWord_list)
        print(f"Lungimea fisierului {fi.name} este de {len(secondaryWord_list)} caractere")
        sortingAndNumberOfApparitions(secondaryWord_list)

Option 2: Automatically Detect Encoding (More Accurate)

For better results, use the chardet library to detect the file's actual encoding. First install it:

pip install chardet

Then modify your function:

import glob
import chardet

word_list = []

def readingAndAddToList():
    list_of_files = glob.glob('./*.txt')
    for file_name in list_of_files:
        with open(file_name, 'rb') as fi:
            file_content = fi.read()
            # Detect the encoding with a confidence score
            detection_result = chardet.detect(file_content)
            detected_encoding = detection_result['encoding']
            confidence = detection_result['confidence']
            
            print(f"Detected encoding for {file_name}: {detected_encoding} (confidence: {confidence:.2f})")
            
            # Decode using the detected encoding, replacing unreadable characters
            stext = file_content.decode(detected_encoding, errors='replace')
        
        # Continue with your processing workflow
        stext_processed = textProcessing(stext)
        secondaryWord_list = stext_processed.split()
        word_list.extend(secondaryWord_list)
        print(f"Lungimea fisierului {fi.name} este de {len(secondaryWord_list)} caractere")
        sortingAndNumberOfApparitions(secondaryWord_list)

Key Notes

  • Using with open(...) is better than manually calling close()—it automatically closes the file even if an error occurs, preventing resource leaks.
  • The errors='replace' parameter replaces unreadable characters with �, so your script doesn't crash. You can also use errors='ignore' to skip those characters entirely, but replacement makes it easier to spot where encoding issues happened.
  • chardet isn't 100% perfect, but it's much more reliable than guessing encodings manually.

内容的提问来源于stack exchange,提问作者Adrian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:28:56