Python读取TXT文件编码报错,求二进制打开实现方法
Hey there, let's work through this encoding headache you're facing—it's super common when dealing with text files from different sources!
Why You're Getting These Errors
The problem is that your .txt files aren't strictly encoded in either cp1252 or utf-8. Characters like 0x9d (from your first error) and 0x92 (from the second) are often smart quotes or special symbols from Windows-based encodings, but some files might have mixed encoding or rare characters that fall outside the codec's map.
Binary Mode Solution: How to Implement It
Reading files in binary mode ('rb') lets you bypass automatic decoding, so you can handle the encoding step manually. Here are a few practical ways to modify your readingAndAddToList() function:
Option 1: Try Common Encodings with Error Handling
First, try decoding with cp1252 (a common Windows encoding), and fall back to utf-8 with error replacement if that fails:
import glob # Make sure word_list is defined (e.g., initialize it outside the function) word_list = [] def readingAndAddToList(): list_of_files = glob.glob('./*.txt') for file_name in list_of_files: # Open in binary mode to read raw bytes with open(file_name, 'rb') as fi: file_content = fi.read() # Attempt to decode with cp1252 first try: stext = file_content.decode('cp1252') except UnicodeDecodeError: # Fall back to utf-8, replacing unreadable characters with � stext = file_content.decode('utf-8', errors='replace') # Your existing processing logic stext_processed = textProcessing(stext) secondaryWord_list = stext_processed.split() word_list.extend(secondaryWord_list) print(f"Lungimea fisierului {fi.name} este de {len(secondaryWord_list)} caractere") sortingAndNumberOfApparitions(secondaryWord_list)
Option 2: Automatically Detect Encoding (More Accurate)
For better results, use the chardet library to detect the file's actual encoding. First install it:
pip install chardet
Then modify your function:
import glob import chardet word_list = [] def readingAndAddToList(): list_of_files = glob.glob('./*.txt') for file_name in list_of_files: with open(file_name, 'rb') as fi: file_content = fi.read() # Detect the encoding with a confidence score detection_result = chardet.detect(file_content) detected_encoding = detection_result['encoding'] confidence = detection_result['confidence'] print(f"Detected encoding for {file_name}: {detected_encoding} (confidence: {confidence:.2f})") # Decode using the detected encoding, replacing unreadable characters stext = file_content.decode(detected_encoding, errors='replace') # Continue with your processing workflow stext_processed = textProcessing(stext) secondaryWord_list = stext_processed.split() word_list.extend(secondaryWord_list) print(f"Lungimea fisierului {fi.name} este de {len(secondaryWord_list)} caractere") sortingAndNumberOfApparitions(secondaryWord_list)
Key Notes
- Using
with open(...)is better than manually callingclose()—it automatically closes the file even if an error occurs, preventing resource leaks. - The
errors='replace'parameter replaces unreadable characters with�, so your script doesn't crash. You can also useerrors='ignore'to skip those characters entirely, but replacement makes it easier to spot where encoding issues happened. chardetisn't 100% perfect, but it's much more reliable than guessing encodings manually.
内容的提问来源于stack exchange,提问作者Adrian

