Twitter数据写入TXT文件遇UnicodeEncodeError问题求解
I'm working on a new project that fetches data from Twitter and writes it to a TXT file. I've run into an encoding error, and even converting the content to a string hasn't fixed it. Here's my code and the error message:
Code:
from TwitterSearch import * cd = open("Registro.txt", "a") try: tso = TwitterSearchOrder() tso.set_language('pt') tso.set_keywords(['rats']) ts = TwitterSearch('') # Removed personal info todo = True next_max_id = 0 while(todo): response = ts.search_tweets(tso) todo = not len(response['content']['statuses']) == 0 for tweet in response['content']['statuses']: tweet_id = tweet['id'] print(f"Seen tweet with ID {tweet_id}") print(tweet['text']) cd.write(tweet['text']) if (tweet_id < next_max_id) or (next_max_id == 0): next_max_id = tweet_id next_max_id -= 1 tso.set_max_id(next_max_id) except TwitterSearchException as e: print(e)
Error Message:
Traceback (most recent call last): File "", line 22, in <module> cd.write(tweet['text']) File "", line 19, in encode return codecs.charmap_encode(input,self.errors,encoding_table)[0] UnicodeEncodeError: 'charmap' codec can't encode characters in position 138-139: character maps to <undefined>
How can I fix this error?
Great question! The issue here is that you're opening your text file without specifying a Unicode-compatible encoding, and your system's default encoding (usually something like cp1252 on Windows) can't handle all the characters from Portuguese tweets—think special accents, emojis, or other Unicode symbols.
The Fix: Specify UTF-8 Encoding When Opening the File
UTF-8 supports all Unicode characters, so updating your open() call to include this encoding will resolve the error immediately. Here's the modified line:
cd = open("Registro.txt", "a", encoding='utf-8')
Bonus: Use a with Statement for Safer File Handling
It’s also a best practice to use Python’s with statement when working with files—it automatically closes the file for you, preventing resource leaks and unexpected behavior. Here’s how you can adjust your full code:
from TwitterSearch import * try: tso = TwitterSearchOrder() tso.set_language('pt') tso.set_keywords(['rats']) ts = TwitterSearch('') # Removed personal info todo = True next_max_id = 0 # Use with to manage the file lifecycle with open("Registro.txt", "a", encoding='utf-8') as cd: while(todo): response = ts.search_tweets(tso) todo = not len(response['content']['statuses']) == 0 for tweet in response['content']['statuses']: tweet_id = tweet['id'] print(f"Seen tweet with ID {tweet_id}") print(tweet['text']) cd.write(tweet['text'] + '\n') # Add newline to separate tweets if (tweet_id < next_max_id) or (next_max_id == 0): next_max_id = tweet_id next_max_id -= 1 tso.set_max_id(next_max_id) except TwitterSearchException as e: print(e)
I also added + '\n' to the write call so each tweet appears on its own line in the text file—makes it way easier to read later!
Edge Case: If You Still Encounter Issues
If for some reason you can’t use UTF-8 (though it’s strongly recommended), you can add an errors parameter to handle unencodable characters:
errors='replace': Replaces unencodable characters with?errors='ignore': Skips unencodable characters entirely
Example:
with open("Registro.txt", "a", encoding='utf-8', errors='replace') as cd:
But again, sticking with UTF-8 is the best solution because it preserves all original content from the tweets.
内容的提问来源于stack exchange,提问作者ArthurOgg

