You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Twitter数据写入TXT文件遇UnicodeEncodeError问题求解

UnicodeEncodeError when writing Twitter tweets to a TXT file

I'm working on a new project that fetches data from Twitter and writes it to a TXT file. I've run into an encoding error, and even converting the content to a string hasn't fixed it. Here's my code and the error message:

Code:

from TwitterSearch import *
cd = open("Registro.txt", "a")
try:
    tso = TwitterSearchOrder()
    tso.set_language('pt')
    tso.set_keywords(['rats'])
    ts = TwitterSearch('') # Removed personal info
    todo = True
    next_max_id = 0
    while(todo):
        response = ts.search_tweets(tso)
        todo = not len(response['content']['statuses']) == 0
        for tweet in response['content']['statuses']:
            tweet_id = tweet['id']
            print(f"Seen tweet with ID {tweet_id}")
            print(tweet['text'])
            cd.write(tweet['text'])
            if (tweet_id < next_max_id) or (next_max_id == 0):
                next_max_id = tweet_id
        next_max_id -= 1
        tso.set_max_id(next_max_id)
except TwitterSearchException as e:
    print(e)

Error Message:

Traceback (most recent call last):
  File "", line 22, in <module>
    cd.write(tweet['text'])
  File "", line 19, in encode
    return codecs.charmap_encode(input,self.errors,encoding_table)[0]
UnicodeEncodeError: 'charmap' codec can't encode characters in position 138-139: character maps to <undefined>

How can I fix this error?


Great question! The issue here is that you're opening your text file without specifying a Unicode-compatible encoding, and your system's default encoding (usually something like cp1252 on Windows) can't handle all the characters from Portuguese tweets—think special accents, emojis, or other Unicode symbols.

The Fix: Specify UTF-8 Encoding When Opening the File

UTF-8 supports all Unicode characters, so updating your open() call to include this encoding will resolve the error immediately. Here's the modified line:

cd = open("Registro.txt", "a", encoding='utf-8')

Bonus: Use a with Statement for Safer File Handling

It’s also a best practice to use Python’s with statement when working with files—it automatically closes the file for you, preventing resource leaks and unexpected behavior. Here’s how you can adjust your full code:

from TwitterSearch import *

try:
    tso = TwitterSearchOrder()
    tso.set_language('pt')
    tso.set_keywords(['rats'])
    ts = TwitterSearch('') # Removed personal info
    todo = True
    next_max_id = 0
    # Use with to manage the file lifecycle
    with open("Registro.txt", "a", encoding='utf-8') as cd:
        while(todo):
            response = ts.search_tweets(tso)
            todo = not len(response['content']['statuses']) == 0
            for tweet in response['content']['statuses']:
                tweet_id = tweet['id']
                print(f"Seen tweet with ID {tweet_id}")
                print(tweet['text'])
                cd.write(tweet['text'] + '\n') # Add newline to separate tweets
                if (tweet_id < next_max_id) or (next_max_id == 0):
                    next_max_id = tweet_id
            next_max_id -= 1
            tso.set_max_id(next_max_id)
except TwitterSearchException as e:
    print(e)

I also added + '\n' to the write call so each tweet appears on its own line in the text file—makes it way easier to read later!

Edge Case: If You Still Encounter Issues

If for some reason you can’t use UTF-8 (though it’s strongly recommended), you can add an errors parameter to handle unencodable characters:

  • errors='replace': Replaces unencodable characters with ?
  • errors='ignore': Skips unencodable characters entirely

Example:

with open("Registro.txt", "a", encoding='utf-8', errors='replace') as cd:

But again, sticking with UTF-8 is the best solution because it preserves all original content from the tweets.


内容的提问来源于stack exchange,提问作者ArthurOgg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 12:42:53