disnake机器人编码错误:无法处理饥荒联机版表情致on_ready中断
问题背景
我开发了一个基于disnake的Discord机器人,负责同步饥荒联机版(Don't Starve Together)的服务器日志与聊天内容。当聊天日志中出现饥荒特殊表情时,会触发UnicodeEncodeError,导致on_ready函数停止运行,但机器人仍保持在线状态。需要忽略该特殊字符,让on_ready函数恢复正常工作。
相关代码
import disnake from disnake.ext import commands import re import asyncio from disnake import Intents bot = commands.Bot(command_prefix="*", intents=disnake.Intents.all()) player_ku = {} isConsole = False isChat = True isId = True @bot.event async def on_ready(): print(f'{bot.user.name} has connected to Discord!') channel1 = bot.get_channel(1105450189688414208) channel2 = bot.get_channel(1105450335947989042) channel3 = bot.get_channel(1105694581800042566) last_line_file1 = '' last_line_file2 = '' while True: console_disabled_message_sended = False if isConsole: with open("server_log.txt", 'r', encoding='cp1251', errors='ignore') as file1: # Пусть к выводу консольи last_line = file1.readlines()[-1].strip() last_line_utf82 = last_line.encode('cp1251', errors='backslashreplace').decode("utf8",errors='ignore') if last_line_utf82 != last_line_file1: await channel1.send(f'```\n{last_line}\n```') last_line_file1 = last_line_utf82 asyncio.sleep(1) if isChat: with open('server_chat_log.txt', 'r', encoding="cp1251", errors='ignore') as file2: # Пусть к выводу чата last_line = file2.readlines()[-1].strip() last_line_utf8 = last_line.encode('cp1251', errors='backslashreplace').decode("utf8",errors='ignore') if last_line_utf8 != last_line_file2: cleaned_message = last_line_utf8.split("]: ")[1].split() clean_message = " ".join(cleaned_message) mega_cleaned_message = re.sub('[:(\)\\[\\]]', '', clean_message).split() ku = mega_cleaned_message[1] nick = mega_cleaned_message[2] if cleaned_message[0] == "[Say]": await channel2.send(f"```\n💬 {clean_message}```") last_line_file2 = last_line_utf8 if not ku in player_ku: player_ku[ku] = nick print(player_ku) with open("savedKu.txt", "a") as file: for key in player_ku: encoded_key = key.encode('cp1251', errors='backslashreplace').decode('utf-8', errors='ignore') file.write(f"\n{encoded_key}: {player_ku[key]}\n") if isId: await channel3.send(f"```\n🆔{encoded_key} = {player_ku[key]}```") elif cleaned_message[0] == "[Whisper]": await channel2.send(f"```\n🤫 {clean_message}```") last_line_file2 = last_line_utf8 if not ku in player_ku: print(player_ku) elif cleaned_message[0] == "[Leave": await channel2.send(f"```\n⬅️ {clean_message}```") last_line_file2 = last_line_utf8 elif cleaned_message[0] == "[Join": await channel2.send(f"```\n➡️ {clean_message}```") last_line_file2 = last_line_utf8 elif cleaned_message[0] == "[Announcement]": await channel2.send(f"```\n📢 {clean_message}```") last_line_file2 = last_line_utf8 elif cleaned_message[0] == "[Death": await channel2.send(f"```\n☠️ {clean_message}```") last_line_file2 = last_line_utf8 else: await channel2.send(f"```\n⬛️ {clean_message}```") last_line_file2 = last_line_utf8 bot.run('---')
报错信息
Ignoring exception in on_ready Traceback (most recent
call last): File "C:\Users\N.Sarychev\Desktop\WTFbotControl\bot.py",
line 66, in on_ready file.write(f"\n{encoded_key}:
{player_ku[key]}\n") File "C:\Program
Files\WindowsApps\PythonSoftwareFoundation.Python.3.10_3.10.3056.0_x64__qbz5n2kfra8p0\lib\encodings\cp1251.py",
line 19, in encode return
codecs.charmap_encode(input,self.errors,encoding_table)[0]
UnicodeEncodeError: 'charmap' codec can't encode character
'\U000f000f' in position 15: character maps to
特殊字符示例:(饥荒联机版内置表情)
触发报错的字符串示例:[02:41:27]: [Say] (KU_nAt72Ax2) 󰀏трэш󰀏:
问题根源
报错出现在写入savedKu.txt文件时:默认打开文件未指定编码,系统使用了cp1251编码,但饥荒的特殊表情不在cp1251的字符范围内,导致编码失败,进而中断on_ready函数的循环。
解决方案
1. 修改文件写入的编码设置
将写入savedKu.txt的代码改为UTF-8编码,并添加errors='ignore'忽略无法编码的字符:
修改前:
with open("savedKu.txt", "a") as file: for key in player_ku: encoded_key = key.encode('cp1251', errors='backslashreplace').decode('utf-8', errors='ignore') file.write(f"\n{encoded_key}: {player_ku[key]}\n")
修改后:
# 指定UTF-8编码打开文件,忽略无法编码的特殊字符 with open("savedKu.txt", "a", encoding='utf-8', errors='ignore') as file: for key in player_ku: # 直接写入处理后的字符串,无需多余的编码转换 file.write(f"\n{key}: {player_ku[key]}\n")
2. 优化编码转换逻辑
读取日志时已经将内容转换为UTF-8字符串,后续无需再用cp1251编码转换,直接使用处理后的字符串即可避免编码混乱。
额外建议
- 所有涉及文件读写的操作统一使用UTF-8编码,避免因编码不一致引发的问题。
- 若需要保留特殊字符的占位符,可将
errors='ignore'替换为errors='backslashreplace',将无法编码的字符转为转义序列保存。
内容的提问来源于stack exchange,提问作者Fanky

