将字典列表保存为CSV时遇UnicodeEncodeError,如何解决?
解决CSV导出时的UnicodeEncodeError错误
爬取并解析数据后,尝试用以下代码将字典列表保存为CSV文件时,出现了Unicode编码错误:
# Example run: if __name__ == "__main__": with httpx.Client(timeout=httpx.Timeout(20.0)) as session: posts = list(scrape_user_posts("1142370320", session, max_posts=25, page_size=3)) print(json.dumps(posts, indent=2, ensure_ascii=False)) keys = posts[0].keys() with open('Lenskart_posts.csv', 'w', newline='') as output_file: dict_writer = csv.DictWriter(output_file, keys) dict_writer.writeheader() dict_writer.writerows(posts)
运行后报错信息:
Traceback (most recent call last): File "C:\Users\arjun\AppData\Local\Programs\Python\Python311\Scrapping_test.py", line 108, in <module> dict_writer.writerows(posts) File "C:\Users\arjun\AppData\Local\Programs\Python\Python311\Lib\csv.py", line 157, in writerows return self.writer.writerows(map(self._dict_to_list, rowdicts)) File "C:\Users\arjun\AppData\Local\Programs\Python\Python311\Lib\encodings\cp1252.py", line 19, in encode return codecs.charmap_encode(input,self.errors,encoding_table)[0] UnicodeEncodeError: 'charmap' codec can't encode character '\U0001f31f' in position 878: character maps to <undefined>
修复方案
错误根源是Windows系统下open()函数默认使用cp1252编码,而爬取的数据包含该编码不支持的Unicode字符(比如示例中的emoji符号\U0001f31f)。只需在打开文件时显式指定utf-8编码即可解决:
修改后的代码:
# Example run: if __name__ == "__main__": with httpx.Client(timeout=httpx.Timeout(20.0)) as session: posts = list(scrape_user_posts("1142370320", session, max_posts=25, page_size=3)) print(json.dumps(posts, indent=2, ensure_ascii=False)) keys = posts[0].keys() # 显式指定utf-8编码 with open('Lenskart_posts.csv', 'w', newline='', encoding='utf-8') as output_file: dict_writer = csv.DictWriter(output_file, keys) dict_writer.writeheader() dict_writer.writerows(posts)
补充说明
utf-8编码兼容几乎所有Unicode字符,能完美处理emoji、特殊符号等内容;- 保留
newline=''可避免CSV文件在不同系统中出现换行符不一致的问题; - 如果后续用Excel打开CSV出现乱码,可尝试将编码改为
utf-8-sig,它会添加BOM头,让Excel正确识别UTF-8编码。
内容的提问来源于stack exchange,提问作者ARJUN RAMASWAMY
相关产品推荐
相关产品推荐

