Python 3.6网页爬虫UnicodeEncodeError问题求助
解决Python爬虫的UnicodeEncodeError问题
嘿,我来帮你搞定这个烦人的编码错误!咱们先理清问题根源,再一步步修正你的脚本。
问题本质
你遇到的UnicodeEncodeError: 'ascii' codec can't encode character '\u2022',核心原因是写入文件时用了默认的ASCII编码——网页里的•(对应Unicode字符\u2022)超出了ASCII的字符范围,自然会编码失败。
初始脚本的问题
你的初始脚本里,打开文件时没指定编码:
f = open(filename, "w")
在macOS环境下,Python默认会用ASCII编码写入文件,这直接导致特殊字符无法处理。另外你尝试的encode("utf-8)")和decode("utf-8)")不仅多余,还存在语法错误(引号没闭合),反而会引发新问题。
修改后脚本的进步与小疏漏
你修改后的脚本已经做了最关键的一步:打开文件时明确指定utf-8编码:
f = open(filename, "w",encoding="utf-8")
这是正确的方向!但里面有两个小问题需要修正:
- 变量名笔误:
un = str(uno_en.text)里的uno_en是不存在的变量,你前面定义的是uno = (one[0].h1),应该改成un = uno.text(甚至不需要str(),因为.text本身就是字符串类型)。 findAll参数格式错误:原脚本里的findAll("div",{"body field"})写法不对,必须明确指定属性键,比如{"class":"body field"},否则BeautifulSoup无法正确匹配元素。
修正后的完整可运行脚本
import bs4 from urllib.request import urlopen as uReq from bs4 import BeautifulSoup as soup my_url = 'https://www.insertsitehere.com/' uClient = uReq(my_url) page_html = uClient.read() uClient.close() page_soup = soup(page_html, "html.parser") # 修正findAll的参数格式,明确指定class属性 one = page_soup.findAll("div", {"class":"large-7 medium-9 columns"}) two = page_soup.findAll("div", {"class":"body field"}) three = page_soup.findAll("div", {"class":"field-ingredients-data field-wrapper"}) four = page_soup.findAll("div", {"class":"field-how-to-use-data field-wrapper"}) uno = one[0].h1 dos = two[1] tres = three[0].p cuatro = four[0].p filename = "example.csv" # 使用with语句自动管理文件,同时指定utf-8编码 with open(filename, "w", encoding="utf-8") as f: headers = "one, two, three, four\n" f.write(headers) # 直接获取文本字符串,无需额外转换 un = uno.text deux = dos.text trois = tres.text quatre = cuatro.text # 用f-string拼接更清晰,避免手动拼接的混乱 f.write(f"{un},{deux},{trois},{quatre}\n")
额外优化提示
- 用with语句管理文件:比手动调用
f.close()更安全,能自动处理文件关闭逻辑,避免遗漏。 - f-string格式化:比字符串拼接更易读,也能减少编码相关的隐性问题。
- 异常处理:可以给
one[0]、two[1]这类索引操作加上try-except,避免网页结构变化导致脚本崩溃。
这样修改后,你的脚本就能正常处理包含•这类特殊字符的内容,不会再出现UnicodeEncodeError啦!
内容的提问来源于stack exchange,提问作者marcoantonio123456
相关产品推荐
相关产品推荐

