在子线程中创建BeautifulSoup对象触发编码错误问题求助
子线程中创建BeautifulSoup触发编码错误的解决方案
我之前也踩过这个坑!这其实是lxml解析器在多线程环境下的编码处理bug搞的鬼。先看看你的问题重现代码:
import requests from bs4 import BeautifulSoup from threading import Thread def test(): r = requests.get('http://zhuanlan.sina.com.cn/') soup = BeautifulSoup(r.content,'lxml') print('run test on main thread') test() print('run test on child thread') t = Thread(target=test) t.start() t.join()
运行后会在子线程执行时抛出编码错误:
encoding error : input conversion failed due to input error, bytes 0x95 0x50 0x22 0x20
encoding error : input conversion...
问题原因
lxml解析器在初始化时会依赖系统的编码上下文,而子线程并没有完全继承主线程的编码配置,或者说在多线程场景下lxml的编码检测逻辑存在竞态问题,导致直接解析二进制的r.content时无法正确识别编码,从而触发转换失败。
解决方案
这里有两个简单有效的修复方式:
方式一:用requests解析好的文本代替二进制内容
requests会自动识别响应的编码并把二进制内容转换成字符串,直接用r.text传给BeautifulSoup就可以绕过lxml的编码检测问题:
def test(): r = requests.get('http://zhuanlan.sina.com.cn/') # 用r.text代替r.content soup = BeautifulSoup(r.text, 'lxml')
方式二:换用Python内置的html.parser解析器
如果不想修改传入的内容,也可以换成Python自带的html.parser,它的多线程编码处理更稳定:
def test(): r = requests.get('http://zhuanlan.sina.com.cn/') soup = BeautifulSoup(r.content, 'html.parser')
两种方式都能解决子线程中的编码错误问题,你可以根据自己的需求选择。
内容的提问来源于stack exchange,提问作者Gao Liang
相关产品推荐
相关产品推荐

