You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在子线程中创建BeautifulSoup对象触发编码错误问题求助

子线程中创建BeautifulSoup触发编码错误的解决方案

我之前也踩过这个坑!这其实是lxml解析器在多线程环境下的编码处理bug搞的鬼。先看看你的问题重现代码:

import requests
from bs4 import BeautifulSoup
from threading import Thread

def test():
    r = requests.get('http://zhuanlan.sina.com.cn/')
    soup = BeautifulSoup(r.content,'lxml')

print('run test on main thread')
test()
print('run test on child thread')
t = Thread(target=test)
t.start()
t.join()

运行后会在子线程执行时抛出编码错误:

encoding error : input conversion failed due to input error, bytes 0x95 0x50 0x22 0x20
encoding error : input conversion...

问题原因

lxml解析器在初始化时会依赖系统的编码上下文,而子线程并没有完全继承主线程的编码配置,或者说在多线程场景下lxml的编码检测逻辑存在竞态问题,导致直接解析二进制的r.content时无法正确识别编码,从而触发转换失败。

解决方案

这里有两个简单有效的修复方式:

方式一:用requests解析好的文本代替二进制内容

requests会自动识别响应的编码并把二进制内容转换成字符串,直接用r.text传给BeautifulSoup就可以绕过lxml的编码检测问题:

def test():
    r = requests.get('http://zhuanlan.sina.com.cn/')
    # 用r.text代替r.content
    soup = BeautifulSoup(r.text, 'lxml')

方式二:换用Python内置的html.parser解析器

如果不想修改传入的内容,也可以换成Python自带的html.parser,它的多线程编码处理更稳定:

def test():
    r = requests.get('http://zhuanlan.sina.com.cn/')
    soup = BeautifulSoup(r.content, 'html.parser')

两种方式都能解决子线程中的编码错误问题,你可以根据自己的需求选择。

内容的提问来源于stack exchange,提问作者Gao Liang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:39:32