Python3使用boilerpipe爬取网页遇Unicode相关错误求助
解决Python3中使用boilerpipe爬取网页时的Unicode错误问题
嘿,我来帮你拆解下你遇到的这两个问题,其实根源都是boilerpipe这个库对Python3的兼容性不太好,加上网页内容的压缩处理导致的:
问题1:NameError: name 'unicode' is not defined
这个错误很直接——你用的boilerpipe版本是为Python2写的,在Python3里已经没有unicode()这个函数了(Python3里的str就对应Python2的unicode)。库的源码里还在调用Python2的API,所以触发了这个错误。
问题2:UnicodeDecodeError: 'utf-8' codec can't decode byte 0x8b in position 1
字节0x8b是gzip压缩内容的标志性开头,说明你爬取的网页返回的是gzip压缩后的内容,但boilerpipe没有正确处理这种压缩格式,直接用utf-8去解码原始的压缩字节,自然就失败了。
具体解决步骤
1. 修复boilerpipe的Python3兼容性
找到报错的文件:/Users/Adrian/anaconda3/lib/python3.6/site-packages/boilerpipe/extract/__init__.py,打开后修改两处代码:
- 把第45行的:
改成:self.data = unicode(self.data, encoding)self.data = str(self.data, encoding) - 把第47行的:
改成带类型判断的代码(因为Python3里只有self.data = self.data.decode(encoding)bytes类型需要decode,str不需要):if isinstance(self.data, bytes): self.data = self.data.decode(encoding)
保存文件后,第一个NameError就解决了。
2. 处理网页的gzip压缩内容
boilerpipe的内置请求逻辑没处理压缩,我们可以自己用requests库先请求网页、解压内容,再传给boilerpipe。修改你的代码如下:
import requests import gzip from io import BytesIO from boilerpipe.extract import Extractor # 假设你的Article类已经正确导入 counter = 1 # 假设counter初始值是1 for urls in allUrls: # 打开文件时指定utf-8编码,避免写入时的编码问题 with open(f'article({counter})', 'w', encoding='utf-8') as fileW: articleDate = Article(urls) articleDate.download() articleDate.parse() print(articleDate.publish_date) # 自行请求网页,处理gzip压缩 headers = {"Accept-Encoding": "gzip"} resp = requests.get(urls, headers=headers) resp.encoding = resp.apparent_encoding # 处理压缩内容 if resp.headers.get("Content-Encoding") == "gzip": decompressed_data = gzip.decompress(resp.content) html_content = decompressed_data.decode(resp.encoding) else: html_content = resp.text # 用处理好的HTML内容创建Extractor,而非直接传url extractor = Extractor(extractor='ArticleExtractor', html=html_content) article_text = extractor.getText() # 拼接内容并写入 content = f"{article_text}\n\n\n{articleDate.publish_date}\n\n\n" fileW.write(content) counter += 1
这里还有个小细节:我把open改成了with语句,这样不用手动调用close(),能确保文件正确关闭,避免内容丢失。
3. 额外注意点
- Python3中打开文件时一定要显式指定
encoding='utf-8',否则会用系统默认编码(比如Windows的gbk),写入中文等非ASCII字符时容易出问题。 - 你原来的代码里
fileW.close少了括号,这会导致文件没有真正关闭,内容可能无法正常写入,用with语句可以避免这个问题。
内容的提问来源于stack exchange,提问作者Adrian Coutsoftides
相关产品推荐
相关产品推荐

