You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3使用boilerpipe爬取网页遇Unicode相关错误求助

解决Python3中使用boilerpipe爬取网页时的Unicode错误问题

嘿,我来帮你拆解下你遇到的这两个问题,其实根源都是boilerpipe这个库对Python3的兼容性不太好,加上网页内容的压缩处理导致的:

问题1:NameError: name 'unicode' is not defined

这个错误很直接——你用的boilerpipe版本是为Python2写的,在Python3里已经没有unicode()这个函数了(Python3里的str就对应Python2的unicode)。库的源码里还在调用Python2的API,所以触发了这个错误。

问题2:UnicodeDecodeError: 'utf-8' codec can't decode byte 0x8b in position 1

字节0x8b是gzip压缩内容的标志性开头,说明你爬取的网页返回的是gzip压缩后的内容,但boilerpipe没有正确处理这种压缩格式,直接用utf-8去解码原始的压缩字节,自然就失败了。


具体解决步骤

1. 修复boilerpipe的Python3兼容性

找到报错的文件:/Users/Adrian/anaconda3/lib/python3.6/site-packages/boilerpipe/extract/__init__.py,打开后修改两处代码:

  • 把第45行的:
    self.data = unicode(self.data, encoding)
    
    改成:
    self.data = str(self.data, encoding)
    
  • 把第47行的:
    self.data = self.data.decode(encoding)
    
    改成带类型判断的代码(因为Python3里只有bytes类型需要decode,str不需要):
    if isinstance(self.data, bytes):
        self.data = self.data.decode(encoding)
    

保存文件后,第一个NameError就解决了。

2. 处理网页的gzip压缩内容

boilerpipe的内置请求逻辑没处理压缩,我们可以自己用requests库先请求网页、解压内容,再传给boilerpipe。修改你的代码如下:

import requests
import gzip
from io import BytesIO
from boilerpipe.extract import Extractor
# 假设你的Article类已经正确导入

counter = 1  # 假设counter初始值是1
for urls in allUrls:
    # 打开文件时指定utf-8编码,避免写入时的编码问题
    with open(f'article({counter})', 'w', encoding='utf-8') as fileW:
        articleDate = Article(urls)
        articleDate.download()
        articleDate.parse()
        print(articleDate.publish_date)
        
        # 自行请求网页,处理gzip压缩
        headers = {"Accept-Encoding": "gzip"}
        resp = requests.get(urls, headers=headers)
        resp.encoding = resp.apparent_encoding
        
        # 处理压缩内容
        if resp.headers.get("Content-Encoding") == "gzip":
            decompressed_data = gzip.decompress(resp.content)
            html_content = decompressed_data.decode(resp.encoding)
        else:
            html_content = resp.text
        
        # 用处理好的HTML内容创建Extractor,而非直接传url
        extractor = Extractor(extractor='ArticleExtractor', html=html_content)
        article_text = extractor.getText()
        
        # 拼接内容并写入
        content = f"{article_text}\n\n\n{articleDate.publish_date}\n\n\n"
        fileW.write(content)
    
    counter += 1

这里还有个小细节:我把open改成了with语句,这样不用手动调用close(),能确保文件正确关闭,避免内容丢失。

3. 额外注意点

  • Python3中打开文件时一定要显式指定encoding='utf-8',否则会用系统默认编码(比如Windows的gbk),写入中文等非ASCII字符时容易出问题。
  • 你原来的代码里fileW.close少了括号,这会导致文件没有真正关闭,内容可能无法正常写入,用with语句可以避免这个问题。

内容的提问来源于stack exchange,提问作者Adrian Coutsoftides

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:11:35