You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取维基百科表格时遭遇AttributeError: 'NoneType' object has no attribute 'find_all'错误求助

解决BeautifulSoup爬取维基百科表格时的AttributeError问题

咱先拆解下你遇到的问题:AttributeError: 'NoneType' object has no attribute 'find_all' 本质是你的代码没找到目标表格——soup1.find('table', {'class':'wikitable sortable jquery-tablesorter'}) 返回了None,之后调用table.find_all()自然就报错了。

问题根源

你用的表格class里,jquery-tablesorter是前端JavaScript动态添加的类,通过requests.get()获取的静态HTML页面里根本不存在这个类名,所以find()方法找不到匹配的表格元素,返回了None。

修正方案

下面是调整后的代码,同时优化了表头提取的逻辑:

import requests
from bs4 import BeautifulSoup

URL = "https://en.wikipedia.org/wiki/List_of_most-viewed_YouTube_videos"
page = requests.get(URL)

# 先确认页面请求是否成功
if page.status_code != 200:
    print(f"请求失败,状态码:{page.status_code}")
else:
    soup1 = BeautifulSoup(page.text, 'lxml')
    # 只使用页面静态渲染时就存在的class定位表格
    table = soup1.find('table', {'class': 'wikitable sortable'})
    
    # 增加容错判断,避免找不到表格直接崩溃
    if not table:
        print("未找到目标表格,可能页面结构已更新")
    else:
        headers = []
        # 定位表头行,遍历<th>标签提取每个表头项
        header_row = table.find('tr')
        for th in header_row.find_all('th'):
            # 清理文本,去掉多余换行、空格
            clean_header = th.get_text(strip=True)
            headers.append(clean_header)
        
        print("提取到的表头:")
        for idx, header in enumerate(headers, 1):
            print(f"{idx}. {header}")

代码改进点说明

  • 增加请求状态检查:避免因为网络问题或页面访问限制导致后续代码报错
  • 修正表格定位class:去掉动态生成的jquery-tablesorter,只保留页面原生的wikitable sortable,确保能找到目标表格
  • 添加容错判断:如果表格不存在(比如页面结构更新),会给出提示而不是直接抛出异常
  • 优化表头提取逻辑:直接遍历表头行的<th>标签,精准提取每个表头项,而不是把整行文本混在一起

内容的提问来源于stack exchange,提问作者bishwajit bhattacharya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 21:22:42