You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页表格爬取报错:AttributeError: 'NoneType'无find_all属性,如何解决?

解决网页表格爬取时的AttributeError错误

你的报错核心原因是:soup.find("tbody")没有找到对应的<tbody>元素,返回了None,后续调用table_body.find_all("tr")自然会抛出AttributeError——因为None对象没有find_all方法。

下面是具体的解决方案:

原因1:网页原始HTML中不存在<tbody>标签

很多现代网页的<tbody>是浏览器渲染DOM时自动添加的,原始HTML源码里其实没有这个标签。这种情况下直接定位表格下的<tr>即可:

修改后的代码:

# 先定位目标表格(如果知道表格的id/class可以更精准,比如soup.find("table", id="movie-table"))
table = soup.find("table")
rows = []

if table:  # 先判断表格是否存在,避免后续报错
    for row in table.find_all("tr"):
        cols = row.find_all("td")
        # 增加列数判断,防止索引越界报错
        if len(cols) >= 5:
            rd = cols[0].text.strip()
            mv = cols[1].text.strip()
            pb = cols[2].text.strip()
            dg = cols[3].text.strip()
            wg = cols[4].text.strip()
            rows.append([rd, mv, pb, dg, wg])

headers = ["Release Date", "Movie", "Production Budget", "Domestic Gross", "Worldwide Gross"]

原因2:强制查找<tbody>但未做判空处理

如果确认网页源码里确实有<tbody>,可以先判断table_body是否有效再执行循环:

修改后的代码:

table_body = soup.find("tbody")
rows = []

# 先检查table_body是否存在
if table_body:
    for row in table_body.find_all("tr"):
        cols = row.find_all("td")
        # 增加列数判断,避免索引越界
        if len(cols) >= 5:
            rd = cols[0].text.strip()
            mv = cols[1].text.strip()
            pb = cols[2].text.strip()
            dg = cols[3].text.strip()
            wg = cols[4].text.strip()
            rows.append([rd, mv, pb, dg, wg])

headers = ["Release Date", "Movie", "Production Budget", "Domestic Gross", "Worldwide Gross"]

额外建议

  • 验证网页结构:右键网页→查看页面源代码,搜索<tbody>确认是否存在。
  • 精准定位表格:如果页面有多个表格,给find("table")加上属性筛选(如class、id),避免找错目标表格。
  • 处理动态内容:如果表格是JavaScript动态渲染的,requests获取的HTML里可能没有数据,需要用selenium或playwright模拟浏览器加载页面。

内容的提问来源于stack exchange,提问作者Hary

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 02:05:12