You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取网页表格单列数据遇tbody为None报错如何解决

问题原因
  • 核心问题:浏览器开发者工具中显示的<tbody>标签是浏览器渲染页面时自动补充的节点,你通过requests获取的原始HTML源码中,table.toplist下没有<tbody>层级,直接嵌套<tr>元素,因此调用journal_list.tbody会返回None类型,触发属性错误。
  • 次要语法问题:你贴出的代码中BeautifulSoup实例化行末尾缺少闭合右括号,运行时会先触发语法错误,注意补全。
可行实现方案

直接定位table.toplist下的所有<tr>元素即可,不需要经过<tbody>层级,示例代码如下:

import requests
from bs4 import BeautifulSoup

URL = "https://ideas.repec.org/top/top.journals.simple.html"
# 添加请求头模拟浏览器,避免被反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}
html_content = requests.get(URL, headers=headers).text
# 补全之前缺失的右括号
soup = BeautifulSoup(html_content, "lxml")

journal_table = soup.find("table", attrs={"class": "toplist"})
# 直接获取所有tr,不需要找tbody
all_rows = journal_table.find_all("tr")

# 先提取表头确认Journal列的索引
headings = [td.b.text.strip() for td in all_rows[0].find_all("td")]
journal_col_index = headings.index("Journal")

# 提取所有期刊名称
journal_names = []
# 跳过表头行,从第二行开始遍历
for row in all_rows[1:]:
    tds = row.find_all("td")
    if len(tds) > journal_col_index:
        journal_name = tds[journal_col_index].text.strip()
        journal_names.append(journal_name)

# 输出结果示例
print(f"共提取到{len(journal_names)}个期刊")
print(journal_names[:10])

内容的提问来源于stack exchange,提问作者lev72748

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 15:15:03