You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup抓取含th和td的表格数据?解决年份提取问题

解决网页抓取中年份(表头<th>)提取问题

首先修正原代码的核心问题,实现年份数据的提取:

核心问题说明

你的代码未处理表格行内的<th>标签(年份信息就存储在这些表头标签里),同时存在变量未定义的小错误:url仅写了字符串但未赋值给变量。

修改后的完整代码

from bs4 import BeautifulSoup
import requests
import pandas as pd

# 定义目标页面URL
url = "https://en.wikipedia.org/wiki/World_population"
data = requests.get(url).text
soup = BeautifulSoup(data, "html.parser")
tables = soup.find_all('table')

table_index = None
# 遍历找到目标表格后立即终止循环,提升效率
for index, table in enumerate(tables):
    if "Global annual population growth" in str(table):
        table_index = index
        break

# 用列表批量存储数据,比反复调用append更高效
population_list = []

# 同时提取行内的<th>和<td>标签
for row in tables[table_index].tbody.find_all('tr'):
    cells = row.find_all(['th', 'td'])
    # 过滤掉无效行,确保是包含年份、人口、增长率的完整数据行
    if len(cells) == 3:
        year = cells[0].text.strip()
        population = cells[1].text.strip()
        growth = cells[2].text.strip()
        population_list.append({"Year": year, "Population": population, "Growth": growth})

# 转换为DataFrame
population_data = pd.DataFrame(population_list)
print(population_data)

关键修改点

  • 补全url变量赋值:原代码仅写了字符串,未赋值给变量会导致请求失败
  • 同时抓取<th>和<td>:通过row.find_all(['th', 'td'])获取行内所有表头和单元格内容,年份就在每个数据行的第一个<th>中
  • 用列表批量收集数据:避免多次调用DataFrame.append(效率低下),最后一次性转换为DataFrame
  • 增加行有效性判断:通过len(cells) == 3过滤空行或结构不符的无效行

内容的提问来源于stack exchange,提问作者Harvey_001

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 00:52:40