如何用BeautifulSoup抓取含th和td的表格数据?解决年份提取问题
解决网页抓取中年份(表头
<th>)提取问题 首先修正原代码的核心问题,实现年份数据的提取:
核心问题说明
你的代码未处理表格行内的<th>标签(年份信息就存储在这些表头标签里),同时存在变量未定义的小错误:url仅写了字符串但未赋值给变量。
修改后的完整代码
from bs4 import BeautifulSoup import requests import pandas as pd # 定义目标页面URL url = "https://en.wikipedia.org/wiki/World_population" data = requests.get(url).text soup = BeautifulSoup(data, "html.parser") tables = soup.find_all('table') table_index = None # 遍历找到目标表格后立即终止循环,提升效率 for index, table in enumerate(tables): if "Global annual population growth" in str(table): table_index = index break # 用列表批量存储数据,比反复调用append更高效 population_list = [] # 同时提取行内的<th>和<td>标签 for row in tables[table_index].tbody.find_all('tr'): cells = row.find_all(['th', 'td']) # 过滤掉无效行,确保是包含年份、人口、增长率的完整数据行 if len(cells) == 3: year = cells[0].text.strip() population = cells[1].text.strip() growth = cells[2].text.strip() population_list.append({"Year": year, "Population": population, "Growth": growth}) # 转换为DataFrame population_data = pd.DataFrame(population_list) print(population_data)
关键修改点
- 补全
url变量赋值:原代码仅写了字符串,未赋值给变量会导致请求失败 - 同时抓取
<th>和<td>:通过row.find_all(['th', 'td'])获取行内所有表头和单元格内容,年份就在每个数据行的第一个<th>中 - 用列表批量收集数据:避免多次调用
DataFrame.append(效率低下),最后一次性转换为DataFrame - 增加行有效性判断:通过
len(cells) == 3过滤空行或结构不符的无效行
内容的提问来源于stack exchange,提问作者Harvey_001
相关产品推荐
相关产品推荐

