使用Pandas抓取维基百科DC动画电影数据时,定位表格返回None求助
解决维基百科DC动画原创电影表格抓取问题
你的代码返回None是因为判断目标表格的逻辑错误:目标"Released films"表格的第一个<th>是"Title",不是"Release date",所以循环里的条件永远不成立,找不到表格。
修正方案
直接通过维基百科表格的class定位目标表格(更可靠,避免表头判断误差),然后提取所需列:
import requests as r from bs4 import BeautifulSoup import pandas as pd # 请求页面 response = r.get("https://en.wikipedia.org/wiki/DC_Universe_Animated_Original_Movies") soup = BeautifulSoup(response.text, "html.parser") # 直接定位"Released films"板块的表格(该表格class为wikitable sortable) target_table = soup.find("table", class_="wikitable sortable") # 提取表格行 rows = target_table.find_all("tr") # 处理表头和数据 data = [] headers = [th.text.strip() for th in rows[0].find_all("th")] # 找到目标列的索引 title_idx = headers.index("Title") release_date_idx = headers.index("Release date") continuity_idx = headers.index("Continuity") for row in rows[1:]: cols = row.find_all("td") # 跳过空行 if len(cols) < len(headers): continue title = cols[title_idx].text.strip() release_date = cols[release_date_idx].text.strip() continuity = cols[continuity_idx].text.strip() data.append({"title": title, "release date": release_date, "continuity": continuity}) # 转换为DataFrame df = pd.DataFrame(data) print(df.head())
关键说明
- 维基百科的可排序表格通常带有
wikitable sortable类,直接用这个定位比遍历所有表格更高效准确。 - 先获取表头的索引,再提取对应列,避免因表格列顺序变化导致的错误。
- 处理空行:表格中可能存在合并单元格的空行,判断列数跳过即可。
内容的提问来源于stack exchange,提问作者Olasubomi
相关产品推荐
相关产品推荐

