网页抓取时维基百科存在多个wikitable如何获取目标表格?
解决维基百科多wikitable抓取问题
方法1:通过索引选择目标表格
你当前用soup.find()只会返回第一个匹配的表格,改用soup.find_all()就能拿到页面上所有带wikitable类的表格,返回结果是一个列表,直接通过索引就能选到你要的表格。
比如你要的是第二个表格,就取索引[1](列表从0开始计数),修改后的代码如下:
import requests from bs4 import BeautifulSoup import pandas as pd url_3 = 'https://en.wikipedia.org/wiki/List_of_Seattle_Seahawks_seasons' response_3 = requests.get(url_3) soup_3 = BeautifulSoup(response_3.text,'html.parser') # 获取所有wikitable表格 all_tables = soup_3.find_all('table', attrs={'class':'wikitable'}) # 选择第二个表格(索引1) target_table = all_tables[1] seahawks_df = pd.read_html(str(target_table))[0] print(seahawks_df)
方法2:通过额外属性精准定位
如果担心页面表格顺序变化导致索引失效,可以观察目标表格的其他特征——比如你要的赛季数据表格还带有sortable类。直接结合多个class属性定位,更稳妥:
import requests from bs4 import BeautifulSoup import pandas as pd url_3 = 'https://en.wikipedia.org/wiki/List_of_Seattle_Seahawks_seasons' response_3 = requests.get(url_3) soup_3 = BeautifulSoup(response_3.text,'html.parser') # 同时匹配wikitable和sortable类的表格 target_table = soup_3.find('table', attrs={'class': ['wikitable', 'sortable']}) seahawks_df = pd.read_html(str(target_table))[0] print(seahawks_df)
这样就能直接拿到你需要的赛季数据表格,不用再受第一个表格的干扰。
内容的提问来源于stack exchange,提问作者Matthew P
相关产品推荐
相关产品推荐

