使用BeautifulSoup抓取维基百科指定欧视表格失败求助
问题:无法定位2022年欧洲歌唱大赛半决赛1参赛结果表格
我在抓取维基百科2022年欧洲歌唱大赛页面时,无法定位到Participants and results of the first semi-final of the Eurovision Song Contest 2022表格,尝试两种方法均失败:
- 使用
find_all('table', class_="wikitable sortable plainrowheaders")后,索引[2]拿到的是目标表格前的分组表,索引[3]是详细评审投票表,直接跳过了目标表格 - 通过标题文本精准匹配时,程序返回“Desired table not found”
解决方案1:精准匹配表格类名
目标表格的完整类名包含jquery-tablesorter,可以用CSS选择器直接匹配所有目标类的组合,避免索引偏移问题:
import requests from bs4 import BeautifulSoup url = 'https://en.wikipedia.org/wiki/Eurovision_Song_Contest_2022' response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") # 用CSS选择器匹配包含所有指定类的表格 target_table = soup.select_one('table.wikitable.sortable.plainrowheaders.jquery-tablesorter')
解决方案2:模糊匹配标题文本
页面表格的标题可能存在隐藏空格、换行符,不要用全量文本匹配,改用关键词包含匹配:
import requests from bs4 import BeautifulSoup url = 'https://en.wikipedia.org/wiki/Eurovision_Song_Contest_2022' response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") # 用关键词匹配标题,避免字符差异 desired_caption_keyword = "Participants and results of the first semi-final" desired_table = None for table in soup.find_all('table'): caption = table.find('caption') if caption and desired_caption_keyword in caption.text.strip(): desired_table = table break if desired_table: print("找到目标表格") # 在这里编写表格解析逻辑 else: print("未找到目标表格")
解决方案3:通过父标题定位
目标表格位于First semi-final标题下方,可以先定位该标题,再获取其后续的第一个表格:
import requests from bs4 import BeautifulSoup url = 'https://en.wikipedia.org/wiki/Eurovision_Song_Contest_2022' response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") # 找到First semi-final的锚点元素 heading_anchor = soup.find('span', id="First_semi-final") if heading_anchor: # 获取锚点所在的h3标签,再找下一个兄弟表格 h3_tag = heading_anchor.parent target_table = h3_tag.find_next_sibling('table', class_="wikitable sortable") if target_table: print("找到目标表格")
内容的提问来源于stack exchange,提问作者Q.Ask
相关产品推荐
相关产品推荐

