You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Survivor维基页面爬取多张表格并解决匹配失败问题?

解决Survivor维基页面表格提取失败问题

问题描述

尝试从Survivor维基页面提取三类特定表格:contestant、season summary和voting history,目前仅能成功获取contestant表格,系统提示无法找到season summary和voting history表格,最终目标是将所有表格合并为一个DataFrame以便清洗和处理。

原代码

import pandas as pd

list_of_seasons = ['41', '42', '43', '44', '45', '46']
season_start = 41
contestants = {}
season_summary = {}
voting_history = {}

for i in list_of_seasons :
    contestants[i] = pd.read_html('https://en.wikipedia.org/wiki/Survivor_' + str(season_start), match='contestants')
    season_summary[i] = pd.read_html('https://en.wikipedia.org/wiki/Survivor_' + str(season_start), match='season summary')
    voting_history[i] = pd.read_html('https://en.wikipedia.org/wiki/Survivor_' + str(season_start), match='voting history')
    season_start = season_start + 1

print(contestants['45'])
print(season_summary['45'])
print(voting_history['45'])

错误信息

Traceback (most recent call last):
  File "c:\Users\bsjes\Documents\Code\Personal Projects\Survivor Data Grabber\SurvivorWikiRipper_0.2.py", line 13, in <module>
    season_summary[i] = pd.read_html('https://en.wikipedia.org/wiki/Survivor_' + str(season_start), match='season summary')        
                        ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^        
  File "C:\Users\bsjes\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\html.py", line 1246, in read_html       
    return _parse(
           ^^^^^^^
  File "C:\Users\bsjes\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\html.py", line 1009, in _parse
    raise retained
  File "C:\Users\bsjes\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\html.py", line 989, in _parse
    tables = p.parse_tables()
             ^^^^^^^^^^^^^^^^
  File "C:\Users\bsjes\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\html.py", line 249, in parse_tables     
    tables = self._parse_tables(self._build_doc(), self.match, self.attrs)
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "C:\Users\bsjes\AppData\Local\Programs\Python\Python312\Lib\site-packages\pandas\io\html.py", line 622, in _parse_tables    
    raise ValueError(f"No tables found matching pattern {repr(match.pattern)}")
ValueError: No tables found matching pattern 'season summary'

解决方案

不需要更换Python工具包,通过调整匹配逻辑和处理页面结构差异即可解决问题,具体调整点如下:

1. 修正匹配关键词的大小写与格式

维基页面的表格标题通常是首字母大写格式(如Season Summary、Voting History),原代码用全小写匹配会找不到目标表格。可以通过正则表达式忽略大小写,或直接匹配正确的标题格式。

2. 增加异常处理,兼容页面结构差异

部分页面的目标表格可能没有直接的caption匹配关键词,这时候可以先获取所有表格,再通过表格的列名特征筛选目标表格。

3. 简化循环变量逻辑

原循环中同时使用list_of_seasons字符串和season_start整数,容易出现变量不一致,直接遍历赛季编号字符串更简洁。

调整后的代码

import pandas as pd
import re

list_of_seasons = ['41', '42', '43', '44', '45', '46']
contestants = {}
season_summary = {}
voting_history = {}

for season in list_of_seasons:
    url = f'https://en.wikipedia.org/wiki/Survivor_{season}'
    
    # 提取参赛者表格:忽略大小写匹配关键词
    contestants[season] = pd.read_html(url, match=r'contestants', flags=re.IGNORECASE)
    
    # 提取赛季总结表格:先尝试匹配标题,失败则按列名筛选
    try:
        season_summary[season] = pd.read_html(url, match=r'Season Summary', flags=re.IGNORECASE)
    except ValueError:
        all_tables = pd.read_html(url)
        for table in all_tables:
            # 赛季总结表格通常包含Episode、Air date等列
            if 'Episode' in table.columns or 'Air date' in table.columns:
                season_summary[season] = [table]
                break
    
    # 提取投票历史表格:同理处理
    try:
        voting_history[season] = pd.read_html(url, match=r'Voting History', flags=re.IGNORECASE)
    except ValueError:
        all_tables = pd.read_html(url)
        for table in all_tables:
            # 投票历史表格通常包含Vote、Eliminated等列
            if 'Vote' in table.columns or 'Eliminated' in table.columns:
                voting_history[season] = [table]
                break

# 示例输出(取第一个表格的前几行)
print("参赛者表格:")
print(contestants['45'][0].head())
if '45' in season_summary:
    print("\n赛季总结表格:")
    print(season_summary['45'][0].head())
if '45' in voting_history:
    print("\n投票历史表格:")
    print(voting_history['45'][0].head())

后续合并DataFrame提示

提取到所有表格后,可按赛季编号将三类表格分别合并,再通过共同字段(如参赛者姓名、赛季编号)关联成最终的DataFrame,例如:

# 合并所有赛季的参赛者表格
all_contestants = pd.concat([df[0].assign(Season=season) for season, df in contestants.items()], ignore_index=True)
# 同理合并赛季总结和投票历史表格,再进行关联

内容的提问来源于stack exchange,提问作者Brandon Jessup

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 00:42:03