You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决抓取含跨行合并单元格的维基百科指定表格问题?

解决维基百科表格跨行单元格的DataFrame填充问题

问题说明

  • 需求:抓取指定维基百科页面的第一个表格,把带有跨行(rowspan>1)的单元格值自动填充到对应的所有行中(比如Contests列首项跨行9行,前9行都要填充“9”)
  • 遇到的问题:第一行里cells[0]对应Contests、cells[1]对应Country、cells[2]对应City,这三个单元格跨行后,第二行的HTML里没有这些单元格,导致后续行的cells索引直接错位(第二行的cells[0]变成了Venue),而且各列的跨行数还不一致,原代码无法处理这种情况
  • 尝试的代码:
import requests
from bs4 import BeautifulSoup
import pandas as pd

url = 'https://en.wikipedia.org/wiki/List_of_Eurovision_Song_Contest_host_cities'
response = requests.get(url)
soup = BeautifulSoup(response.content, "html.parser")

# Create an empty DataFrame with desired column headers
df = pd.DataFrame(columns=['Contests', 'Country', 'City', 'Venue', 'Year', 'Ref'])

for index, row in enumerate(soup.find_all('tr')):
    if index == 0:  # Skip the first header row
        continue

    cells = row.find_all(['td', 'th'])
    
    country_value = None
    if cells[0].has_attr('rowspan'):
        contests_value = cells[0].get_text(strip=True)
        contests_rowspan = int(cells[0]['rowspan'])
        contests_values = [contests_value] * contests_rowspan # Replicate the value the required number of time
        df = df.append(pd.DataFrame({'Contests': contests_values}), ignore_index=True)

    if cells[1].has_attr('rowspan'):
        country_value = cells[1].get_text(strip=True)
        country_rowspan = int(cells[1]['rowspan'])
        country_values = [country_value] * country_rowspan
        df = df.append(pd.DataFrame({'Country': country_values}), ignore_index=True)

    if cells[2].has_attr('rowspan'):
        print(cells[2])
        city_value = cells[2].get_text(strip=True)
        city_rowspan = int(cells[2]['rowspan'])
        city_values = [city_value] * city_rowspan
        df = df.append(pd.DataFrame({'City': city_values}), ignore_index=True)
    
    venue_value = cells[3].get_text(strip=True)
    year_value = cells[4].get_text(strip=True)
    ref_value = cells[5].get_text(strip=True)
    
    for _ in range(max(contests_rowspan, country_rowspan, city_rowspan)):
            df = df.append({'Venue': venue_value, 'Year': year_value, 'Ref': ref_value}, ignore_index=True)

df.head()

解决方案

核心思路是维护一个待填充状态字典,记录当前哪些列还需要延续跨行的值,以及剩余需要填充的行数。每处理一行时,先把待填充的列值填入当前行,再处理当前行的单元格,更新待填充状态,这样就能避免索引错位的问题。

完整代码:

import requests
from bs4 import BeautifulSoup
import pandas as pd

url = 'https://en.wikipedia.org/wiki/List_of_Eurovision_Song_Contest_host_cities'
response = requests.get(url)
soup = BeautifulSoup(response.content, "html.parser")

# 定义表格列顺序
columns = ['Contests', 'Country', 'City', 'Venue', 'Year', 'Ref']
# 初始化待填充字典:键为列名,值为(剩余行数, 填充值)
pending_cols = {}
rows_data = []

# 获取表格的所有行(跳过表头行)
table_rows = soup.find('table', class_='wikitable').find_all('tr')[1:]

for row in table_rows:
    current_row = {col: None for col in columns}
    # 第一步:填充待延续的跨行值
    for col in pending_cols:
        current_row[col] = pending_cols[col][1]
        # 剩余行数减1,减到0就移除
        pending_cols[col] = (pending_cols[col][0] - 1, pending_cols[col][1])
        if pending_cols[col][0] == 0:
            del pending_cols[col]
    
    # 第二步:处理当前行的单元格,匹配对应的列
    cells = row.find_all(['td', 'th'])
    col_index = 0
    for cell in cells:
        # 跳过已经被待填充字典处理过的列
        while col_index < len(columns) and columns[col_index] in pending_cols:
            col_index += 1
        if col_index >= len(columns):
            break
        
        cell_text = cell.get_text(strip=True)
        current_col = columns[col_index]
        current_row[current_col] = cell_text
        
        # 如果单元格有跨行,更新待填充字典
        if cell.has_attr('rowspan'):
            rowspan = int(cell['rowspan'])
            # 剩余行数要减1,因为当前行已经用了一次
            pending_cols[current_col] = (rowspan - 1, cell_text)
        
        col_index += 1
    
    rows_data.append(current_row)

# 转成DataFrame
df = pd.DataFrame(rows_data)
print(df.head(10))

代码说明

  1. pending_cols字典:用来追踪需要延续的跨行列,比如当一个单元格rowspan=9时,就会记录该列还需要填充8行(因为当前行已经用了1行)
  2. 每行处理逻辑:先填充待延续的列值,再处理当前行的单元格,自动跳过已经被待填充覆盖的列,避免索引错位
  3. 适配不同跨行数:不管各列的跨行数是否一致,都能正确跟踪和填充,不会出现数据错位的情况

内容的提问来源于stack exchange,提问作者Q.Ask

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 13:55:00