如何解决抓取含跨行合并单元格的维基百科指定表格问题?
解决维基百科表格跨行单元格的DataFrame填充问题
问题说明
- 需求:抓取指定维基百科页面的第一个表格,把带有跨行(
rowspan>1)的单元格值自动填充到对应的所有行中(比如Contests列首项跨行9行,前9行都要填充“9”) - 遇到的问题:第一行里
cells[0]对应Contests、cells[1]对应Country、cells[2]对应City,这三个单元格跨行后,第二行的HTML里没有这些单元格,导致后续行的cells索引直接错位(第二行的cells[0]变成了Venue),而且各列的跨行数还不一致,原代码无法处理这种情况 - 尝试的代码:
import requests from bs4 import BeautifulSoup import pandas as pd url = 'https://en.wikipedia.org/wiki/List_of_Eurovision_Song_Contest_host_cities' response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") # Create an empty DataFrame with desired column headers df = pd.DataFrame(columns=['Contests', 'Country', 'City', 'Venue', 'Year', 'Ref']) for index, row in enumerate(soup.find_all('tr')): if index == 0: # Skip the first header row continue cells = row.find_all(['td', 'th']) country_value = None if cells[0].has_attr('rowspan'): contests_value = cells[0].get_text(strip=True) contests_rowspan = int(cells[0]['rowspan']) contests_values = [contests_value] * contests_rowspan # Replicate the value the required number of time df = df.append(pd.DataFrame({'Contests': contests_values}), ignore_index=True) if cells[1].has_attr('rowspan'): country_value = cells[1].get_text(strip=True) country_rowspan = int(cells[1]['rowspan']) country_values = [country_value] * country_rowspan df = df.append(pd.DataFrame({'Country': country_values}), ignore_index=True) if cells[2].has_attr('rowspan'): print(cells[2]) city_value = cells[2].get_text(strip=True) city_rowspan = int(cells[2]['rowspan']) city_values = [city_value] * city_rowspan df = df.append(pd.DataFrame({'City': city_values}), ignore_index=True) venue_value = cells[3].get_text(strip=True) year_value = cells[4].get_text(strip=True) ref_value = cells[5].get_text(strip=True) for _ in range(max(contests_rowspan, country_rowspan, city_rowspan)): df = df.append({'Venue': venue_value, 'Year': year_value, 'Ref': ref_value}, ignore_index=True) df.head()
解决方案
核心思路是维护一个待填充状态字典,记录当前哪些列还需要延续跨行的值,以及剩余需要填充的行数。每处理一行时,先把待填充的列值填入当前行,再处理当前行的单元格,更新待填充状态,这样就能避免索引错位的问题。
完整代码:
import requests from bs4 import BeautifulSoup import pandas as pd url = 'https://en.wikipedia.org/wiki/List_of_Eurovision_Song_Contest_host_cities' response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") # 定义表格列顺序 columns = ['Contests', 'Country', 'City', 'Venue', 'Year', 'Ref'] # 初始化待填充字典:键为列名,值为(剩余行数, 填充值) pending_cols = {} rows_data = [] # 获取表格的所有行(跳过表头行) table_rows = soup.find('table', class_='wikitable').find_all('tr')[1:] for row in table_rows: current_row = {col: None for col in columns} # 第一步:填充待延续的跨行值 for col in pending_cols: current_row[col] = pending_cols[col][1] # 剩余行数减1,减到0就移除 pending_cols[col] = (pending_cols[col][0] - 1, pending_cols[col][1]) if pending_cols[col][0] == 0: del pending_cols[col] # 第二步:处理当前行的单元格,匹配对应的列 cells = row.find_all(['td', 'th']) col_index = 0 for cell in cells: # 跳过已经被待填充字典处理过的列 while col_index < len(columns) and columns[col_index] in pending_cols: col_index += 1 if col_index >= len(columns): break cell_text = cell.get_text(strip=True) current_col = columns[col_index] current_row[current_col] = cell_text # 如果单元格有跨行,更新待填充字典 if cell.has_attr('rowspan'): rowspan = int(cell['rowspan']) # 剩余行数要减1,因为当前行已经用了一次 pending_cols[current_col] = (rowspan - 1, cell_text) col_index += 1 rows_data.append(current_row) # 转成DataFrame df = pd.DataFrame(rows_data) print(df.head(10))
代码说明
- pending_cols字典:用来追踪需要延续的跨行列,比如当一个单元格
rowspan=9时,就会记录该列还需要填充8行(因为当前行已经用了1行) - 每行处理逻辑:先填充待延续的列值,再处理当前行的单元格,自动跳过已经被待填充覆盖的列,避免索引错位
- 适配不同跨行数:不管各列的跨行数是否一致,都能正确跟踪和填充,不会出现数据错位的情况
内容的提问来源于stack exchange,提问作者Q.Ask
相关产品推荐
相关产品推荐

