VBA转Python开发者求助:表格爬取后表头与数据写入二维数组问题
表格爬取与数据整理解决方案
核心实现思路
- 直接跳过无效的第一行数据
- 按行交替识别表头(原表格偶数行)和数据行(原表格奇数行)
- 将数据行的驼峰式地图名称拆分后,每个地图对应一行表头+地图的组合,最终生成二维数组
修正后的完整代码
import requests import re from bs4 import BeautifulSoup url = 'https://modernarmor.worldoftanks.com/en/cms/guides/map-labels-tiers-eras/' response = requests.get(url) soup = BeautifulSoup(response.content, 'html.parser') # 初始化结果二维数组 result_matrix = [] table = soup.find('table') rows = table.find_all('tr') # 跳过第一行无效数据,从第二行开始遍历 for idx, row in enumerate(rows[1:]): cols = row.find_all(['td', 'th']) # 提取单元格文本并去除空白,过滤空字符串 coldata = [ele.text.strip() for ele in cols if ele.text.strip()] # 移除[MERGE]标记 cleaned_data = [item.replace('[MERGE]', '') for item in coldata] # 跳过空行 if not cleaned_data: continue # 偶数索引(原表格的偶数行)作为表头 if idx % 2 == 0: current_header = cleaned_data # 奇数索引(原表格的奇数行)作为数据行 else: # 将驼峰式名称拆分为逗号分隔的地图列表 map_str = ''.join(cleaned_data) split_maps = re.sub(r'([a-z])([A-Z])', r'\1,\2', map_str).split(',') # 过滤空字符串和过短的无效项 valid_maps = [map_name.strip() for map_name in split_maps if map_name.strip() and len(map_name.strip()) > 1] # 每个地图对应一行,与表头组合后加入结果数组 for map_item in valid_maps: result_matrix.append(current_header + [map_item]) # 打印验证结果 for row in result_matrix: print(row)
关键代码说明
- 跳过无效行:通过
rows[1:]直接跳过原表格第一行,避免处理空内容 - 表头识别:利用
enumerate的索引,idx%2==0对应原表格的偶数行,保存为当前表头 - 数据拆分:用正则
re.sub(r'([a-z])([A-Z])', r'\1,\2', map_str)拆分驼峰命名,过滤掉空字符串和过短无效项 - 二维数组构建:遍历拆分后的地图列表,将每个地图与当前表头拼接,添加到结果数组中,确保每行数据与表头一一对应
内容的提问来源于stack exchange,提问作者Bob
相关产品推荐
相关产品推荐

