You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

VBA转Python开发者求助:表格爬取后表头与数据写入二维数组问题

表格爬取与数据整理解决方案

核心实现思路

  • 直接跳过无效的第一行数据
  • 按行交替识别表头(原表格偶数行)和数据行(原表格奇数行)
  • 将数据行的驼峰式地图名称拆分后,每个地图对应一行表头+地图的组合,最终生成二维数组

修正后的完整代码

import requests
import re
from bs4 import BeautifulSoup

url = 'https://modernarmor.worldoftanks.com/en/cms/guides/map-labels-tiers-eras/'

response = requests.get(url)
soup = BeautifulSoup(response.content, 'html.parser')

# 初始化结果二维数组
result_matrix = []

table = soup.find('table')
rows = table.find_all('tr')

# 跳过第一行无效数据,从第二行开始遍历
for idx, row in enumerate(rows[1:]):
    cols = row.find_all(['td', 'th'])
    # 提取单元格文本并去除空白,过滤空字符串
    coldata = [ele.text.strip() for ele in cols if ele.text.strip()]
    # 移除[MERGE]标记
    cleaned_data = [item.replace('[MERGE]', '') for item in coldata]
    
    # 跳过空行
    if not cleaned_data:
        continue
    
    # 偶数索引(原表格的偶数行)作为表头
    if idx % 2 == 0:
        current_header = cleaned_data
    # 奇数索引(原表格的奇数行)作为数据行
    else:
        # 将驼峰式名称拆分为逗号分隔的地图列表
        map_str = ''.join(cleaned_data)
        split_maps = re.sub(r'([a-z])([A-Z])', r'\1,\2', map_str).split(',')
        # 过滤空字符串和过短的无效项
        valid_maps = [map_name.strip() for map_name in split_maps if map_name.strip() and len(map_name.strip()) > 1]
        
        # 每个地图对应一行,与表头组合后加入结果数组
        for map_item in valid_maps:
            result_matrix.append(current_header + [map_item])

# 打印验证结果
for row in result_matrix:
    print(row)

关键代码说明

  • 跳过无效行:通过rows[1:]直接跳过原表格第一行,避免处理空内容
  • 表头识别:利用enumerate的索引,idx%2==0对应原表格的偶数行,保存为当前表头
  • 数据拆分:用正则re.sub(r'([a-z])([A-Z])', r'\1,\2', map_str)拆分驼峰命名,过滤掉空字符串和过短无效项
  • 二维数组构建:遍历拆分后的地图列表,将每个地图与当前表头拼接,添加到结果数组中,确保每行数据与表头一一对应

内容的提问来源于stack exchange,提问作者Bob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 10:25:17