You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python+Selenium爬取wotstars网站表格数据时,TD数据写入二维数组出现行数列数颠倒问题

使用Python+Selenium爬取wotstars网站表格数据时,TD数据写入二维数组出现行数列数颠倒问题

看起来你已经搞定了Selenium触发「View More」按钮的部分,很棒!但在提取表格数据的时候确实踩了个小坑,导致行和列搞反了,我来帮你梳理下问题出在哪,再给你修正后的代码。

背景回顾

你想要爬取wotstars上《坦克世界主机版》的玩家对战数据,一开始用Excel Power Query只能拿到5条近期对战数据,因为网站需要点击「View More」按钮才能加载最近100条记录(该按钮是加载更多数据的触发入口)。你觉得这是学习Python网页爬取的好机会,已经成功用Selenium触发了按钮,能看到完整的100行表格,但在提取<td>数据到二维数组时,出现了行数列数颠倒的问题,无法正常导出成CSV。

问题分析

看你的代码,问题出在遍历表格子元素的逻辑上:

  • soup.find_all('table')[3].children 会包含表格里的所有子节点,包括换行符这类非标签文本节点,而不仅仅是<tr>行标签。
  • 当你遍历这些非<tr>的节点时,比如换行符,for td in child会把字符串拆成单个字符来遍历,导致你收集的不是单元格数据,而是乱码一样的单个字符。
  • 另外,你在每个<td>遍历后就把row添加到rows里,这会导致一个完整的行被拆成多个不完整的行,最终数组结构完全混乱,出现你说的“100列而非100行”的问题。

修正后的代码

我调整了遍历逻辑,只处理<tr>标签,并且每个<tr>对应一行,收集完所有<td>后再添加到rows里,同时用Pandas轻松导出CSV:

from selenium import webdriver
from selenium.webdriver.common.by import By
import time
from bs4 import BeautifulSoup
import pandas as pd

# Set up the WebDriver
driver = webdriver.Chrome()
# Famous player's page
url = 'https://www.wotstars.com/xbox/6757320'

# Open the target page
driver.get(url)
time.sleep(5)

login_button = driver.find_elements(By.CLASS_NAME, "_button_1gcqp_2")
for login in login_button:
    print(login.text)

# Handle the "Start tracking" related button to load more data
if login_button[3].text == 'Start tracking':
    login_button[4].click()

print("Button activated, loading more match data...")
time.sleep(2)

# Parse page source with BeautifulSoup
soup = BeautifulSoup(driver.page_source, 'html.parser')

# Locate the target table and extract all <tr> rows
target_table = soup.find_all('table')[3]
rows = []

# Iterate through each row in the table
for tr in target_table.find_all('tr'):
    row_data = []
    # Iterate through each cell in the current row
    for td in tr.find_all('td'):
        # Clean up the text by removing newlines and extra spaces
        clean_text = td.text.strip().replace('\n', '').replace('\r', '')
        row_data.append(clean_text)
    # Only add non-empty rows to avoid headers or empty lines
    if row_data:
        rows.append(row_data)

# Create DataFrame with first row as column headers
df = pd.DataFrame(rows[1:], columns=rows[0])
# Export to CSV file
df.to_csv('wotstars_match_data.csv', index=False, encoding='utf-8')
print(f"Successfully exported {len(df)} match records to wotstars_match_data.csv")

# Close the browser properly
driver.quit()

关键改进点

  1. 精准定位行元素:用target_table.find_all('tr')直接获取所有行标签,过滤掉无关的文本节点,避免无效遍历。
  2. 完整收集一行数据:先把当前行的所有<td>文本收集到row_data里,确认行非空后再添加到rows数组,保证每行是完整的单元格集合。
  3. 文本清理优化:用strip()去掉首尾空格,再替换掉换行符,让导出的数据更整洁。
  4. Pandas便捷导出:直接用DataFrame处理表头和数据,一键导出CSV,省去手动处理二维数组的麻烦,还能避免格式错误。

额外小贴士

  • 尽量避免用time.sleep()固定等待时间,可以换成Selenium的显式等待(WebDriverWait),比如等待表格行数加载到100,或者“View More”按钮消失,这样代码更稳定,不会因为网络慢导致数据加载不完整。
  • 确认table[3]确实是你要的目标表格,有时候网页结构变化会导致索引变化,最好给表格加个更精准的定位(比如根据class属性),避免后续网页更新导致代码失效。

备注:内容来源于stack exchange,提问作者Bob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 09:27:56