使用Python+Selenium爬取wotstars网站表格数据时,TD数据写入二维数组出现行数列数颠倒问题
使用Python+Selenium爬取wotstars网站表格数据时,TD数据写入二维数组出现行数列数颠倒问题
看起来你已经搞定了Selenium触发「View More」按钮的部分,很棒!但在提取表格数据的时候确实踩了个小坑,导致行和列搞反了,我来帮你梳理下问题出在哪,再给你修正后的代码。
背景回顾
你想要爬取wotstars上《坦克世界主机版》的玩家对战数据,一开始用Excel Power Query只能拿到5条近期对战数据,因为网站需要点击「View More」按钮才能加载最近100条记录(该按钮是加载更多数据的触发入口)。你觉得这是学习Python网页爬取的好机会,已经成功用Selenium触发了按钮,能看到完整的100行表格,但在提取<td>数据到二维数组时,出现了行数列数颠倒的问题,无法正常导出成CSV。
问题分析
看你的代码,问题出在遍历表格子元素的逻辑上:
soup.find_all('table')[3].children会包含表格里的所有子节点,包括换行符这类非标签文本节点,而不仅仅是<tr>行标签。- 当你遍历这些非
<tr>的节点时,比如换行符,for td in child会把字符串拆成单个字符来遍历,导致你收集的不是单元格数据,而是乱码一样的单个字符。 - 另外,你在每个
<td>遍历后就把row添加到rows里,这会导致一个完整的行被拆成多个不完整的行,最终数组结构完全混乱,出现你说的“100列而非100行”的问题。
修正后的代码
我调整了遍历逻辑,只处理<tr>标签,并且每个<tr>对应一行,收集完所有<td>后再添加到rows里,同时用Pandas轻松导出CSV:
from selenium import webdriver from selenium.webdriver.common.by import By import time from bs4 import BeautifulSoup import pandas as pd # Set up the WebDriver driver = webdriver.Chrome() # Famous player's page url = 'https://www.wotstars.com/xbox/6757320' # Open the target page driver.get(url) time.sleep(5) login_button = driver.find_elements(By.CLASS_NAME, "_button_1gcqp_2") for login in login_button: print(login.text) # Handle the "Start tracking" related button to load more data if login_button[3].text == 'Start tracking': login_button[4].click() print("Button activated, loading more match data...") time.sleep(2) # Parse page source with BeautifulSoup soup = BeautifulSoup(driver.page_source, 'html.parser') # Locate the target table and extract all <tr> rows target_table = soup.find_all('table')[3] rows = [] # Iterate through each row in the table for tr in target_table.find_all('tr'): row_data = [] # Iterate through each cell in the current row for td in tr.find_all('td'): # Clean up the text by removing newlines and extra spaces clean_text = td.text.strip().replace('\n', '').replace('\r', '') row_data.append(clean_text) # Only add non-empty rows to avoid headers or empty lines if row_data: rows.append(row_data) # Create DataFrame with first row as column headers df = pd.DataFrame(rows[1:], columns=rows[0]) # Export to CSV file df.to_csv('wotstars_match_data.csv', index=False, encoding='utf-8') print(f"Successfully exported {len(df)} match records to wotstars_match_data.csv") # Close the browser properly driver.quit()
关键改进点
- 精准定位行元素:用
target_table.find_all('tr')直接获取所有行标签,过滤掉无关的文本节点,避免无效遍历。 - 完整收集一行数据:先把当前行的所有
<td>文本收集到row_data里,确认行非空后再添加到rows数组,保证每行是完整的单元格集合。 - 文本清理优化:用
strip()去掉首尾空格,再替换掉换行符,让导出的数据更整洁。 - Pandas便捷导出:直接用DataFrame处理表头和数据,一键导出CSV,省去手动处理二维数组的麻烦,还能避免格式错误。
额外小贴士
- 尽量避免用
time.sleep()固定等待时间,可以换成Selenium的显式等待(WebDriverWait),比如等待表格行数加载到100,或者“View More”按钮消失,这样代码更稳定,不会因为网络慢导致数据加载不完整。 - 确认
table[3]确实是你要的目标表格,有时候网页结构变化会导致索引变化,最好给表格加个更精准的定位(比如根据class属性),避免后续网页更新导致代码失效。
备注:内容来源于stack exchange,提问作者Bob
相关产品推荐
相关产品推荐

