You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取数据解析为类表格格式并适配Excel的技术求助

解决方案

步骤1:解析表格数据(表头+内容)

你已经定位到了tbody元素,接下来需要遍历其中的每一行<tr>,提取每行内<td>的文本内容。同时建议抓取表头,让最终导出的Excel结构更完整。

修改后的代码如下:

import chromedriver_autoinstaller
from selenium import webdriver
from bs4 import BeautifulSoup
import pandas as pd

chromedriver_autoinstaller.install()

driver = webdriver.Chrome()
driver.get('https://results.advancedeventsystems.com/event/PTAwMDAwMjkwMjQ90/divisions/131313/standings')

html = driver.page_source
soup = BeautifulSoup(html, 'html.parser')

# 抓取表头文本
thead = soup.find('thead', 'k-table-thead')
headers = [th.get_text(strip=True) for th in thead.find_all('th')]

# 抓取表格主体数据
tbody = soup.find('tbody', 'k-table-tbody')  # 使用find而非find_all,页面中目标tbody仅一个
rows = []
for tr in tbody.find_all('tr'):
    # 提取每行单元格的文本,strip去除多余空格换行
    row_data = [td.get_text(strip=True) for td in tr.find_all('td')]
    rows.append(row_data)

driver.quit()  # 关闭浏览器实例

步骤2:导出数据到Excel

借助pandas库将整理好的表头和数据转换为DataFrame,再导出为Excel文件:

# 创建DataFrame结构
df = pd.DataFrame(rows, columns=headers)

# 导出到Excel,index=False不保留行索引
df.to_excel('赛事成绩表.xlsx', index=False, engine='openpyxl')
print("Excel文件已成功导出")

注意事项

  • 若目标页面存在分页或动态加载内容,需额外处理滚动、分页点击逻辑,但当前示例URL的表格为一次性加载,上述代码可直接运行。
  • 运行前需安装依赖库:pip install pandas openpyxl

内容的提问来源于stack exchange,提问作者phoenix06sa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 12:24:51