You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手求助:如何爬取aspx.net网站点击后生成的表格?

解决ASP.NET网站表格爬取问题(Selenium实现)

我是Python新手,尝试从网站https://tradereport.moc.go.th/Report/ReportEng.aspx?Report=HarmonizeCommodity&Lang=Eng&ImExType=1&Option=1下载表格。该网站基于ASP.NET,难点在于页面源码里看不到生成报告所需点击的按钮。我先做简化测试,只点击“ReviewReport”按钮获取结果表格,网页能显示表格,但无法从页面源码中爬取到数据。以下是我写的Selenium代码,请求帮忙解决表格爬取问题:

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.support.ui import Select
from selenium.common.exceptions import NoSuchElementException
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.keys import Keys
import pandas as pd

Year = ('2021')
Month = ('11')
HScode = ('721990')

options = Options()
#options.headless = True
#driver = webdriver.Chrome('C:\\Python\Scripts\chromedriver.exe', options=options)
driver = webdriver.Firefox(executable_path=r'C:\Python\geckodriver.exe')
driver.get("https://tradereport.moc.go.th/Report/ReportEng.aspx?      Report=HarmonizeCommodity&Lang=Eng&ImExType=1&Option=1")
time.sleep(5)
driver.get("https://tradereport.moc.go.th/Report/ReportEng.aspx?    Report=HarmonizeCommodity&Lang=Eng&ImExType=1&Option=1") #reenter the url again to get the table page
time.sleep(5)  


driver.find_element_by_id("ddlYear").send_keys(Year)  #Working fine
driver.find_element_by_id("ddlMonth").send_keys(Month)  #Working fine
driver.find_element_by_id("txtHsCode").send_keys(HScode) #Working fine
submitbttn = driver.find_element_by_id("btnSubmit")       #Working fine
submitbttn.click()

time.sleep(5)

f=open("d:\\page.txt","w")
f.write(driver.page_source)
f.close()

print ("************* Scrapping data done**********************")
driver.quit`

问题分析

页面显示的表格大概率是动态加载在iframe中,或是通过AJAX异步渲染的,直接获取driver.page_source无法拿到表格数据。另外代码存在几个小问题:

  1. 重复调用driver.get()完全没必要,首次加载页面后无需重复访问
  2. 未导入time模块,执行time.sleep()会报错
  3. driver.quit缺少括号,正确写法是driver.quit()

解决方案步骤

1. 检查并切换到iframe

很多ASP.NET网站会把报表放在iframe里,先切换到iframe再操作:

# 点击提交后等待iframe加载完成并切换
WebDriverWait(driver, 10).until(EC.frame_to_be_available_and_switch_to_it((By.TAG_NAME, "iframe")))

2. 用显式等待定位表格并提取数据

改用显式等待定位表格元素,再通过Pandas直接读取表格内容:

# 等待表格加载完成,可替换为表格实际ID或其他定位器
table = WebDriverWait(driver, 15).until(
    EC.presence_of_element_located((By.CSS_SELECTOR, "table"))
)
# Pandas读取表格数据
df = pd.read_html(table.get_attribute('outerHTML'))[0]
# 保存为CSV文件
df.to_csv("d:\\trade_report.csv", index=False)

3. 修正代码错误

  • 新增import time导入模块
  • 删除重复的driver.get()调用
  • 将driver.quit改为driver.quit()

完整修正代码示例

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.support.ui import Select
from selenium.common.exceptions import NoSuchElementException
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.keys import Keys
import pandas as pd
import time

Year = '2021'
Month = '11'
HScode = '721990'

options = Options()
# options.headless = True
# driver = webdriver.Chrome('C:\\Python\Scripts\chromedriver.exe', options=options)
driver = webdriver.Firefox(executable_path=r'C:\Python\geckodriver.exe')
driver.get("https://tradereport.moc.go.th/Report/ReportEng.aspx?Report=HarmonizeCommodity&Lang=Eng&ImExType=1&Option=1")
time.sleep(5)

# 填写表单
driver.find_element_by_id("ddlYear").send_keys(Year)
driver.find_element_by_id("ddlMonth").send_keys(Month)
driver.find_element_by_id("txtHsCode").send_keys(HScode)
submitbttn = driver.find_element_by_id("btnSubmit")
submitbttn.click()

# 尝试切换到iframe(表格可能在其中)
try:
    WebDriverWait(driver, 10).until(EC.frame_to_be_available_and_switch_to_it((By.TAG_NAME, "iframe")))
except:
    pass

# 等待表格加载并提取数据
table = WebDriverWait(driver, 15).until(
    EC.presence_of_element_located((By.CSS_SELECTOR, "table"))
)
df = pd.read_html(table.get_attribute('outerHTML'))[0]
print(df.head())
df.to_csv("d:\\trade_report.csv", index=False)

# 保存页面源码(可选)
with open("d:\\page.txt","w", encoding='utf-8') as f:
    f.write(driver.page_source)

print("************* 数据爬取完成**********************")
driver.quit()

注意事项

  • 如果不知道表格ID,可用By.CSS_SELECTOR, "table"或By.TAG_NAME, "table"定位所有表格,再从Pandas返回的列表中选择正确表格
  • 显式等待WebDriverWait比time.sleep()更可靠,能避免网络延迟导致的元素未加载问题
  • 确保浏览器驱动(geckodriver/chromedriver)版本与浏览器版本匹配

内容的提问来源于stack exchange,提问作者Newbier_RP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 13:55:38