Python新手求助:如何爬取aspx.net网站点击后生成的表格?
解决ASP.NET网站表格爬取问题(Selenium实现)
我是Python新手,尝试从网站https://tradereport.moc.go.th/Report/ReportEng.aspx?Report=HarmonizeCommodity&Lang=Eng&ImExType=1&Option=1下载表格。该网站基于ASP.NET,难点在于页面源码里看不到生成报告所需点击的按钮。我先做简化测试,只点击“ReviewReport”按钮获取结果表格,网页能显示表格,但无法从页面源码中爬取到数据。以下是我写的Selenium代码,请求帮忙解决表格爬取问题:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.support.ui import Select from selenium.common.exceptions import NoSuchElementException from selenium.webdriver.common.keys import Keys from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.keys import Keys import pandas as pd Year = ('2021') Month = ('11') HScode = ('721990') options = Options() #options.headless = True #driver = webdriver.Chrome('C:\\Python\Scripts\chromedriver.exe', options=options) driver = webdriver.Firefox(executable_path=r'C:\Python\geckodriver.exe') driver.get("https://tradereport.moc.go.th/Report/ReportEng.aspx? Report=HarmonizeCommodity&Lang=Eng&ImExType=1&Option=1") time.sleep(5) driver.get("https://tradereport.moc.go.th/Report/ReportEng.aspx? Report=HarmonizeCommodity&Lang=Eng&ImExType=1&Option=1") #reenter the url again to get the table page time.sleep(5) driver.find_element_by_id("ddlYear").send_keys(Year) #Working fine driver.find_element_by_id("ddlMonth").send_keys(Month) #Working fine driver.find_element_by_id("txtHsCode").send_keys(HScode) #Working fine submitbttn = driver.find_element_by_id("btnSubmit") #Working fine submitbttn.click() time.sleep(5) f=open("d:\\page.txt","w") f.write(driver.page_source) f.close() print ("************* Scrapping data done**********************") driver.quit`
问题分析
页面显示的表格大概率是动态加载在iframe中,或是通过AJAX异步渲染的,直接获取driver.page_source无法拿到表格数据。另外代码存在几个小问题:
- 重复调用
driver.get()完全没必要,首次加载页面后无需重复访问 - 未导入
time模块,执行time.sleep()会报错 driver.quit缺少括号,正确写法是driver.quit()
解决方案步骤
1. 检查并切换到iframe
很多ASP.NET网站会把报表放在iframe里,先切换到iframe再操作:
# 点击提交后等待iframe加载完成并切换 WebDriverWait(driver, 10).until(EC.frame_to_be_available_and_switch_to_it((By.TAG_NAME, "iframe")))
2. 用显式等待定位表格并提取数据
改用显式等待定位表格元素,再通过Pandas直接读取表格内容:
# 等待表格加载完成,可替换为表格实际ID或其他定位器 table = WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.CSS_SELECTOR, "table")) ) # Pandas读取表格数据 df = pd.read_html(table.get_attribute('outerHTML'))[0] # 保存为CSV文件 df.to_csv("d:\\trade_report.csv", index=False)
3. 修正代码错误
- 新增
import time导入模块 - 删除重复的
driver.get()调用 - 将
driver.quit改为driver.quit()
完整修正代码示例
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.support.ui import Select from selenium.common.exceptions import NoSuchElementException from selenium.webdriver.common.keys import Keys from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.keys import Keys import pandas as pd import time Year = '2021' Month = '11' HScode = '721990' options = Options() # options.headless = True # driver = webdriver.Chrome('C:\\Python\Scripts\chromedriver.exe', options=options) driver = webdriver.Firefox(executable_path=r'C:\Python\geckodriver.exe') driver.get("https://tradereport.moc.go.th/Report/ReportEng.aspx?Report=HarmonizeCommodity&Lang=Eng&ImExType=1&Option=1") time.sleep(5) # 填写表单 driver.find_element_by_id("ddlYear").send_keys(Year) driver.find_element_by_id("ddlMonth").send_keys(Month) driver.find_element_by_id("txtHsCode").send_keys(HScode) submitbttn = driver.find_element_by_id("btnSubmit") submitbttn.click() # 尝试切换到iframe(表格可能在其中) try: WebDriverWait(driver, 10).until(EC.frame_to_be_available_and_switch_to_it((By.TAG_NAME, "iframe"))) except: pass # 等待表格加载并提取数据 table = WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.CSS_SELECTOR, "table")) ) df = pd.read_html(table.get_attribute('outerHTML'))[0] print(df.head()) df.to_csv("d:\\trade_report.csv", index=False) # 保存页面源码(可选) with open("d:\\page.txt","w", encoding='utf-8') as f: f.write(driver.page_source) print("************* 数据爬取完成**********************") driver.quit()
注意事项
- 如果不知道表格ID,可用
By.CSS_SELECTOR, "table"或By.TAG_NAME, "table"定位所有表格,再从Pandas返回的列表中选择正确表格 - 显式等待
WebDriverWait比time.sleep()更可靠,能避免网络延迟导致的元素未加载问题 - 确保浏览器驱动(geckodriver/chromedriver)版本与浏览器版本匹配
内容的提问来源于stack exchange,提问作者Newbier_RP
相关产品推荐
相关产品推荐

