使用Python抓取JSP网站中的表格数据
抓取指定JSP页面表格数据的实现
我要抓取https://www.eprocure.gov.bd/resources/common/SearcheCMS.jsp页面上的表格,参考相关示例写了以下代码:
from selenium import webdriver from selenium.webdriver.firefox.options import Options import time from bs4 import BeautifulSoup import pandas as pd options = Options() options.add_argument('--headless') driver = webdriver.Firefox(executable_path="C:/Users/DefaultUser/AppData/geckodriver.exe") driver.get("https://www.eprocure.gov.bd/resources/common/SearcheCMS.jsp") time.sleep(5) res = driver.execute_script("return document.documentElement.outerHTML") driver.quit() soup = BeautifulSoup(res, 'html.parser') table_rows = soup.find_all('table')[1].find_all('tr') rows = [] for tr in table_rows: td = tr.find_all('td') rows.append([i.text.strip() for i in td]) # 去除文本首尾空白 delaydata = rows[3:] df = pd.DataFrame(delaydata, columns = [ '序号', '部委、部门、采购执行机构', '采购性质、类型及方式', '招标/提案编号、参考号、标题及发布日期', '中标单位', '企业唯一ID', '经验证书编号', '合同金额', '合同起止日期', '工作状态' ]) print(df)
补充说明
- 补上了原代码缺失的导入语句,避免运行报错
- 提取文本时增加
strip()处理,去除多余空格和换行,让数据更整洁 - 固定等待
time.sleep(5)可替换为Selenium显式等待(比如等待表格元素加载完成),能提升代码稳定性和效率
内容的提问来源于stack exchange,提问作者Nazmul
相关产品推荐
相关产品推荐

