You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python抓取JSP网站中的表格数据

抓取指定JSP页面表格数据的实现

我要抓取https://www.eprocure.gov.bd/resources/common/SearcheCMS.jsp页面上的表格,参考相关示例写了以下代码:

from selenium import webdriver
from selenium.webdriver.firefox.options import Options
import time
from bs4 import BeautifulSoup
import pandas as pd

options = Options()
options.add_argument('--headless')

driver = webdriver.Firefox(executable_path="C:/Users/DefaultUser/AppData/geckodriver.exe")
driver.get("https://www.eprocure.gov.bd/resources/common/SearcheCMS.jsp")
time.sleep(5)
res = driver.execute_script("return document.documentElement.outerHTML")
driver.quit()

soup = BeautifulSoup(res, 'html.parser')
table_rows = soup.find_all('table')[1].find_all('tr')
rows = []
for tr in table_rows:
    td = tr.find_all('td')
    rows.append([i.text.strip() for i in td])  # 去除文本首尾空白
delaydata = rows[3:]
df = pd.DataFrame(delaydata, columns = [
    '序号', 
    '部委、部门、采购执行机构', 
    '采购性质、类型及方式', 
    '招标/提案编号、参考号、标题及发布日期', 
    '中标单位', 
    '企业唯一ID', 
    '经验证书编号', 
    '合同金额', 
    '合同起止日期', 
    '工作状态'
])
print(df)

补充说明

  • 补上了原代码缺失的导入语句,避免运行报错
  • 提取文本时增加strip()处理,去除多余空格和换行,让数据更整洁
  • 固定等待time.sleep(5)可替换为Selenium显式等待(比如等待表格元素加载完成),能提升代码稳定性和效率

内容的提问来源于stack exchange,提问作者Nazmul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 19:35:24