批量自动打开发票门户并下载PDF的技术实现问询
批量自动下载发票PDF的实现方案
完全可以实现这个批量操作,我们可以用Web自动化工具结合Python完成整个流程,以下是具体实现方案:
核心思路
通过自动化浏览器模拟人工操作流程:
- 循环遍历发票编号列表,构造对应门户URL并打开页面
- 定位并点击"See PDF"按钮触发PDF下载
- 从页面提取正式发票编号,重命名下载的PDF文件到指定文件夹
工具选择
推荐使用Selenium(Python版):它是成熟的Web自动化工具,支持主流浏览器,能轻松模拟点击、页面元素定位等操作;也可以选择Playwright,它的自动等待机制更适配React这类动态渲染页面,代码逻辑更简洁。
具体实现步骤(以Selenium为例)
1. 环境准备
- 安装Selenium:
pip install selenium - 下载对应浏览器的驱动(比如ChromeDriver),确保驱动版本与浏览器版本一致,将驱动放在Python可执行路径或指定路径
2. 代码实现
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import os import time def setup_chrome_driver(download_folder): """配置Chrome浏览器,设置自动下载PDF到指定文件夹""" chrome_options = webdriver.ChromeOptions() prefs = { "download.default_directory": os.path.abspath(download_folder), "download.prompt_for_download": False, # 关闭下载确认弹窗 "download.directory_upgrade": True, "plugins.always_open_pdf_externally": True # 直接下载PDF,不打开预览 } chrome_options.add_experimental_option("prefs", prefs) # 可选:启用无头模式,不显示浏览器窗口 # chrome_options.add_argument("--headless=new") driver = webdriver.Chrome(options=chrome_options) return driver def download_invoice_pdfs(invoice_numbers, download_folder): """批量下载发票PDF并按编号命名""" # 创建目标文件夹(不存在则自动创建) os.makedirs(download_folder, exist_ok=True) driver = setup_chrome_driver(download_folder) for invoice_num in invoice_numbers: try: # 构造发票门户URL portal_url = f"https://invoice_portal/statements/{invoice_num}" driver.get(portal_url) # 等待"See PDF"按钮加载完成并点击 see_pdf_btn = WebDriverWait(driver, 15).until( EC.element_to_be_clickable((By.CLASS_NAME, "_1iKuo")) ) see_pdf_btn.click() # 等待发票编号元素加载,提取正式编号 invoice_title_elem = WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.CLASS_NAME, "_35Wsy")) ) actual_invoice_num = invoice_title_elem.text.split("#")[-1].strip() # 等待下载完成(可根据网络情况调整等待时间,或用文件监听优化) time.sleep(6) # 获取下载文件夹中最新的PDF文件 download_files = [os.path.join(download_folder, f) for f in os.listdir(download_folder) if f.endswith(".pdf")] latest_pdf = max(download_files, key=os.path.getctime) # 重命名文件为发票编号 target_path = os.path.join(download_folder, f"{actual_invoice_num}.pdf") os.rename(latest_pdf, target_path) print(f"已完成:{actual_invoice_num}.pdf") except Exception as e: print(f"处理发票{invoice_num}失败:{str(e)}") continue driver.quit() # 调用示例 if __name__ == "__main__": # 替换为你的发票编号列表 invoice_list = ["INV-225", "INV-226", "INV-227"] # 替换为你要保存的文件夹路径 save_directory = "./invoice_pdfs" download_invoice_pdfs(invoice_list, save_directory)
关键细节说明
- 浏览器配置:通过
plugins.always_open_pdf_externally设置直接下载PDF,避免打开预览窗口;download.prompt_for_download关闭下载确认弹窗 - 元素定位:使用
WebDriverWait等待元素加载完成,避免因React动态渲染导致的元素未找到问题 - 下载等待:示例中用
time.sleep简单等待,若要更精准,可监听下载文件夹的文件变化,或通过浏览器日志捕获PDF的下载URL,直接用requests库下载(更高效) - 异常处理:捕获单个发票处理的异常,避免一个失败导致整个批量任务中断
替代方案:Playwright
如果页面是复杂的React应用,Playwright的自动等待机制更适配,代码会更简洁,示例片段:
from playwright.sync_api import sync_playwright import os def download_with_playwright(invoice_numbers, download_folder): os.makedirs(download_folder, exist_ok=True) with sync_playwright() as p: browser = p.chromium.launch(headless=False) context = browser.new_context( accept_downloads=True, downloads_path=download_folder ) page = context.new_page() for invoice_num in invoice_numbers: try: page.goto(f"https://invoice_portal/statements/{invoice_num}") # 等待按钮并点击 page.click("button._1iKuo") # 等待下载完成 download = page.wait_for_download() # 提取发票编号 invoice_title = page.locator("h5._35Wsy").text_content() actual_invoice_num = invoice_title.split("#")[-1].strip() # 重命名保存文件 download.save_as(os.path.join(download_folder, f"{actual_invoice_num}.pdf")) print(f"已完成:{actual_invoice_num}.pdf") except Exception as e: print(f"处理发票{invoice_num}失败:{str(e)}") continue browser.close()
内容的提问来源于stack exchange,提问作者user14070248
相关产品推荐
相关产品推荐

