You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python和Selenium在pdfcompressor.com上传文件失败求助

解决Selenium上传PDF到pdfcompressor.com失败的问题

你遇到的核心问题是用动态生成的ID定位上传元素——那个html5_1cciqvn90sehe7rachs1c3m03类型的ID是页面每次加载时随机生成的,下次刷新页面ID就变了,自然定位不到元素。

给你几个可靠的解决方案:

方案1:通过type="file"属性定位上传输入框

上传文件的input标签几乎都会带有type="file"属性,这是静态不变的,用它来定位最稳妥。加上显式等待确保元素加载完成,代码如下:

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

file_path = "/home/gugli/Documents/script_py/Dainik_Jagron/h2.pdf"
browser = webdriver.Firefox()
url = 'http://pdfcompressor.com/'
browser.get(url)

# 显式等待10秒,直到上传input出现
upload_input = WebDriverWait(browser, 10).until(
    EC.presence_of_element_located((By.XPATH, "//input[@type='file']"))
)
# 发送文件路径
upload_input.send_keys(file_path)

方案2:结合父元素缩小定位范围

如果页面有多个type="file"的input,可以观察上传区域的父容器类名,进一步缩小定位范围,示例:

# 假设上传input在class为"upload-container"的父容器下
upload_input = WebDriverWait(browser, 10).until(
    EC.presence_of_element_located((By.CSS_SELECTOR, ".upload-container input[type='file']"))
)
upload_input.send_keys(file_path)

关键注意点

  • 放弃动态ID定位:这类ID是前端框架自动生成的,每次加载都会变化,完全不可靠。
  • 必须用显式等待:直接调用find_element_*可能会因为页面未加载完成导致元素找不到,显式等待能有效避免这种时序问题。

内容的提问来源于stack exchange,提问作者Prince

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:42:37