You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Selenium获取Frame中嵌入PDF的源URL失败问题求助

Fixing the NoSuchFrameException & Getting Embedded PDF URLs with Selenium

Hey there! Let's work through this problem together—you're super close, just a few tweaks needed to get that PDF URL successfully.

First, Let's Diagnose the Error

The NoSuchFrameException happens because your code is trying to switch to the mBottomFrame before it's fully loaded on the page. time.sleep(2) is a guess, but sometimes pages take longer to load frames, especially with dynamic content like generated PDFs. We'll replace that with a smarter wait that waits until the frame exists before trying to interact with it.

Also, Quick Fixes in Your Existing Code

Before we get to the frame issue, let's fix two small but critical problems:

  • URLs need to be full: driver.get("www.xxxxxx.com") should include https:// (e.g., driver.get("https://www.xxxxxx.com"))—otherwise, Selenium will treat it as a relative path and fail to load the site.
  • Path escaping: In Python, backslashes in file paths need to be escaped. Change PATH = "C:\Program Files (x86)\chromedriver.exe" to PATH = r"C:\Program Files (x86)\chromedriver.exe" (the r makes it a raw string, avoiding escape issues).

Step-by-Step Solution Code

Here's your updated code with fixes and proper waits. We'll cover both getting the frame's src (which is the PDF URL you saw in the frameset) and the embed's src if you need that too:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

# Fix path with raw string
PATH = r"C:\Program Files (x86)\chromedriver.exe"
driver = webdriver.Chrome(PATH)
myUsername="xxxx"
myPassword="xxxx"

# Fix URL with full https://
driver.get("https://www.xxxxxx.com")

# Login flow
driver.find_element(By.XPATH, "//*[@id='tbUserName']").send_keys(myUsername)
driver.find_element(By.XPATH, "//*[@id='tbPassword']").send_keys(myPassword)
driver.find_element(By.XPATH, "//*[@id='ctl00_cp_Content_spLogin']").click()

# Use explicit wait instead of time.sleep for reliability
wait = WebDriverWait(driver, 10) # Wait up to 10 seconds for elements

# Select report
driver.get("https://www.xxxxxx.com")
wait.until(EC.element_to_be_clickable((By.XPATH, "//*[@id='Repeater_ReportCategory_ctl00_LinkButton_ReportCategory']"))).click()
wait.until(EC.element_to_be_clickable((By.XPATH, "//*[@id='Repeater_AdditionalReports_ctl06_LinkButton_AdditionalReportName']"))).click()

# Sort options
wait.until(EC.element_to_be_clickable((By.XPATH, "//*[@id='RadioButton_WordCountSortByWCHighToLow']"))).click()
wait.until(EC.element_to_be_clickable((By.XPATH, "//*[@id='mButton_Next']"))).click()

# Get the PDF URL from the mBottomFrame
# Wait for the frame to exist first
bottom_frame = wait.until(EC.presence_of_element_located((By.ID, "mBottomFrame")))
# Get the frame's src BEFORE switching into it (once inside, you can't access the frame tag itself)
frame_pdf_url = bottom_frame.get_attribute("src")
print("Frame PDF URL:", frame_pdf_url)

# If you also need the embed's src (from the <embed> tag inside the frame):
driver.switch_to.frame(bottom_frame) # Switch into the frame context
embed_element = wait.until(EC.presence_of_element_located((By.ID, "plugin")))
embed_pdf_url = embed_element.get_attribute("src")
print("Embed PDF URL:", embed_pdf_url)

# Switch back to the main page context if you need to do more actions later
driver.switch_to.default_content()

# Cleanup
driver.quit()

Key Changes Explained

  1. Explicit Waits: WebDriverWait waits until the element is ready (either present or clickable) instead of using fixed sleep times. This makes your code way more reliable, especially on slow-loading pages.
  2. Get Frame Src Before Switching: Once you switch into a frame, your Selenium driver's context is inside that frame's HTML—you can't access the <frame> tag itself anymore. So we grab the src first while we're still in the main page context.
  3. Accessing the Embed: If you need the URL from the <embed> tag inside the frame, we switch into the frame first, then wait for the embed element to load before grabbing its src.

Why This Works

Your original code failed because it tried to switch to the frame too early (before it loaded) and used an incorrect way to get the frame's src (chaining switch_to.frame().get_attribute() doesn't work because switch_to.frame() returns a SwitchTo object, not the frame element).

This code fixes both issues and gives you both possible PDF URLs you mentioned in your question.

内容的提问来源于stack exchange,提问作者Aristotle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 08:17:44