使用Python Selenium获取Frame中嵌入PDF的源URL失败问题求助
Hey there! Let's work through this problem together—you're super close, just a few tweaks needed to get that PDF URL successfully.
First, Let's Diagnose the Error
The NoSuchFrameException happens because your code is trying to switch to the mBottomFrame before it's fully loaded on the page. time.sleep(2) is a guess, but sometimes pages take longer to load frames, especially with dynamic content like generated PDFs. We'll replace that with a smarter wait that waits until the frame exists before trying to interact with it.
Also, Quick Fixes in Your Existing Code
Before we get to the frame issue, let's fix two small but critical problems:
- URLs need to be full:
driver.get("www.xxxxxx.com")should includehttps://(e.g.,driver.get("https://www.xxxxxx.com"))—otherwise, Selenium will treat it as a relative path and fail to load the site. - Path escaping: In Python, backslashes in file paths need to be escaped. Change
PATH = "C:\Program Files (x86)\chromedriver.exe"toPATH = r"C:\Program Files (x86)\chromedriver.exe"(thermakes it a raw string, avoiding escape issues).
Step-by-Step Solution Code
Here's your updated code with fixes and proper waits. We'll cover both getting the frame's src (which is the PDF URL you saw in the frameset) and the embed's src if you need that too:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time # Fix path with raw string PATH = r"C:\Program Files (x86)\chromedriver.exe" driver = webdriver.Chrome(PATH) myUsername="xxxx" myPassword="xxxx" # Fix URL with full https:// driver.get("https://www.xxxxxx.com") # Login flow driver.find_element(By.XPATH, "//*[@id='tbUserName']").send_keys(myUsername) driver.find_element(By.XPATH, "//*[@id='tbPassword']").send_keys(myPassword) driver.find_element(By.XPATH, "//*[@id='ctl00_cp_Content_spLogin']").click() # Use explicit wait instead of time.sleep for reliability wait = WebDriverWait(driver, 10) # Wait up to 10 seconds for elements # Select report driver.get("https://www.xxxxxx.com") wait.until(EC.element_to_be_clickable((By.XPATH, "//*[@id='Repeater_ReportCategory_ctl00_LinkButton_ReportCategory']"))).click() wait.until(EC.element_to_be_clickable((By.XPATH, "//*[@id='Repeater_AdditionalReports_ctl06_LinkButton_AdditionalReportName']"))).click() # Sort options wait.until(EC.element_to_be_clickable((By.XPATH, "//*[@id='RadioButton_WordCountSortByWCHighToLow']"))).click() wait.until(EC.element_to_be_clickable((By.XPATH, "//*[@id='mButton_Next']"))).click() # Get the PDF URL from the mBottomFrame # Wait for the frame to exist first bottom_frame = wait.until(EC.presence_of_element_located((By.ID, "mBottomFrame"))) # Get the frame's src BEFORE switching into it (once inside, you can't access the frame tag itself) frame_pdf_url = bottom_frame.get_attribute("src") print("Frame PDF URL:", frame_pdf_url) # If you also need the embed's src (from the <embed> tag inside the frame): driver.switch_to.frame(bottom_frame) # Switch into the frame context embed_element = wait.until(EC.presence_of_element_located((By.ID, "plugin"))) embed_pdf_url = embed_element.get_attribute("src") print("Embed PDF URL:", embed_pdf_url) # Switch back to the main page context if you need to do more actions later driver.switch_to.default_content() # Cleanup driver.quit()
Key Changes Explained
- Explicit Waits:
WebDriverWaitwaits until the element is ready (either present or clickable) instead of using fixed sleep times. This makes your code way more reliable, especially on slow-loading pages. - Get Frame Src Before Switching: Once you switch into a frame, your Selenium driver's context is inside that frame's HTML—you can't access the
<frame>tag itself anymore. So we grab thesrcfirst while we're still in the main page context. - Accessing the Embed: If you need the URL from the
<embed>tag inside the frame, we switch into the frame first, then wait for the embed element to load before grabbing itssrc.
Why This Works
Your original code failed because it tried to switch to the frame too early (before it loaded) and used an incorrect way to get the frame's src (chaining switch_to.frame().get_attribute() doesn't work because switch_to.frame() returns a SwitchTo object, not the frame element).
This code fixes both issues and gives you both possible PDF URLs you mentioned in your question.
内容的提问来源于stack exchange,提问作者Aristotle

