Selenium多窗口爬取异常求助:窗口切换与链接遍历问题
Fixing Your Selenium Multi-Window Scraping Issues
Hey there, let's work through those two key issues you're hitting with your ScienceDirect scraping script—we'll get this sorted out!
First, let's break down what's going wrong:
- Wrong page data after window switch: Your
get_data()function wasn't receiving the specificlinkyou wanted to process, so it was reusing a stale reference. Plus, you weren't waiting for the new window to fully load before scraping, which might have led to accidentally pulling data from the original window. - All links opening at once: Without passing
linkas a parameter toget_data(), the function ended up referencing the last link in your list by the time the loop ran, and missing proper waits caused batch window openings.
Here's the revised code with fixes:
from selenium import webdriver from selenium.webdriver import ActionChains from selenium.webdriver.common.by import By from selenium.webdriver.common.keys import Keys from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException, TimeoutException driver = webdriver.Chrome(executable_path='C:/chromedriver.exe') actions = ActionChains(driver) search_term = input("Enter your search term :") url = f'https://www.sciencedirect.com/search?qs={search_term}&years=2021%2C2020%2C2019&lastSelectedFacet=years' driver.get(url) driver.maximize_window() # Handle cookie popup with fallback try: WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.XPATH,'/html/body/div[3]/div/div/div/button/span'))).click() except TimeoutException: print("Cookie popup not found, continuing...") # Collect all result links divs = driver.find_elements(By.CLASS_NAME, 'result-item-content') links = [] for div in divs: link = div.find_element(By.TAG_NAME, 'a') links.append(link) def get_data(link): # Store original window state original_window = driver.current_window_handle original_window_count = len(driver.window_handles) # Open link in new tab actions.key_down(Keys.CONTROL).click(link).key_up(Keys.CONTROL).perform() # Wait for new window to load try: WebDriverWait(driver, 10).until(lambda d: len(d.window_handles) > original_window_count) except TimeoutException: print(f"Failed to open new window for link: {link.get_attribute('href')}") return # Switch to the new window for window_handle in driver.window_handles: if window_handle != original_window: driver.switch_to.window(window_handle) break # Wait for target content to load in new tab try: WebDriverWait(driver, 15).until(EC.presence_of_element_located((By.ID, 'author-group'))) except TimeoutException: print("Author section not found in new tab, closing...") driver.close() driver.switch_to.window(original_window) return # Scrape author data author_group = driver.find_element(By.ID, 'author-group') for author in author_group.find_elements(By.CSS_SELECTOR, "a.author"): try: given_name = author.find_element(By.CSS_SELECTOR, ".given-name").text surname = author.find_element(By.CSS_SELECTOR, ".surname").text except NoSuchElementException: print("Could not extract first or last name") continue try: mail_icon = author.find_element(By.CSS_SELECTOR, ".icon-envelope") mail_icon.click() # Wait for email to render mail = WebDriverWait(driver, 5).until(EC.presence_of_element_located((By.CLASS_NAME, 'e-address'))) print(f"{given_name} {surname}: {mail.text}") except NoSuchElementException: print(f"Author {given_name} {surname}: Mail icon not present") # Clean up and switch back driver.close() driver.switch_to.window(original_window) # Process links one at a time for link in links: get_data(link) driver.quit()
Key fixes explained:
- Pass
linktoget_data(): Ensures each iteration processes the correct link instead of reusing a stale reference. - Wait for new window to open: Uses
WebDriverWaitto confirm the window count increases, so we don't switch tabs prematurely. - Wait for page elements: Added waits for the author section and email address, preventing attempts to scrape before the page is ready.
- Robust error handling: Added fallback for missing popups or elements to keep the script running smoothly.
- Modern locator syntax: Updated deprecated
find_elements_by_*methods to the currentfind_elements(By.*, ...)format.
Why this works:
- No more wrong-page scraping: By waiting for the new tab's content to load and explicitly switching to the correct window handle, we guarantee all scraping happens in the right tab. Closing the new tab and switching back keeps your main search page intact.
- Links open sequentially: Each link is processed fully (scraped, tab closed, switched back) before moving to the next—no more batch openings.
内容的提问来源于stack exchange,提问作者Abhishek Rai
相关产品推荐
相关产品推荐

