You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium多窗口爬取异常求助:窗口切换与链接遍历问题

Fixing Your Selenium Multi-Window Scraping Issues

Hey there, let's work through those two key issues you're hitting with your ScienceDirect scraping script—we'll get this sorted out!

First, let's break down what's going wrong:

  • Wrong page data after window switch: Your get_data() function wasn't receiving the specific link you wanted to process, so it was reusing a stale reference. Plus, you weren't waiting for the new window to fully load before scraping, which might have led to accidentally pulling data from the original window.
  • All links opening at once: Without passing link as a parameter to get_data(), the function ended up referencing the last link in your list by the time the loop ran, and missing proper waits caused batch window openings.

Here's the revised code with fixes:

from selenium import webdriver
from selenium.webdriver import ActionChains
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import NoSuchElementException, TimeoutException

driver = webdriver.Chrome(executable_path='C:/chromedriver.exe')
actions = ActionChains(driver)
search_term = input("Enter your search term :")
url = f'https://www.sciencedirect.com/search?qs={search_term}&years=2021%2C2020%2C2019&lastSelectedFacet=years'
driver.get(url)
driver.maximize_window()

# Handle cookie popup with fallback
try:
    WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.XPATH,'/html/body/div[3]/div/div/div/button/span'))).click()
except TimeoutException:
    print("Cookie popup not found, continuing...")

# Collect all result links
divs = driver.find_elements(By.CLASS_NAME, 'result-item-content')
links = []
for div in divs:
    link = div.find_element(By.TAG_NAME, 'a')
    links.append(link)

def get_data(link):
    # Store original window state
    original_window = driver.current_window_handle
    original_window_count = len(driver.window_handles)
    
    # Open link in new tab
    actions.key_down(Keys.CONTROL).click(link).key_up(Keys.CONTROL).perform()
    
    # Wait for new window to load
    try:
        WebDriverWait(driver, 10).until(lambda d: len(d.window_handles) > original_window_count)
    except TimeoutException:
        print(f"Failed to open new window for link: {link.get_attribute('href')}")
        return
    
    # Switch to the new window
    for window_handle in driver.window_handles:
        if window_handle != original_window:
            driver.switch_to.window(window_handle)
            break
    
    # Wait for target content to load in new tab
    try:
        WebDriverWait(driver, 15).until(EC.presence_of_element_located((By.ID, 'author-group')))
    except TimeoutException:
        print("Author section not found in new tab, closing...")
        driver.close()
        driver.switch_to.window(original_window)
        return
    
    # Scrape author data
    author_group = driver.find_element(By.ID, 'author-group')
    for author in author_group.find_elements(By.CSS_SELECTOR, "a.author"):
        try:
            given_name = author.find_element(By.CSS_SELECTOR, ".given-name").text
            surname = author.find_element(By.CSS_SELECTOR, ".surname").text
        except NoSuchElementException:
            print("Could not extract first or last name")
            continue
        
        try:
            mail_icon = author.find_element(By.CSS_SELECTOR, ".icon-envelope")
            mail_icon.click()
            # Wait for email to render
            mail = WebDriverWait(driver, 5).until(EC.presence_of_element_located((By.CLASS_NAME, 'e-address')))
            print(f"{given_name} {surname}: {mail.text}")
        except NoSuchElementException:
            print(f"Author {given_name} {surname}: Mail icon not present")
    
    # Clean up and switch back
    driver.close()
    driver.switch_to.window(original_window)

# Process links one at a time
for link in links:
    get_data(link)

driver.quit()

Key fixes explained:

  • Pass link to get_data(): Ensures each iteration processes the correct link instead of reusing a stale reference.
  • Wait for new window to open: Uses WebDriverWait to confirm the window count increases, so we don't switch tabs prematurely.
  • Wait for page elements: Added waits for the author section and email address, preventing attempts to scrape before the page is ready.
  • Robust error handling: Added fallback for missing popups or elements to keep the script running smoothly.
  • Modern locator syntax: Updated deprecated find_elements_by_* methods to the current find_elements(By.*, ...) format.

Why this works:

  1. No more wrong-page scraping: By waiting for the new tab's content to load and explicitly switching to the correct window handle, we guarantee all scraping happens in the right tab. Closing the new tab and switching back keeps your main search page intact.
  2. Links open sequentially: Each link is processed fully (scraped, tab closed, switched back) before moving to the next—no more batch openings.

内容的提问来源于stack exchange,提问作者Abhishek Rai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:19:43