You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python技术咨询:列表非整数元素遍历与爬虫代码相关问题

Hey there! Let's walk through how to handle this for your Python web scraping project using BeautifulSoup and Selenium.

Understanding Your Current Setup First

First, let's recap what you've already built:

  • You're using driver.find_elements_by_link_text('Fall/Winter 2017-2018 Course Schedule') to grab all matching link elements (this returns a list of Selenium WebElement objects, not integers)
  • You're checking the number of links with len(link_click)
  • You've defined a get_course_info() function to extract details from the course page
Solution: Traversing Non-Integer Elements in Your List

Since the link_click list contains WebElement objects (not integers), you can iterate over them directly with standard Python loops. Here's how to integrate this with your existing workflow:

Step 1: Loop Through the WebElement List

Instead of just counting links, loop through each element to interact with it and call your course info function:

# Your existing code to retrieve links
link_click = driver.find_elements_by_link_text('Fall/Winter 2017-2018 Course Schedule')
print(f"Found {len(link_click)} course schedule links")

# Iterate over each non-integer (WebElement) in the list
for link in link_click:
    # Click the link to navigate to the course page
    link.click()
    
    # Call your function to extract course information
    get_course_info()
    
    # Navigate back to the original page to process the next link
    driver.back()

Step 2: Refine Your get_course_info() Function (Example)

Make sure your function can parse the page correctly. Here's an enhanced version that combines Selenium and BeautifulSoup:

from bs4 import BeautifulSoup

def get_course_info():
    current_url = driver.current_url
    print(f"Extracting data from: {current_url}")
    
    # Use BeautifulSoup to parse the current page's source
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    
    # Add your custom parsing logic here (adjust selectors to match the site)
    course_cards = soup.find_all('div', class_='course-card')
    for card in course_cards:
        course_code = card.find('span', class_='course-code').text.strip()
        course_name = card.find('h2', class_='course-name').text.strip()
        print(f"Course: {course_code} - {course_name}")

Pro Tips to Avoid Common Issues

  • Add Explicit Waits: Links might not load immediately, so use waits to prevent errors:
    from selenium.webdriver.common.by import By
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    
    # Wait up to 10 seconds for links to become available
    link_click = WebDriverWait(driver, 10).until(
        EC.presence_of_all_elements_located((By.LINK_TEXT, 'Fall/Winter 2017-2018 Course Schedule'))
    )
    
  • Handle Stale Elements: If navigating back causes elements to become "stale" (unreachable), re-fetch the link list after returning to the original page.

内容的提问来源于stack exchange,提问作者express_v2

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:52:43