Python技术咨询:列表非整数元素遍历与爬虫代码相关问题
Hey there! Let's walk through how to handle this for your Python web scraping project using BeautifulSoup and Selenium.
Understanding Your Current Setup First
First, let's recap what you've already built:
- You're using
driver.find_elements_by_link_text('Fall/Winter 2017-2018 Course Schedule')to grab all matching link elements (this returns a list of Selenium WebElement objects, not integers) - You're checking the number of links with
len(link_click) - You've defined a
get_course_info()function to extract details from the course page
Solution: Traversing Non-Integer Elements in Your List
Since the link_click list contains WebElement objects (not integers), you can iterate over them directly with standard Python loops. Here's how to integrate this with your existing workflow:
Step 1: Loop Through the WebElement List
Instead of just counting links, loop through each element to interact with it and call your course info function:
# Your existing code to retrieve links link_click = driver.find_elements_by_link_text('Fall/Winter 2017-2018 Course Schedule') print(f"Found {len(link_click)} course schedule links") # Iterate over each non-integer (WebElement) in the list for link in link_click: # Click the link to navigate to the course page link.click() # Call your function to extract course information get_course_info() # Navigate back to the original page to process the next link driver.back()
Step 2: Refine Your get_course_info() Function (Example)
Make sure your function can parse the page correctly. Here's an enhanced version that combines Selenium and BeautifulSoup:
from bs4 import BeautifulSoup def get_course_info(): current_url = driver.current_url print(f"Extracting data from: {current_url}") # Use BeautifulSoup to parse the current page's source soup = BeautifulSoup(driver.page_source, 'html.parser') # Add your custom parsing logic here (adjust selectors to match the site) course_cards = soup.find_all('div', class_='course-card') for card in course_cards: course_code = card.find('span', class_='course-code').text.strip() course_name = card.find('h2', class_='course-name').text.strip() print(f"Course: {course_code} - {course_name}")
Pro Tips to Avoid Common Issues
- Add Explicit Waits: Links might not load immediately, so use waits to prevent errors:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Wait up to 10 seconds for links to become available link_click = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.LINK_TEXT, 'Fall/Winter 2017-2018 Course Schedule')) ) - Handle Stale Elements: If navigating back causes elements to become "stale" (unreachable), re-fetch the link list after returning to the original page.
内容的提问来源于stack exchange,提问作者express_v2
相关产品推荐
相关产品推荐

