爬虫新手使用Selenium获取课程名时遭遇CSS选择器无效选择器异常
Fixing the Invalid Selector Exception in Your Selenium Web Scraper
Hey there! As someone who’s been in your shoes as a new web scraper, let’s walk through how to fix that Invalid Selector Exception you’re hitting. First, let’s break down what’s likely going wrong and adjust your code to get it working.
What’s Causing the Exception?
From your code snippet, it looks like you started looping through odd rows but didn’t add the actual CSS selector logic to grab course names. This error usually pops up for two main reasons:
- Your CSS selector has syntax mistakes (like typos or unsupported syntax)
- You’re trying to select an element that either doesn’t exist, or hasn’t finished loading when your code runs
Updated, Working Code
Let’s rewrite your script with proper element targeting and error handling, tailored to the UIUC CS course page structure:
from selenium import webdriver from selenium.webdriver.firefox.firefox_binary import FirefoxBinary from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException # Initialize Firefox driver binary = FirefoxBinary(r'C:\Program Files\Mozilla Firefox\firefox.exe') driver = webdriver.Firefox(firefox_binary=binary) url = 'https://courses.illinois.edu/schedule/2018/fall/CS' driver.get(url) course_names = [] # Wait for the schedule table to load (prevents race conditions) wait = WebDriverWait(driver, 10) schedule_table = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'table.schedule-table'))) # Grab all odd rows (the page uses alternating rows for course details) course_rows = schedule_table.find_elements_by_css_selector('tr:nth-child(odd)') for row in course_rows: try: # Target the course name (it lives in the 2nd <td> of each row, inside an <a> tag) course_name = row.find_element_by_css_selector('td:nth-child(2) a').text course_names.append(course_name) except NoSuchElementException: print(f"Skipping row - couldn't find the course name element") # Print results and clean up print("Scraped course names:") for name in course_names: print(f"- {name}") driver.quit()
Key Fixes & Improvements
- Dynamic Row Targeting: Instead of hardcoding a range of 69 rows, we dynamically grab all odd rows from the loaded table. This is way more resilient if the page updates with more or fewer courses.
- Explicit Waits: We use
WebDriverWaitto make sure the table loads completely before scraping. This eliminates errors from trying to select elements that aren’t yet rendered on the page. - Valid CSS Selectors: The selectors
table.schedule-tableandtr:nth-child(odd)correctly target the elements on the UIUC page with no syntax errors. We then drill down to the 2nd table cell where the course name is stored. - Error Handling: We catch
NoSuchElementExceptionto handle any unexpected row structures without crashing the whole script.
Quick Tips to Avoid Future Selector Issues
- Always inspect the page’s HTML (right-click > Inspect) to confirm element classes or positions before writing selectors.
- Avoid overly specific selectors (like long nested paths) — stick to stable class names or structural pseudo-classes like
:nth-child. - Test selectors in your browser’s dev tools (Console tab:
document.querySelectorAll("your-selector")) to verify they return the correct elements.
内容的提问来源于stack exchange,提问作者Lily
相关产品推荐
相关产品推荐

