基于Python与Selenium爬取词典网站动词释义的技术需求问询
解决方案:用Selenium爬取Dictionary.com的动词释义
Hey there! Since you're new to Python web scraping with Selenium, let's break down exactly how to target and extract verb definitions from dictionary.com. I'll walk you through each step with clear code and explanations that make sense for a beginner.
准备工作
Before we jump into code, let's get the basics sorted:
- Install Selenium via pip:
pip install selenium - Download the matching WebDriver for your browser (like ChromeDriver for Google Chrome) — make sure it's either in your system PATH or you specify its file path directly in the code.
完整代码实现
Here's a working script that does exactly what you need. Just swap out "run" with your target word when you test it:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.keys import Keys from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time def get_verb_definitions(target_word): # Initialize Chrome browser (swap with Firefox/Edge driver if you prefer) driver = webdriver.Chrome() try: # Navigate to dictionary.com's homepage driver.get("https://www.dictionary.com/") # Wait for the search input to load, then type in our word search_box = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "global-search-input")) ) search_box.send_keys(target_word) # Press Enter to trigger search (easier than locating the search icon) search_box.send_keys(Keys.ENTER) # Give the page a second to load results (we can use WebDriverWait here too for robustness) time.sleep(2) # Find all sections labeled "verb" on the results page verb_sections = driver.find_elements(By.XPATH, "//div[contains(@class, 'css-1urpfgu') and .//span[text()='verb']]") if not verb_sections: print(f"No verb definitions found for '{target_word}'.") return [] # Pull out all the definition text from each verb section verb_defs = [] for section in verb_sections: definition_items = section.find_elements(By.XPATH, ".//div[contains(@class, 'css-10n3ydx e1hk9ate0')]") for idx, def_text in enumerate(definition_items, 1): verb_defs.append(f"{idx}. {def_text.text.strip()}") return verb_defs except Exception as e: print(f"Oops, something went wrong: {str(e)}") return [] finally: # Make sure the browser closes even if we hit an error driver.quit() # Example usage word_to_check = "run" results = get_verb_definitions(word_to_check) if results: print(f"Verb definitions for '{word_to_check}':") for item in results: print(item)
关键部分解释
Let's walk through the important bits so you understand what's happening:
- WebDriverWait: We use this instead of just
time.sleep()whenever possible to wait for elements to load — this makes the script way more reliable if the site is slow to load. - XPath Selectors: These help us zero in on the exact "verb" sections and their definitions. Heads up: Website HTML can change over time, so you might need to tweak these selectors if dictionary.com updates their layout.
- Error Handling: The
try/finallyblock ensures the browser always closes, even if something goes wrong mid-scrape. - Alternative to Clicking Search Icon: Pressing Enter after typing the word is simpler than locating and clicking the search button — it works just as well!
新手注意事项
- Double-check that your WebDriver version matches your browser version (e.g., ChromeDriver 118 for Chrome 118) — mismatches cause weird errors.
- If the site blocks your script, try adding longer delays or using a headless browser (add
options.add_argument("--headless=new")when initializing Chrome to run it without a visible window). - Always respect the website's terms of service and
robots.txtwhen scraping — don't hammer their servers with too many requests too quickly.
内容的提问来源于stack exchange,提问作者Sameeresque
相关产品推荐
相关产品推荐

