You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python与Selenium爬取词典网站动词释义的技术需求问询

解决方案:用Selenium爬取Dictionary.com的动词释义

Hey there! Since you're new to Python web scraping with Selenium, let's break down exactly how to target and extract verb definitions from dictionary.com. I'll walk you through each step with clear code and explanations that make sense for a beginner.

准备工作

Before we jump into code, let's get the basics sorted:

  • Install Selenium via pip: pip install selenium
  • Download the matching WebDriver for your browser (like ChromeDriver for Google Chrome) — make sure it's either in your system PATH or you specify its file path directly in the code.

完整代码实现

Here's a working script that does exactly what you need. Just swap out "run" with your target word when you test it:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

def get_verb_definitions(target_word):
    # Initialize Chrome browser (swap with Firefox/Edge driver if you prefer)
    driver = webdriver.Chrome()
    try:
        # Navigate to dictionary.com's homepage
        driver.get("https://www.dictionary.com/")
        
        # Wait for the search input to load, then type in our word
        search_box = WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.ID, "global-search-input"))
        )
        search_box.send_keys(target_word)
        # Press Enter to trigger search (easier than locating the search icon)
        search_box.send_keys(Keys.ENTER)
        
        # Give the page a second to load results (we can use WebDriverWait here too for robustness)
        time.sleep(2)
        
        # Find all sections labeled "verb" on the results page
        verb_sections = driver.find_elements(By.XPATH, "//div[contains(@class, 'css-1urpfgu') and .//span[text()='verb']]")
        
        if not verb_sections:
            print(f"No verb definitions found for '{target_word}'.")
            return []
        
        # Pull out all the definition text from each verb section
        verb_defs = []
        for section in verb_sections:
            definition_items = section.find_elements(By.XPATH, ".//div[contains(@class, 'css-10n3ydx e1hk9ate0')]")
            for idx, def_text in enumerate(definition_items, 1):
                verb_defs.append(f"{idx}. {def_text.text.strip()}")
        
        return verb_defs
    
    except Exception as e:
        print(f"Oops, something went wrong: {str(e)}")
        return []
    finally:
        # Make sure the browser closes even if we hit an error
        driver.quit()

# Example usage
word_to_check = "run"
results = get_verb_definitions(word_to_check)
if results:
    print(f"Verb definitions for '{word_to_check}':")
    for item in results:
        print(item)

关键部分解释

Let's walk through the important bits so you understand what's happening:

  • WebDriverWait: We use this instead of just time.sleep() whenever possible to wait for elements to load — this makes the script way more reliable if the site is slow to load.
  • XPath Selectors: These help us zero in on the exact "verb" sections and their definitions. Heads up: Website HTML can change over time, so you might need to tweak these selectors if dictionary.com updates their layout.
  • Error Handling: The try/finally block ensures the browser always closes, even if something goes wrong mid-scrape.
  • Alternative to Clicking Search Icon: Pressing Enter after typing the word is simpler than locating and clicking the search button — it works just as well!

新手注意事项

  • Double-check that your WebDriver version matches your browser version (e.g., ChromeDriver 118 for Chrome 118) — mismatches cause weird errors.
  • If the site blocks your script, try adding longer delays or using a headless browser (add options.add_argument("--headless=new") when initializing Chrome to run it without a visible window).
  • Always respect the website's terms of service and robots.txt when scraping — don't hammer their servers with too many requests too quickly.

内容的提问来源于stack exchange,提问作者Sameeresque

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:34:19