使用Python LXML解析牛津词典获取词源遇阻求助
Fixing Your Oxford Dictionary Etymology Scraper with LXML
Let's break down what's going wrong with your code and fix it step by step—your core idea is solid, but you're missing a few critical pieces around anti-bypass, reliable element targeting, and text extraction.
Key Issues in Your Original Code
- Fragile XPath: Your hardcoded path (
//div[1]/div[2]/div/div/div/div[1]/section[5]/div/p) relies on exact page structure, which Oxford frequently updates. This is almost certainly why you're not getting results. - No Request Headers: Oxford's site blocks raw
urlopenrequests that don't look like they're coming from a browser. - Unused Exception: You defined
SkipExceptionbut never actually throw it, and you don't handle common web scraping errors (like 404s or network issues). - Missing Text Extraction: Even if you found the right elements, you don't pull the actual text content from them.
Fixed & Improved Code
import lxml.html from urllib.request import urlopen, Request from urllib.error import URLError, HTTPError class SkipException(Exception): def __init__(self, value): self.value = value def fetch_oxford_etymology(target_word): base_url = "https://en.oxforddictionaries.com/definition/{word}" url = base_url.format(word=target_word) # Mimic a browser request to avoid anti-scraping blocks browser_headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } try: # Create a request with headers and fetch the page request = Request(url, headers=browser_headers) page_response = urlopen(request) doc = lxml.html.fromstring(page_response.read()) # Use semantic class-based XPath to find the etymology section # This is far more reliable than hardcoding div/section positions etymology_sections = doc.xpath("//section[contains(@class, 'etymology')]//p") if not etymology_sections: raise SkipException(f"No etymology entry found for '{target_word}'") # Extract clean text from all matching paragraphs etymology_content = "\n\n".join( para.text_content().strip() for para in etymology_sections ) return etymology_content # Handle common web errors except HTTPError as err: if err.code == 404: raise SkipException(f"Word '{target_word}' not found in Oxford Dictionary") else: raise SkipException(f"HTTP Error {err.code}: {err.reason}") except URLError as err: raise SkipException(f"Network Error: {str(err)}") except Exception as err: raise SkipException(f"Unexpected error: {str(err)}") # Example usage if __name__ == "__main__": try: word_etymology = fetch_oxford_etymology("good") print(f"Etymology of 'good':\n{word_etymology}") except SkipException as e: print(f"Error: {e.value}")
What Changed & Why
- Browser Headers: The
User-Agentheader tricks Oxford's server into thinking the request comes from a real browser, avoiding immediate blocks. - Reliable XPath: Instead of guessing section positions, we target the
etymologyclass directly—this will keep working even if Oxford rearranges other page elements. - Proper Exception Handling: We now throw your
SkipExceptionin meaningful scenarios (missing etymology, 404s, network issues) and handle standard web scraping errors. - Text Extraction: Using
text_content()pulls all text from the paragraph elements (including any nested spans or links), and we join them into a clean, readable string.
Note
Oxford's terms of service may restrict automated scraping. Make sure you're using this code for personal, non-commercial use and don't send too many requests in a short period to avoid being permanently blocked.
内容的提问来源于stack exchange,提问作者Núria Bosch
相关产品推荐
相关产品推荐

