如何通过XPath从currys.co.uk爬取智能手表的重量值?
Hey there, let's tackle this XPath problem you're facing. The issue with your original query is that you're targeting the text node containing "Weight" instead of the parent <th> element—so trying to traverse to a <td> from that text node won't work. Here are a few reliable solutions tailored to the structure you shared, plus safeguards for inconsistent page layouts:
1. Direct Sibling Traversal (Matches Your Example Structure)
Use this XPath to target the <th> element that contains "Weight", then grab its immediate following sibling <td>'s text:
//th[contains(normalize-space(text()), 'Weight')]/following-sibling::td/text()
normalize-space()handles any extra whitespace around "Weight" (like newlines or tabs)following-sibling::tddirectly targets the adjacent<td>holding the weight value
2. Row-Level Targeting (More Robust for Inconsistent Layouts)
If some pages wrap the <th> and <td> in different nested structures but still share the same <tr>, this version first finds the entire row containing the "Weight" header, then extracts the <td> content:
//tr[th[contains(normalize-space(text()), 'Weight')]]/td/text()
This works even if there are extra elements between the <th> and <td> within the row.
3. Case-Insensitive Match (Handles Variations Like "weight" or "WEIGHT")
Currys might use inconsistent capitalization across product pages. Use this XPath to ignore case:
//th[contains(translate(normalize-space(text()), 'WEIGHT', 'weight'), 'weight')]/following-sibling::td/text()
translate()converts all uppercase letters in the text to lowercase, so matches any capitalization of "Weight"
Bonus: Check for Dynamic Content
If none of these XPaths return results, the weight data might be loaded dynamically via JavaScript (common on e-commerce sites). In that case, you'll need to use a tool like Selenium or Playwright to render the page fully before scraping, rather than just fetching static HTML.
内容的提问来源于stack exchange,提问作者sophocles

