使用Python+lxml通过XPath提取评论有用数返回空值的问题求助
Hey there! Let's figure out why your XPath query is returning an empty list even though the helpfulness count exists on the page. Here are the most common issues and fixes to try:
1. First, Verify the Raw HTML Content You're Parsing
Sometimes the HTML you get with your request (stored in page) doesn't match what you see in your browser. This could be due to:
- Missing user-agent header: Many sites serve different content to bots vs. real browsers. Add a proper user-agent when fetching the page.
- Anti-scraping measures: The site might block your request or return minimal content.
- Login requirements: If the page requires being logged in to see the counts, your request won't include session cookies.
Quick check: Save the page content to a file and open it in a browser to see if the <span class="brand-find-useful__count"> exists there. Example code to save it:
with open('page_source.html', 'w', encoding='utf-8') as f: f.write(page)
If the span isn't in this file, your problem is with fetching the page, not the XPath.
2. Adjust Your XPath Query
Your current XPath looks right, but sometimes small tweaks help:
- Maybe the span has multiple class names? Use
contains()to match part of the class:helpfulness = tree.xpath('//span[contains(@class, "brand-find-useful__count")]/text()') - Ensure there's no extra whitespace in the class name in the HTML. You can also trim whitespace from the result later if needed.
3. Check for Dynamic Content (JavaScript Loading)
If the span is loaded dynamically with JavaScript after the initial page load, lxml (which only parses static HTML) won't see it. In this case, you need to use a tool that renders JavaScript, like Selenium or Playwright.
Example with Selenium:
from selenium import webdriver from selenium.webdriver.common.by import By driver = webdriver.Chrome() driver.get("your_page_url_here") # Wait for the element to load (adjust timeout as needed) driver.implicitly_wait(10) helpfulness_elem = driver.find_element(By.CLASS_NAME, "brand-find-useful__count") helpfulness = helpfulness_elem.text if helpfulness_elem else [] driver.quit() print(helpfulness)
4. Check if the Element is Inside an iframe
Occasionally, content is embedded in an iframe. If that's the case, you need to switch to the iframe first before querying the element. With Selenium, you'd do something like:
driver.switch_to.frame("iframe_id_or_name") # Then find the element as before
Try these steps one by one—start with verifying the raw HTML, since that's the easiest check. Most likely, either your request isn't getting the full page content, or the content is loaded dynamically.
内容的提问来源于stack exchange,提问作者Principia

