使用Chrome生成的Scrapy XPath选择器失效问题求助
Hey there! Let's troubleshoot why your Chrome-generated XPath is returning an empty list in Scrapy and get that target text extracted.
Common Issues & Solutions
Remove the auto-added
<tbody>tag
Browsers like Chrome automatically insert a<tbody>element into tables when rendering, but this tag often doesn't exist in the raw HTML response Scrapy receives. Your original XPath includes/tbody/— remove that part, so it becomes://*[@id="ires"]/ol/div/div[1]/table/tr/td[1]/a/div[2]/text()Use a more flexible XPath based on the "START" marker
Instead of relying on strict DOM hierarchy (which can break if Google tweaks their layout), target the element relative to the "START" text. Try this://div[contains(text(), 'START')]/following-sibling::table/tr/td[1]/a/div[2]/text()This finds the div with "START" and grabs the target text from the immediately following table.
Test in Scrapy Shell first
Load your saved HTML into Scrapy Shell to debug interactively:scrapy shell file:///path/to/your/saved/google_page.htmlThen run
response.xpath("your_xpath_here").get()to see if it returns the text. This helps you quickly tweak the selector without running the full spider.Try CSS selectors as an alternative
CSS selectors can be more resilient sometimes. For your target text, this might work:#ires ol div div table tr td a div:nth-child(2)::textUse it in Scrapy like
response.css("your_css_selector_here").get()
Pro Tip
Google's page structure changes frequently, so avoid relying on deep nested selectors or IDs that might change. Look for stable attributes (like specific class names) or text patterns to make your selectors more durable.
内容的提问来源于stack exchange,提问作者wprins

