Scrapy爬取Moneycontrol印度指数页面时,正确XPath/CSS选择器无返回结果的问题求助
Let's break down and fix the scraping issues you're facing with the Moneycontrol Indian Indices page:
First, Fix the Escaped Characters in Your Selectors
Looking at your code, you’re using HTML-escaped characters (" instead of ", > instead of >) which won’t work with Scrapy’s XPath/CSS selectors. These are HTML entities, and Scrapy expects raw syntax for selector queries.
Here are the corrected versions of your code:
- Fixed XPath:
countries = response.xpath('//*[@id="indices_stocks"]/div[2]/div/div/div/div[2]/table/tbody[2]/tr/td[1]/p/a/text()').getall() - Fixed CSS Selector:
countries = response.css('#indices_stocks > div.indices > div > div > div > div.pinned > table > tbody:nth-child(2) > tr > td:nth-child(1) > p > a::text').getall()
Check if the Content is Dynamically Loaded
Moneycontrol often loads table data via JavaScript (AJAX calls or client-side rendering), which means Scrapy’s default response only captures the initial static HTML—not the data you’re targeting.
To verify this:
- Run
view(response)in your Scrapy spider to open the captured HTML in a browser. If the table is missing, you’re dealing with dynamic content. - For dynamic content, use Scrapy’s integration with tools like Playwright to render the page fully. Here’s a quick example setup:
from scrapy_playwright.page import PageCoroutine class MoneycontrolSpider(scrapy.Spider): name = 'moneycontrol' start_urls = ['https://www.moneycontrol.com/markets/indian-indices/'] custom_settings = { 'DOWNLOAD_HANDLERS': { "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", }, 'PLAYWRIGHT_LAUNCH_OPTIONS': { 'headless': True, }, } def start_requests(self): for url in self.start_urls: yield scrapy.Request( url, meta={ 'playwright': True, 'playwright_page_coroutines': [ PageCoroutine('wait_for_selector', '#indices_stocks table'), ], }, ) def parse(self, response): # Use your corrected selectors here countries = response.xpath('//*[@id="indices_stocks"]/div[2]/div/div/div/div[2]/table/tbody[2]/tr/td[1]/p/a/text()').getall() yield {'countries': countries}
Verify Selector Validity in Browser DevTools
Before relying on your selectors, test them directly in your browser’s Developer Tools:
- Open the Moneycontrol page, right-click the element you want to scrape, and select "Inspect".
- In the DevTools console, test your XPath with
$x('your-xpath-here')or CSS selector with$$('your-css-selector-here'). - Note: Browsers sometimes auto-add
<tbody>tags to tables even if they aren’t in the raw HTML. If your selector includestbodybut it’s missing from the actual page source, try removing it (e.g., use//table/trinstead of//table/tbody/tr).
内容的提问来源于stack exchange,提问作者Dhinesh Mathiyalagan

