You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取Moneycontrol印度指数页面时,正确XPath/CSS选择器无返回结果的问题求助

Troubleshooting Web Scraping Issues on Moneycontrol's Indian Indices Page

Let's break down and fix the scraping issues you're facing with the Moneycontrol Indian Indices page:

First, Fix the Escaped Characters in Your Selectors

Looking at your code, you’re using HTML-escaped characters (" instead of ", > instead of >) which won’t work with Scrapy’s XPath/CSS selectors. These are HTML entities, and Scrapy expects raw syntax for selector queries.

Here are the corrected versions of your code:

  • Fixed XPath:
    countries = response.xpath('//*[@id="indices_stocks"]/div[2]/div/div/div/div[2]/table/tbody[2]/tr/td[1]/p/a/text()').getall()
    
  • Fixed CSS Selector:
    countries = response.css('#indices_stocks > div.indices > div > div > div > div.pinned > table > tbody:nth-child(2) > tr > td:nth-child(1) > p > a::text').getall()
    

Check if the Content is Dynamically Loaded

Moneycontrol often loads table data via JavaScript (AJAX calls or client-side rendering), which means Scrapy’s default response only captures the initial static HTML—not the data you’re targeting.

To verify this:

  1. Run view(response) in your Scrapy spider to open the captured HTML in a browser. If the table is missing, you’re dealing with dynamic content.
  2. For dynamic content, use Scrapy’s integration with tools like Playwright to render the page fully. Here’s a quick example setup:
    from scrapy_playwright.page import PageCoroutine
    
    class MoneycontrolSpider(scrapy.Spider):
        name = 'moneycontrol'
        start_urls = ['https://www.moneycontrol.com/markets/indian-indices/']
    
        custom_settings = {
            'DOWNLOAD_HANDLERS': {
                "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
            },
            'PLAYWRIGHT_LAUNCH_OPTIONS': {
                'headless': True,
            },
        }
    
        def start_requests(self):
            for url in self.start_urls:
                yield scrapy.Request(
                    url,
                    meta={
                        'playwright': True,
                        'playwright_page_coroutines': [
                            PageCoroutine('wait_for_selector', '#indices_stocks table'),
                        ],
                    },
                )
    
        def parse(self, response):
            # Use your corrected selectors here
            countries = response.xpath('//*[@id="indices_stocks"]/div[2]/div/div/div/div[2]/table/tbody[2]/tr/td[1]/p/a/text()').getall()
            yield {'countries': countries}
    

Verify Selector Validity in Browser DevTools

Before relying on your selectors, test them directly in your browser’s Developer Tools:

  1. Open the Moneycontrol page, right-click the element you want to scrape, and select "Inspect".
  2. In the DevTools console, test your XPath with $x('your-xpath-here') or CSS selector with $$('your-css-selector-here').
  3. Note: Browsers sometimes auto-add <tbody> tags to tables even if they aren’t in the raw HTML. If your selector includes tbody but it’s missing from the actual page source, try removing it (e.g., use //table/tr instead of //table/tbody/tr).

内容的提问来源于stack exchange,提问作者Dhinesh Mathiyalagan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 03:17:29