You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取NepseAlpha网站无数据返回问题求助

解决Scrapy爬取NEPSE Alpha IPO日历空响应问题

问题原因

目标网站的IPO日历表格是JavaScript动态渲染的,Scrapy默认请求仅会获取静态HTML源码,而表格数据是页面加载完成后通过AJAX请求从后端接口拉取并渲染生成的,静态源码中不存在DataTables_Table_0这个表格节点,因此XPath查询返回空结果。


方案一:直接抓取后端API接口(推荐)

通过浏览器开发者工具的「Network」面板(筛选XHR/Fetch请求),可以定位到表格数据的接口,直接请求该接口即可获取结构化JSON数据,效率远高于JS渲染。

爬虫代码示例:

import scrapy
from scrapy import Request


class ShareInfoSpider(scrapy.Spider):
    name = 'share'

    def start_requests(self):
        # 直接调用数据接口
        api_url = "https://nepsealpha.com/api/calendar/ipo"
        yield Request(api_url, callback=self.parse_api)

    def parse_api(self, response):
        # 解析返回的JSON数据
        raw_data = response.json()
        for item in raw_data:
            # 按需提取字段,可根据接口返回的字段名调整
            yield {
                '公司名称': item.get('company_name'),
                '股票代码': item.get('scrip'),
                '申购开始日期': item.get('open_date'),
                '申购结束日期': item.get('close_date'),
                '发行类型': item.get('issue_type'),
                '发行价格': item.get('issue_price'),
                '发行份额': item.get('total_share')
            }

方案二:使用Playwright渲染JavaScript

若不想分析接口,可通过Scrapy结合Playwright实现页面JS渲染,获取完整的渲染后HTML。

步骤1:安装依赖

pip install scrapy-playwright

步骤2:修改Scrapy配置(settings.py)

# 启用Playwright下载处理器
DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}

# 配置Playwright启动选项
PLAYWRIGHT_LAUNCH_OPTIONS = {
    "headless": True,  # 无头模式,不显示浏览器窗口
}

步骤3:修改爬虫代码

import scrapy
from scrapy import Request


class ShareInfoSpider(scrapy.Spider):
    name = 'share'

    def start_requests(self):
        target_url = "https://nepsealpha.com/investment-calandar/ipo"
        # 标记该请求需要Playwright渲染
        yield Request(target_url, callback=self.parse, meta={"playwright": True})

    def parse(self, response):
        # 提取渲染后的表格行数据
        table_rows = response.xpath("//table[@id='DataTables_Table_0']//tbody/tr")
        for row in table_rows:
            yield {
                '公司名称': row.xpath(".//td[1]/text()").get().strip(),
                '股票代码': row.xpath(".//td[2]/text()").get().strip(),
                '申购开始日期': row.xpath(".//td[3]/text()").get().strip(),
                '申购结束日期': row.xpath(".//td[4]/text()").get().strip(),
                '发行类型': row.xpath(".//td[5]/text()").get().strip()
            }

内容的提问来源于stack exchange,提问作者astro geek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 13:30:52