You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Scrapy中发送POST请求后如何获取HTML响应而非JSON响应

解决方案

你调用的SubmitSearch接口本身设计就是仅返回JSON格式的搜索结果数据,页面HTML是浏览器拿到JSON后通过前端JS动态渲染生成的,接口本身不会返回HTML内容,有两种可行方案解决需求:

方案1:直接解析JSON提取数据(优先推荐)

这是效率最高、开发成本最低的方案,不需要额外获取渲染后的HTML,搜索所需的所有字段都已经包含在返回的JSON中:

  • 调整回调函数直接解析响应即可:
import json
def start(self, response):
    # 将响应转为Python字典,直接提取所需字段
    search_result = json.loads(response.body)
    # 打印字段结构确认需要的内容位置
    print(search_result.keys())
    # 举例:提取机构列表后遍历处理
    for org_item in search_result.get("orgList", []):
        yield {
            "org_name": org_item.get("name"),
            "ein": org_item.get("ein"),
            # 其他需要的字段
        }

方案2:获取动态渲染后的完整HTML

如果确实需要渲染完成的HTML源码(比如要提取复杂样式关联的内容、验证页面展示效果等),可以使用scrapy-playwright对接无头浏览器模拟用户操作,直接获取渲染后的页面:

  1. 先安装依赖:
pip install scrapy-playwright
playwright install chromium
  1. 修改爬虫代码:
import scrapy
from scrapy_playwright.page import PageCoroutine

class NonprofitSpider(scrapy.Spider):
    name = 'nonprofit'
    custom_settings = {
        "DOWNLOAD_HANDLERS": {
            "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
            "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
        },
        "PLAYWRIGHT_LAUNCH_OPTIONS": {
            "headless": True,
        }
    }
    def start_requests(self):
        yield scrapy.Request(
            url="https://www.guidestar.org/search",
            meta={
                "playwright": True,
                "playwright_page_coroutines": [
                    # 模拟选择阿拉斯加州的操作,选择器根据页面实际元素调整
                    PageCoroutine("select_option", 'select[name="State"]', value="Alaska"),
                    # 点击提交搜索按钮
                    PageCoroutine("click", '#search-submit'),
                    # 等待搜索结果元素加载完成,选择器根据页面实际元素调整
                    PageCoroutine("wait_for_selector", ".search-result-item", timeout=10000),
                ]
            },
            callback=self.parse_rendered_page
        )
    
    def parse_rendered_page(self, response):
        # 此处response为渲染完成的完整页面,直接用xpath/css提取内容即可
        print(response.text)
        # 后续提取逻辑

注意说明

  • 无特殊需求优先使用方案1,接口返回的JSON已经是结构化数据,不需要额外做HTML解析,运行效率和稳定性都更高
  • 页面元素的选择器需要你根据站点实际的页面代码微调,替换示例中的占位选择器即可

内容的提问来源于stack exchange,提问作者Albert

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 22:15:05