You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取MDPI期刊编辑数据仅获10条,如何获取全部数据?

问题根源

MDPI的期刊编辑页面采用懒加载机制:页面初始化时仅渲染前10位编辑的数据,剩余内容需要滚动页面触发AJAX请求后才会加载。Scrapy直接爬取初始页面,自然只能获取到这10条数据。

解决办法

1. 定位全量数据的API接口

打开浏览器开发者工具(F12)切换到「网络」标签,滚动编辑列表页面,会捕获到格式如下的API请求:

GET https://www.mdpi.com/journal/[期刊名]/editors/general_editors?offset=0&limit=100
  • offset:设置为0即可从第一条开始获取全量数据
  • limit:设置一个大于实际编辑数量的值(比如100),确保一次性拿到所有数据

2. 改写Scrapy爬虫代码

直接请求API接口获取完整数据,同时简化选择器逻辑,无需区分带图/不带图的编辑:

import scrapy
from urllib.parse import urljoin

class MdpjournalSpider(scrapy.Spider):
    name = 'mdpi_editors'
    start_urls = ["https://www.mdpi.com/journal/agrochemicals/editors"]

    def parse(self, response):
        # 从当前URL提取期刊名称
        journal_name = response.url.split('/')[-2]
        # 构造全量编辑数据的API地址
        api_url = urljoin(response.url, f"editors/general_editors?offset=0&limit=100")
        
        yield scrapy.Request(
            url=api_url,
            callback=self.parse_editors,
            meta={'journal': journal_name}
        )

    def parse_editors(self, response):
        # API返回HTML片段,直接解析即可
        for editor_div in response.css("div.editor-div"):
            editor_name = editor_div.css("div.editor-div__content b::text").get()
            if editor_name:
                yield {
                    "期刊": response.meta['journal'],
                    "编辑姓名": editor_name.strip(),
                    "角色": "编辑委员会成员"
                }

3. 批量爬取所有期刊编辑数据

如果需要爬取全部MDPI期刊的编辑信息,先从期刊列表页提取所有期刊标识,再批量请求对应API:

import scrapy
from urllib.parse import urljoin

class AllMdpiEditorsSpider(scrapy.Spider):
    name = 'all_mdpi_editors'
    start_urls = ["https://www.mdpi.com/about/journals"]

    def parse(self, response):
        # 提取所有期刊的链接,获取期刊标识
        journal_links = response.css("div.journals-list__item a::attr(href)").getall()
        for link in journal_links:
            journal_slug = link.split('/')[-1]
            # 构造对应期刊的编辑API地址
            editor_api = urljoin(response.url, f"/journal/{journal_slug}/editors/general_editors?offset=0&limit=100")
            yield scrapy.Request(
                url=editor_api,
                callback=self.parse_editors,
                meta={'journal': journal_slug}
            )

    def parse_editors(self, response):
        for editor_div in response.css("div.editor-div"):
            editor_name = editor_div.css("div.editor-div__content b::text").get()
            if editor_name:
                yield {
                    "期刊": response.meta['journal'],
                    "编辑姓名": editor_name.strip(),
                    "角色": "编辑委员会成员"
                }
优化说明
  • 直接调用官方API绕开前端懒加载的DOM渲染问题,无需模拟页面滚动
  • 简化选择器逻辑:无需区分带图/不带图的编辑,单个选择器即可精准提取编辑姓名
  • 批量爬取时先获取所有期刊标识再批量请求API,大幅提升爬取效率

内容的提问来源于stack exchange,提问作者Harikrishnan V

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 13:03:16