Scrapy爬取MDPI期刊编辑数据仅获10条,如何获取全部数据?
问题根源
MDPI的期刊编辑页面采用懒加载机制:页面初始化时仅渲染前10位编辑的数据,剩余内容需要滚动页面触发AJAX请求后才会加载。Scrapy直接爬取初始页面,自然只能获取到这10条数据。
解决办法
1. 定位全量数据的API接口
打开浏览器开发者工具(F12)切换到「网络」标签,滚动编辑列表页面,会捕获到格式如下的API请求:
GET https://www.mdpi.com/journal/[期刊名]/editors/general_editors?offset=0&limit=100
offset:设置为0即可从第一条开始获取全量数据limit:设置一个大于实际编辑数量的值(比如100),确保一次性拿到所有数据
2. 改写Scrapy爬虫代码
直接请求API接口获取完整数据,同时简化选择器逻辑,无需区分带图/不带图的编辑:
import scrapy from urllib.parse import urljoin class MdpjournalSpider(scrapy.Spider): name = 'mdpi_editors' start_urls = ["https://www.mdpi.com/journal/agrochemicals/editors"] def parse(self, response): # 从当前URL提取期刊名称 journal_name = response.url.split('/')[-2] # 构造全量编辑数据的API地址 api_url = urljoin(response.url, f"editors/general_editors?offset=0&limit=100") yield scrapy.Request( url=api_url, callback=self.parse_editors, meta={'journal': journal_name} ) def parse_editors(self, response): # API返回HTML片段,直接解析即可 for editor_div in response.css("div.editor-div"): editor_name = editor_div.css("div.editor-div__content b::text").get() if editor_name: yield { "期刊": response.meta['journal'], "编辑姓名": editor_name.strip(), "角色": "编辑委员会成员" }
3. 批量爬取所有期刊编辑数据
如果需要爬取全部MDPI期刊的编辑信息,先从期刊列表页提取所有期刊标识,再批量请求对应API:
import scrapy from urllib.parse import urljoin class AllMdpiEditorsSpider(scrapy.Spider): name = 'all_mdpi_editors' start_urls = ["https://www.mdpi.com/about/journals"] def parse(self, response): # 提取所有期刊的链接,获取期刊标识 journal_links = response.css("div.journals-list__item a::attr(href)").getall() for link in journal_links: journal_slug = link.split('/')[-1] # 构造对应期刊的编辑API地址 editor_api = urljoin(response.url, f"/journal/{journal_slug}/editors/general_editors?offset=0&limit=100") yield scrapy.Request( url=editor_api, callback=self.parse_editors, meta={'journal': journal_slug} ) def parse_editors(self, response): for editor_div in response.css("div.editor-div"): editor_name = editor_div.css("div.editor-div__content b::text").get() if editor_name: yield { "期刊": response.meta['journal'], "编辑姓名": editor_name.strip(), "角色": "编辑委员会成员" }
优化说明
- 直接调用官方API绕开前端懒加载的DOM渲染问题,无需模拟页面滚动
- 简化选择器逻辑:无需区分带图/不带图的编辑,单个选择器即可精准提取编辑姓名
- 批量爬取时先获取所有期刊标识再批量请求API,大幅提升爬取效率
内容的提问来源于stack exchange,提问作者Harikrishnan V
相关产品推荐
相关产品推荐

