如何处理多层嵌套HTML标签 使用Scrapy抓取影院全量电影排期数据
解决方案
核心思路
- 针对无固定层数的
movieDay嵌套结构,使用递归遍历替代固定层数的硬编码判断,可自动适配任意层级的节点嵌套 - 同一部电影的名称、分级为全局共享字段,仅需提取一次,不同放映类型、排片时间从各层
movieDay节点中单独提取
完整实现代码
# -*- coding: utf-8 -*- import scrapy class TesttwoSpider(scrapy.Spider): name = 'testtwo' allowed_domains = ['www.vscinemas.com.tw'] start_urls = ['https://www.vscinemas.com.tw/vsweb/theater/detail.aspx?id=16'] # 递归处理任意层级的movieDay节点 def parse_movie_day(self, node, base_info): # 提取当前层放映类型、排片时间 movie_type = node.xpath('./h4/text()').extract_first().strip() movie_time = node.xpath('./ul/li/a/text()').extract() # 输出结果,可替换为yield Item存入数据集 print(base_info['name']) print(base_info['grade']) print(movie_type) print(movie_time) print() # 检查是否存在子层movieDay,存在则递归处理 child_nodes = node.xpath('./div[@class="movieDay"]') for child in child_nodes: self.parse_movie_day(child, base_info) def parse(self, response): print('************ Start parse ************') # 定位到目标日期的电影列表容器 target_area = response.xpath('//div[@class="theaterTime"]/article[@id="movieTime-1508017172"]') # 遍历所有首层movieDay节点 for first_movie_day in target_area.xpath('./div[@class="movieDay"]'): # 提取当前电影公共共享字段:名称、分级 base_info = { 'name': first_movie_day.xpath('./preceding-sibling::h2[1]/text()').extract_first().strip(), 'grade': first_movie_day.xpath('./preceding-sibling::h1[1]/span/@class').extract_first() } # 触发递归遍历所有嵌套层级 self.parse_movie_day(first_movie_day, base_info)
适配说明
- 若需要将结果存入Scrapy的Item,仅需将
parse_movie_day方法中的print逻辑替换为yield Item即可 - 代码自动适配任意层数的
movieDay嵌套,无需新增层级判断逻辑
内容的提问来源于stack exchange,提问作者Morton
相关产品推荐
相关产品推荐

