如何从无class属性的span标签提取结构化房源数据?
解决Scrapy抓取无class的span房源信息结构化问题
完全可以在抓取阶段就完成拆分,不用等到后续处理,直接输出结构化的rooms、bathrooms、sqm2字段,效率更高。
修改后的代码实现
import re def parse(self, response): # 正则匹配数字,兼容content属性或纯文本里的数值 num_pattern = re.compile(r'\d+') for product in response.css('div.listing.listing-card'): # 初始化结构化字段,缺失时返回None rooms = None bathrooms = None sqm2 = None # 遍历每个span标签,按关键词匹配对应字段 for span in product.css('span'): span_text = span.css('::text').get(default='').strip() span_content = span.attrib.get('content', '') # 匹配房间数(habitaciones) if 'habitaciones' in span_text: rooms = span_content if span_content else (num_pattern.search(span_text).group() if num_pattern.search(span_text) else None) # 匹配浴室数(baños) elif 'baños' in span_text: bathrooms = span_content if span_content else (num_pattern.search(span_text).group() if num_pattern.search(span_text) else None) # 匹配面积(m²) elif 'm²' in span_text: sqm2 = span_content if span_content else (num_pattern.search(span_text).group() if num_pattern.search(span_text) else None) yield { 'name': product.css('div.listing-card__title::text').get(default='').strip(), 'location': product.css('div.listing-card__location::text').get(default='').strip(), 'link': product.attrib.get('data-href', ''), 'rooms': rooms, 'bathrooms': bathrooms, 'sqm2': sqm2 }
核心逻辑说明
- 兼容两种span格式:优先读取
content属性的数值(更准确),没有的话用正则从文本中提取数字 - 关键词匹配:通过
habitaciones、baños、m²这些关键词定位对应的字段 - 缺失值处理:用
get(default='')和条件判断确保缺失信息的字段返回None,不会报错 - 结构化输出:直接生成带有明确字段的字典,后续无需额外清洗
内容的提问来源于stack exchange,提问作者Jonathan Montaluisa
相关产品推荐
相关产品推荐

