使用BeautifulSoup爬取房产网站 保留嵌套结构提取租金文本
实现代码
monthly_list = [] # 遍历所有建筑节点 for building_node in rent: # 提取当前建筑下所有租金标签 monthly_tags = building_node.find_all('span',{'class':'cassetteitem_other-emphasis ui-text--bold'}) # 批量提取标签文本,解决无法批量调用.text的问题 price_list = [tag.text for tag in monthly_tags] # 按长度判断存储格式,保留嵌套结构 if len(price_list) == 1: monthly_list.append(price_list[0]) else: monthly_list.append(price_list) print(monthly_list)
改动说明
- 针对BeautifulSoup标签列表无法批量提取文本的问题:用列表推导式遍历标签列表,逐个调用
.text方法,一次性生成当前建筑所有租金的文本列表,写法简洁高效。 - 针对保留嵌套结构的需求:每处理完一栋建筑的租金列表后判断长度,单房源直接存租金字符串,多房源就存整段租金子列表,即可实现建筑和对应房源价格的关联,输出结果与你期望的格式完全匹配。
内容的提问来源于stack exchange,提问作者Jason Park
相关产品推荐
相关产品推荐

