使用BeautifulSoup提取子元素数据时同class标签冗余内容过滤问题
问题解决思路
两类功能完全不同的span使用了相同的re-offer-type类名,全局检索会同时命中两类数据,只需要通过父容器限定检索范围即可精准提取中介名称,完全过滤掉“出售”类无效文本。
最优修改方案
直接调整publisher字段的检索规则,通过CSS选择器限定仅提取seller容器下的re-offer-type标签,该容器是中介信息的专属父级,不会混入“продава”出售标签:
把原代码中提取publisher的循环替换为以下代码:
for publish in soup.find('ul', {'class': 'list-view real-estates'}).select('div.seller span.re-offer-type'): publish_value = publish.get_text().strip() publisher.append(publish_value)
该方案比直接过滤掉
продава字符串更稳定,不受网站文本翻译、标签文案修改的影响。
额外优化建议
原代码全局检索所有字段存在数据错位风险:如果某套房源缺失任意字段,多个列表的顺序就会无法对应。建议改为先遍历单套房源的父节点,再在节点内提取对应字段,修改后的完整函数如下:
def get_prices(urls): # 先拿到所有房源的父节点 items = soup.find('ul', {'class': 'list-view real-estates'}).find_all('li', class_=['odd', 'even']) for item in items: # 提取价格 price_text = item.find('strong', {'class': 'price'}).get_text() final_price = ''.join(re.findall('[0-9]+', price_text)) prices.append(final_price) # 提取房屋基础信息 property_type_text = item.find('div', {'class': 'inline-group'}).get_text() # 提取房屋类型 property_type_value = ' '.join(property_type_text.split(',')[0].split()[1:3]) type_of_property.append(property_type_value) # 提取面积 sqm_value = property_type_text.split(',')[1].split()[0] sqm_area.append(sqm_value) # 提取位置 location_value = property_type_text.split(',')[-1].strip() locations.append(location_value) # 提取中介名称 publish_value = item.select_one('div.seller span.re-offer-type').get_text().strip() publisher.append(publish_value) return prices, type_of_property, sqm_area, locations, publisher
内容的提问来源于stack exchange,提问作者tsetsko
相关产品推荐
相关产品推荐

