You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取子元素数据时同class标签冗余内容过滤问题

问题解决思路

两类功能完全不同的span使用了相同的re-offer-type类名,全局检索会同时命中两类数据,只需要通过父容器限定检索范围即可精准提取中介名称,完全过滤掉“出售”类无效文本。

最优修改方案

直接调整publisher字段的检索规则,通过CSS选择器限定仅提取seller容器下的re-offer-type标签,该容器是中介信息的专属父级,不会混入“продава”出售标签:
把原代码中提取publisher的循环替换为以下代码:

for publish in soup.find('ul', {'class': 'list-view real-estates'}).select('div.seller span.re-offer-type'):
    publish_value = publish.get_text().strip()
    publisher.append(publish_value)

该方案比直接过滤掉продава字符串更稳定,不受网站文本翻译、标签文案修改的影响。

额外优化建议

原代码全局检索所有字段存在数据错位风险:如果某套房源缺失任意字段,多个列表的顺序就会无法对应。建议改为先遍历单套房源的父节点,再在节点内提取对应字段,修改后的完整函数如下:

def get_prices(urls):
    # 先拿到所有房源的父节点
    items = soup.find('ul', {'class': 'list-view real-estates'}).find_all('li', class_=['odd', 'even'])
    for item in items:
        # 提取价格
        price_text = item.find('strong', {'class': 'price'}).get_text()
        final_price = ''.join(re.findall('[0-9]+', price_text))
        prices.append(final_price)
        # 提取房屋基础信息
        property_type_text = item.find('div', {'class': 'inline-group'}).get_text()
        # 提取房屋类型
        property_type_value = ' '.join(property_type_text.split(',')[0].split()[1:3])
        type_of_property.append(property_type_value)
        # 提取面积
        sqm_value = property_type_text.split(',')[1].split()[0]
        sqm_area.append(sqm_value)
        # 提取位置
        location_value = property_type_text.split(',')[-1].strip()
        locations.append(location_value)
        # 提取中介名称
        publish_value = item.select_one('div.seller span.re-offer-type').get_text().strip()
        publisher.append(publish_value)
    return prices, type_of_property, sqm_area, locations, publisher

内容的提问来源于stack exchange,提问作者tsetsko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 22:57:03