如何从Immobilienscout24爬虫中仅提取房屋室内面积数据?
解决方案:仅提取Immobilienscout24的室内面积
当前代码抓取面积时会同时获取室内(Wohnfläche)和室外(Grundstücksfläche)面积,以下两种方法可以帮你只提取目标室内面积:
方法一:基于位置索引筛选(适配你描述的页面结构)
如果室内面积确实出现在奇数位置(按你描述的数值列表顺序,对应Python的0-based奇数索引:1、3、5...),可以在遍历过程中通过索引判断是否处理当前面积:
def extract_information(soup): information = soup.find_all('dd', class_='font-highlight font-tabular') links = [info.find_previous('div', class_='grid-item').find('a')['href'] for info in information] # 加入索引记录,判断当前面积是否为室内面积 for idx, (link, info) in enumerate(zip(links, information)): house_link = f"https://www.immobilienscout24.de{link}" text = info.getText().strip() try: if '€' in text: price = int(text.replace('€', '').replace('.', '').replace(',', '').strip()) elif 'm²' in text: # 仅处理奇数索引的面积(若实际位置是偶数索引,改为idx % 2 == 0) if idx % 2 == 1: square_meter_str = text.replace('m²', '').replace(',', '').strip() # 保留原有的千分位/小数点处理逻辑 if '.' in square_meter_str and square_meter_str.count('.') == 1: square_meter_str = square_meter_str.replace('.', '') elif '.' in square_meter_str and square_meter_str.count('.') > 1: square_meter_str = square_meter_str.rsplit('.', 1)[0] + square_meter_str.rsplit('.', 1)[1].replace('.', '') square_meter = float(square_meter_str) if square_meter_str else 0.0 if square_meter != 0: price_per_square_meter = price / square_meter if 2000 <= price_per_square_meter <= 3000: house_data['House link'].append(house_link) house_data['Price per square meter [€]'].append(price_per_square_meter) except Exception as e: print(f"处理出错: {str(e)}") return house_data
方法二:通过上下文标签判断(更稳定,不受页面结构变化影响)
网页中每个面积对应的<dt>标签会标注面积类型(比如Wohnfläche代表室内面积),直接通过这个标签文本判断,比依赖位置更可靠:
def extract_information(soup): information = soup.find_all('dd', class_='font-highlight font-tabular') links = [info.find_previous('div', class_='grid-item').find('a')['href'] for info in information] for link, info in zip(links, information): house_link = f"https://www.immobilienscout24.de{link}" text = info.getText().strip() try: if '€' in text: price = int(text.replace('€', '').replace('.', '').replace(',', '').strip()) elif 'm²' in text: # 获取面积类型标签,确认是否为室内面积 area_type = info.find_previous('dt').getText().strip() if area_type == 'Wohnfläche': square_meter_str = text.replace('m²', '').replace(',', '').strip() # 保留原有的千分位/小数点处理逻辑 if '.' in square_meter_str and square_meter_str.count('.') == 1: square_meter_str = square_meter_str.replace('.', '') elif '.' in square_meter_str and square_meter_str.count('.') > 1: square_meter_str = square_meter_str.rsplit('.', 1)[0] + square_meter_str.rsplit('.', 1)[1].replace('.', '') square_meter = float(square_meter_str) if square_meter_str else 0.0 if square_meter != 0: price_per_square_meter = price / square_meter if 2000 <= price_per_square_meter <= 3000: house_data['House link'].append(house_link) house_data['Price per square meter [€]'].append(price_per_square_meter) except Exception as e: print(f"处理出错: {str(e)}") return house_data
关键提示
- 方法一需要你确认室内面积的实际索引位置,若页面结构变化需调整判断条件;
- 方法二更推荐,因为面积类型的标签文本通常不会轻易修改,稳定性更强;
- 原代码的空
except块已改为捕获具体异常并打印信息,方便排查问题。
内容的提问来源于stack exchange,提问作者NewUser
相关产品推荐
相关产品推荐

