网页爬取求助:如何提取同类别名元素中的房屋产权类型与EPC评级信息
网页爬取求助:如何提取同类别名元素中的房屋产权类型与EPC评级信息
嗨,我看到你在爬取Zoopla的房产数据时遇到了个头疼的小问题——房屋产权类型(比如Freehold)和EPC评级的元素用了同一个类名,导致没法直接区分提取对吧?别担心,咱们一步步来解决这个问题~
问题根源分析
你之前的代码里直接用find_element(By.CLASS_NAME, "_14bi3x30")去提取信息,但find_element只会返回当前房屋下第一个匹配到的元素。如果产权类型在HTML结构里排在EPC前面,那你拿到的其实是产权而不是EPC,反之亦然,这就导致信息提取完全错位了。
解决方案:通过内容特征区分同类别名元素
我们可以先获取当前房屋下所有带_14bi3x30类名的元素,然后根据它们的文本特征来区分:
- 产权类型的文本通常是
Freehold或Leasehold这类固定关键词 - EPC评级是单个字母(A-G),长度为1
基于这个逻辑,我修改了你的代码,现在可以同时准确提取这两个字段:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium import webdriver import pandas as pd import time # Initialize WebDriver driver = webdriver.Chrome() # Open URL url = "https://www.zoopla.co.uk/house-prices/england/?new_homes=include&q=england+&orig_q=united+kingdom&view_type=list&pn=1" driver.get(url) # Wait for the main content to load (adjust time as needed) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "_17smgnt0")) ) # Initialize result list to store data result = [] # Find all house elements houses = driver.find_elements(By.CLASS_NAME, "_1hzil3o0") # Extract and print addresses for house in houses: try: # 先提取基础信息 address = house.find_element(By.XPATH, './/a/h2').text date_last_sold = house.find_element(By.CSS_SELECTOR, "._1hzil3o9._1hzil3o8._194zg6t7").text num_rooms = house.find_element(By.CLASS_NAME, "_1pbf8i53").text # 提取所有同类别名的属性元素,区分产权和EPC property_attrs = house.find_elements(By.CLASS_NAME, "_14bi3x30") freehold_type = None epc_rating = None for attr in property_attrs: attr_text = attr.text.strip() # 判断产权类型 if attr_text in ["Freehold", "Leasehold"]: freehold_type = attr_text # 判断EPC评级(单个字母A-G) elif len(attr_text) == 1 and attr_text.upper() in "ABCDEFG": epc_rating = attr_text.upper() item = { "address": address, "DateLast_sold": date_last_sold, "Number of Rooms": num_rooms, "Property Type (Freehold/Leasehold)": freehold_type, "EPC Rating": epc_rating } result.append(item) # Append to the result list except Exception as e: print(f"Error extracting data for house: {e}") # Store the result into a dataframe after the loop df = pd.DataFrame(result) # Show the result print(df) # Close the driver driver.quit()
额外小提醒
- 网站的类名(比如
_14bi3x30、_1hzil3o0)属于动态生成的,可能会随时更新,如果之后代码突然失效,记得重新检查网页的HTML结构,更新定位符。 - 有些房屋可能没有EPC评级或者产权信息,代码里已经用
None来处理这种情况,你也可以根据需求改成默认值(比如"Unknown")。
备注:内容来源于stack exchange,提问作者Chioma Okoroafor
相关产品推荐
相关产品推荐

