You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Selenium爬取Houzz平台页面中的站点链接及商户信息

问题描述

我正在尝试从目标页面爬取website link,但不清楚具体的抓取逻辑,示例目标页面链接为:https://www.houzz.com/professionals/general-contractors/capital-remodeling-pfvwus-pf~663981496
我目前已编写的Selenium爬取代码如下,仅实现了商户地址、电话信息的抓取,请问如何调整代码可完成页面中站点链接的抓取?

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
import time
import pandas as pd

# chromedriver.exe存放路径
driver = webdriver.Chrome("D:/chromedriver/94/chromedriver.exe")

# 打开目标站点
driver.get("https://www.houzz.com/professionals/general-contractor")

# 等待DIV元素加载完成
WebDriverWait(driver, 60).until(EC.presence_of_element_located((By.XPATH, '//div[@class="hz-pro-search-result__right-info"]')))

# 获取所有匹配的DIV元素
info_divs = driver.find_elements(By.XPATH,  '//div[@class="hz-pro-search-result__right-info"]')

house_details = {
    "address": [],
    "phone": []
}

for row in info_divs:
    try:
        address = row.find_element(By.CLASS_NAME, "hz-pro-search-result__right-info__full-address")
        phone = row.find_element(By.CLASS_NAME, "hz-pro-search-result__right-info__contact-info")
        phone.click()
        time.sleep(0.5)
        house_details['address'].append(address.text)
        house_details['phone'].append(phone.text)
    except Exception as ex:
        print(f"出现错误:{ex}")
        
# 存储为DataFrame
house_df = pd.DataFrame.from_dict(house_details)
# 导出为CSV文件
house_df.to_csv('house_details.csv', index=False)

print(house_df)
解决方案

你只需要对原有代码做三处调整即可实现站点链接抓取:

  • 给存储结果的字典新增website字段,用于存储抓取到的站点链接
  • 遍历商户信息块时,定位到站点链接的a标签元素,提取其href属性作为链接值
  • 异常捕获分支补充空值写入逻辑,避免各字段长度不一致导致的DataFrame转换报错

调整后的完整代码如下:

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
import time
import pandas as pd

# chromedriver.exe存放路径
driver = webdriver.Chrome("D:/chromedriver/94/chromedriver.exe")

# 打开目标站点
driver.get("https://www.houzz.com/professionals/general-contractor")

# 等待DIV元素加载完成
WebDriverWait(driver, 60).until(EC.presence_of_element_located((By.XPATH, '//div[@class="hz-pro-search-result__right-info"]')))

# 获取所有匹配的DIV元素
info_divs = driver.find_elements(By.XPATH,  '//div[@class="hz-pro-search-result__right-info"]')

house_details = {
    "address": [],
    "phone": [],
    "website": [] # 新增站点链接存储字段
}

for row in info_divs:
    try:
        address = row.find_element(By.CLASS_NAME, "hz-pro-search-result__right-info__full-address")
        phone = row.find_element(By.CLASS_NAME, "hz-pro-search-result__right-info__contact-info")
        # 新增站点链接抓取逻辑
        website_ele = row.find_element(By.CSS_SELECTOR, 'a.hz-pro-search-result__right-info__website')
        website = website_ele.get_attribute('href')
        phone.click()
        time.sleep(0.5)
        house_details['address'].append(address.text)
        house_details['phone'].append(phone.text)
        house_details['website'].append(website)
    except Exception as ex:
        print(f"出现错误:{ex}")
        # 异常时对应字段补空值,避免数据长度不一致
        house_details['address'].append('')
        house_details['phone'].append('')
        house_details['website'].append('')
        
# 存储为DataFrame
house_df = pd.DataFrame.from_dict(house_details)
# 导出为CSV文件
house_df.to_csv('house_details.csv', index=False)

print(house_df)

内容的提问来源于stack exchange,提问作者GX Mentor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 18:15:04