You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何匹配动态单字母前缀的div class实现网页数据爬取?

解决Immoweb动态Class前缀的爬虫优化方案

Hey there, great job simplifying your code already! The dynamic class prefixes (xl-, l-, m-) you're facing are super common in responsive websites, and we can make your scraper even more robust and maintainable than your current approach.

核心问题分析

Your optimized code uses lists of class names to match variations, but this still requires updating the list if new prefixes (like s- for small screens) are added later. Instead, we can match elements based on the fixed suffix part of the class name using CSS attribute selectors or regex, which is far more flexible.

优化方案1:使用CSS属性包含选择器(推荐)

BeautifulSoup supports CSS selectors, and we can use the *= operator to match elements whose class attribute contains a specific substring. This works perfectly since your target elements all share fixed suffixes:

  • Results: result- + prefix → match any class starting with result-
  • Price: -price rangePrice → match any class containing this substring
  • Surface: -surface-ch → match any class containing this substring
  • Description: -desc → match any class containing this substring

Here's the revised code with this approach, plus error handling for missing elements:

import pandas as pd
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup
from fake_useragent import UserAgent
import time

# Setup Chrome options
options = Options()
options.add_argument("window-size=1400,600")
ua = UserAgent()
user_agent = ua.random
print(user_agent)
options.add_argument(f'user-agent={user_agent}')

# Initialize driver (note: consider using webdriver-manager for auto chromedriver management)
driver = webdriver.Chrome('/Users/raduulea/Documents/chromedriver', options=options)
driver.get('https://www.immoweb.be/fr/recherche/immeuble-de-rapport/a-vendre/liege/4000')
time.sleep(10)

# Initialize data lists
Title = []
address = []
price = []
surface = []
desc = []
page = 2

while True:
    time.sleep(10)
    html = driver.page_source
    soup = BeautifulSoup(html, 'html.parser')
    
    if page <= 6:
        # Match any result block with class starting with "result-"
        results = soup.select('div[class^="result-"]')
        
        for result in results:
            # Extract title
            title_elem = result.find("div", class_="title-bar-left")
            Title.append(title_elem.get_text().strip() if title_elem else 'N/A')
            
            # Extract address (fixed class name, no variation)
            addr_elem = result.find("span", class_="result-adress")
            address.append(addr_elem.get_text().strip() if addr_elem else 'N/A')
            
            # Extract price: match any class containing "-price rangePrice"
            price_elem = result.select_one('[class*="-price rangePrice"]')
            price.append(price_elem.get_text().strip() if price_elem else 'N/A')
            
            # Extract surface: match any class containing "-surface-ch"
            surface_elem = result.select_one('[class*="-surface-ch"]')
            surface.append(surface_elem.get_text().strip() if surface_elem else 'N/A')
            
            # Extract description: match any class containing "-desc"
            desc_elem = result.select_one('[class*="-desc"]')
            desc.append(desc_elem.get_text().strip() if desc_elem else 'N/A')
        
        # Navigate to next page if available
        if len(driver.find_elements_by_css_selector("a.next")) > 0:
            url = f"https://www.immoweb.be/fr/recherche/immeuble-de-rapport/a-vendre/liege/4000/?page={page}"
            driver.get(url)
            page += 1
        else:
            break
    else:
        break

# Save to CSV
df = pd.DataFrame({
    "Title": Title, 
    "Address": address, 
    "Price": price, 
    "Surface": surface, 
    "Description": desc
})
df.to_csv("immoweb_scrap_optimized.csv", index=False)
driver.quit()

关键优化点

  • Flexible element matching: No more hardcoding all possible class prefixes — this will automatically work for any new prefixes the site adds (like s- for mobile).
  • Error handling: Added checks for None elements to prevent your scraper from crashing if a field is missing on a listing.
  • Cleaner syntax: Used f-strings for URL construction, and simplified class matching logic.

优化方案2:使用正则表达式匹配Class

If you prefer regex, you can use BeautifulSoup's class_ parameter with a regex pattern to match the dynamic classes:

import re

# Match result blocks
results = soup.find_all("div", class_=re.compile(r'result-(xl|l|m)'))

# Match price element
price_elem = result.find("div", class_=re.compile(r'(xl|l|m)-price rangePrice'))

This works well too, but the CSS selector approach is more concise for this use case.

额外建议

  • Replace time.sleep() with WebDriverWait (from selenium.webdriver.support.ui) to wait for specific elements to load, instead of using fixed delays. This makes your scraper faster and more reliable.
  • Use webdriver-manager to automatically manage your ChromeDriver version, so you don't have to update the path manually. Install it with pip install webdriver-manager, then initialize the driver like this:
    from webdriver_manager.chrome import ChromeDriverManager
    driver = webdriver.Chrome(ChromeDriverManager().install(), options=options)
    

内容的提问来源于stack exchange,提问作者mr-kim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:57:10