如何匹配动态单字母前缀的div class实现网页数据爬取?
Hey there, great job simplifying your code already! The dynamic class prefixes (xl-, l-, m-) you're facing are super common in responsive websites, and we can make your scraper even more robust and maintainable than your current approach.
核心问题分析
Your optimized code uses lists of class names to match variations, but this still requires updating the list if new prefixes (like s- for small screens) are added later. Instead, we can match elements based on the fixed suffix part of the class name using CSS attribute selectors or regex, which is far more flexible.
优化方案1:使用CSS属性包含选择器(推荐)
BeautifulSoup supports CSS selectors, and we can use the *= operator to match elements whose class attribute contains a specific substring. This works perfectly since your target elements all share fixed suffixes:
- Results:
result-+ prefix → match any class starting withresult- - Price:
-price rangePrice→ match any class containing this substring - Surface:
-surface-ch→ match any class containing this substring - Description:
-desc→ match any class containing this substring
Here's the revised code with this approach, plus error handling for missing elements:
import pandas as pd from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup from fake_useragent import UserAgent import time # Setup Chrome options options = Options() options.add_argument("window-size=1400,600") ua = UserAgent() user_agent = ua.random print(user_agent) options.add_argument(f'user-agent={user_agent}') # Initialize driver (note: consider using webdriver-manager for auto chromedriver management) driver = webdriver.Chrome('/Users/raduulea/Documents/chromedriver', options=options) driver.get('https://www.immoweb.be/fr/recherche/immeuble-de-rapport/a-vendre/liege/4000') time.sleep(10) # Initialize data lists Title = [] address = [] price = [] surface = [] desc = [] page = 2 while True: time.sleep(10) html = driver.page_source soup = BeautifulSoup(html, 'html.parser') if page <= 6: # Match any result block with class starting with "result-" results = soup.select('div[class^="result-"]') for result in results: # Extract title title_elem = result.find("div", class_="title-bar-left") Title.append(title_elem.get_text().strip() if title_elem else 'N/A') # Extract address (fixed class name, no variation) addr_elem = result.find("span", class_="result-adress") address.append(addr_elem.get_text().strip() if addr_elem else 'N/A') # Extract price: match any class containing "-price rangePrice" price_elem = result.select_one('[class*="-price rangePrice"]') price.append(price_elem.get_text().strip() if price_elem else 'N/A') # Extract surface: match any class containing "-surface-ch" surface_elem = result.select_one('[class*="-surface-ch"]') surface.append(surface_elem.get_text().strip() if surface_elem else 'N/A') # Extract description: match any class containing "-desc" desc_elem = result.select_one('[class*="-desc"]') desc.append(desc_elem.get_text().strip() if desc_elem else 'N/A') # Navigate to next page if available if len(driver.find_elements_by_css_selector("a.next")) > 0: url = f"https://www.immoweb.be/fr/recherche/immeuble-de-rapport/a-vendre/liege/4000/?page={page}" driver.get(url) page += 1 else: break else: break # Save to CSV df = pd.DataFrame({ "Title": Title, "Address": address, "Price": price, "Surface": surface, "Description": desc }) df.to_csv("immoweb_scrap_optimized.csv", index=False) driver.quit()
关键优化点
- Flexible element matching: No more hardcoding all possible class prefixes — this will automatically work for any new prefixes the site adds (like
s-for mobile). - Error handling: Added checks for
Noneelements to prevent your scraper from crashing if a field is missing on a listing. - Cleaner syntax: Used f-strings for URL construction, and simplified class matching logic.
优化方案2:使用正则表达式匹配Class
If you prefer regex, you can use BeautifulSoup's class_ parameter with a regex pattern to match the dynamic classes:
import re # Match result blocks results = soup.find_all("div", class_=re.compile(r'result-(xl|l|m)')) # Match price element price_elem = result.find("div", class_=re.compile(r'(xl|l|m)-price rangePrice'))
This works well too, but the CSS selector approach is more concise for this use case.
额外建议
- Replace
time.sleep()with WebDriverWait (fromselenium.webdriver.support.ui) to wait for specific elements to load, instead of using fixed delays. This makes your scraper faster and more reliable. - Use
webdriver-managerto automatically manage your ChromeDriver version, so you don't have to update the path manually. Install it withpip install webdriver-manager, then initialize the driver like this:from webdriver_manager.chrome import ChromeDriverManager driver = webdriver.Chrome(ChromeDriverManager().install(), options=options)
内容的提问来源于stack exchange,提问作者mr-kim

