使用Python+BeautifulSoup爬取edgeprop获取地块面积数据失败如何解决
问题根因
- 目标页面数据由JavaScript动态渲染生成,
requests库仅能获取静态HTML源码,无法拿到渲染后加载的Land Size相关数据 - 现有代码使用的
detail-title__text类选择器不匹配实际页面元素定位规则
修复后实现方案
推荐使用Selenium模拟浏览器渲染页面后再提取数据,具体实现代码如下:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup import time # 配置无头浏览器模式 chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-blink-features=AutomationControlled") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") query_string = 'https://www.edgeprop.sg/condo-apartment/aquarius-by-the-park' driver = webdriver.Chrome(options=chrome_options) driver.get(query_string) # 等待页面渲染完成,可根据网络情况调整时长 time.sleep(3) soup = BeautifulSoup(driver.page_source, 'html.parser') driver.quit() try: # 定位概览区域所有信息项 overview_items = soup.find_all("div", class_="detail-property-oview__item") for item in overview_items: title = item.find("p", class_="title") if title and "Land Size" in title.text: # 提取sqm对应数值 land_size = item.find("p", class_="value").text.split("(")[1].replace(" sqm)", "").strip() print("Land Size (sqm) 是:", land_size) break except Exception as e: print("提取失败:", e)
依赖说明
运行代码前需要先安装对应依赖:pip install selenium beautifulsoup4,同时需要安装和本地Chrome版本匹配的ChromeDriver。
内容的提问来源于stack exchange,提问作者Chris lee
相关产品推荐
相关产品推荐

