You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup解析公寓页面时部分同class的房源图片URL随机无法获取

使用BeautifulSoup解析公寓页面时部分同class的房源图片URL随机无法获取

看起来你遇到的问题是部分房源的户型图URL抓不到——明明用了相同的class选择器,结果却时灵时不灵对吧?我帮你分析下可能的原因,再给几个可行的修复方案:

可能的问题根源

  • 图片懒加载机制:很多网站会采用懒加载策略,只有当元素滚动到视口范围内时,才会把真实图片地址填充到src属性中。你当前的Selenium滚动逻辑可能没触发出部分房源图片的加载,导致抓取时src为空、是占位图地址,甚至真实地址可能存在data-src这类备用属性而非src。
  • DOM结构的细微差异:虽然表面上class名称一致,但部分房源的图片标签可能存在嵌套结构、隐藏状态的差异,导致BeautifulSoup无法匹配到目标img标签。
  • 页面快照时机不对:你是直接获取driver.page_source交给BeautifulSoup解析,但可能在获取源码时,部分图片的DOM还未完全渲染完成。

针对性修复方案

方案1:强制触发所有图片的懒加载

在获取页面源码前,执行一段JavaScript,强制将所有懒加载图片的真实地址填充到src中,确保静态解析时能拿到有效地址:

# 在点击完所有「显示更多」按钮后,添加这段代码
print("Triggering all lazy-loaded images...")
driver.execute_script("""
    // 遍历所有带有data-src属性的图片,替换为真实地址
    document.querySelectorAll('img[data-src]').forEach(img => {
        img.src = img.dataset.src;
    });
""")
# 预留1-2秒等待地址替换完成
time.sleep(2)

# 再获取页面源码
container = driver.page_source
driver.quit()

同时,你可以在点击完最后一个「显示更多」后,把页面滚动到底部并等待几秒,确保所有房源都进入过视口。

方案2:扩展图片属性的检查范围

很多懒加载图片会把真实地址存在data-src、data-original这类属性里,而非直接存在src中。修改你的图片提取逻辑,多检查几个常见的懒加载属性:

# 替换原有的图片提取代码块
plan_img_tag = result.find('img', class_="v-img__img v-img__img--contain")
if plan_img_tag:
    # 依次检查src、data-src、data-original属性,取第一个非空值
    plan_url = plan_img_tag.get('src') or plan_img_tag.get('data-src') or plan_img_tag.get('data-original')
    if plan_url:
        if plan_url.startswith('/'):
            plan_url = f"https://etalongroup.ru{plan_url}"
    else:
        plan_url = "No plan URL"
        print(f"No images found for apartment: {title}")
else:
    plan_url = "No plan URL"
    print(f"No images found for apartment: {title}")

方案3:直接用Selenium提取元素属性(最可靠)

既然已经用了Selenium,不如直接通过它来提取每个房源的属性——Selenium操作的是渲染后的实时DOM,能避开静态解析的局限性:

# 替换原有的解析流程,不再使用BeautifulSoup
apartments = []
print("Extracting data via Selenium...")
result_containers = driver.find_elements(By.CSS_SELECTOR, "div.bg-white.relative")

for result in result_containers:
    # 提取链接
    link_tag = result.find_element(By.TAG_NAME, 'a')
    link = f"https://etalongroup.ru{link_tag.get_attribute('href')}"
    
    # 提取价格
    try:
        price_tag = result.find_element(By.CLASS_NAME, "th-h2")
        price = price_tag.text.strip()
    except:
        price = "No price"
    
    # 提取标题
    try:
        title_tag = result.find_element(By.CLASS_NAME, "th-h4")
        title = title_tag.text.strip()
    except:
        title = "No title"
    
    # 提取户型图URL
    plan_url = "No plan URL"
    try:
        plan_img_tag = result.find_element(By.CLASS_NAME, "v-img__img.v-img__img--contain")
        # 多属性检查
        plan_url = plan_img_tag.get_attribute('src') or plan_img_tag.get_attribute('data-src')
        if plan_url and plan_url.startswith('/'):
            plan_url = f"https://etalongroup.ru{plan_url}"
    except:
        print(f"No images found for apartment: {title}")
    
    # 提取其他字段(面积、楼层等,同样用Selenium方式)
    area = "No area"
    floor = "No floor"
    completion_year = "No completion year"
    discount = "No discount"
    
    try:
        th_b1_tags = result.find_elements(By.CLASS_NAME, "th-b1-regular")
        for tag in th_b1_tags:
            text = tag.text.strip()
            if "м²" in text and "|" in text:
                area, floor = text.split(" | ")
            elif "м²" in text:
                area = text
            elif "этаж" in text:
                floor = text
            elif "квартал" in text or "год" in text:
                completion_year = text
            elif "%" in text:
                discount = text
    except:
        pass
    
    # 组装公寓数据
    apartment = {
        "link": link,
        "price": price,
        "title": title,
        "area": area,
        "floor": floor,
        "completion_year": completion_year.strip(),
        "discount": discount,
        "plan_url": plan_url
    }
    apartments.append(apartment)

# 关闭driver
driver.quit()

推荐先尝试方案1+方案2的组合,这个改动最小,能解决大部分懒加载导致的问题。如果还是有遗漏,再切换到方案3——直接用Selenium提取元素属性的方式,能最大程度保证数据的准确性,毕竟它是直接和浏览器渲染后的DOM交互的。

备注:内容来源于stack exchange,提问作者Danny Mxxre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.15 03:43:06