You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python Selenium爬取多网页标签时如何处理部分页面标签缺失问题

实现方案

核心调整点

  • 每次进入新页面的循环时,先将当前页的字段变量初始化为空值,避免上一页的结果残留到下一页
  • 简化标签取值逻辑,直接匹配特征文本对应的字段,匹配不到的字段保持默认空值
  • 优化driver启动关闭逻辑,避免重复启动driver浪费资源

修正后代码

from selenium import webdriver
import time
import random

par2 = [
    'https://www.inmuebles24.com/propiedades/bosques-de-las-lomas-departamento-a-la-venta-en-bosque-62355654.html',
    'https://www.inmuebles24.com/propiedades/tu-mejor-lugar-en-tulum-hyd-60498140.html'
]
links = []
titulo = []
sup_total = []
superficie_cons = []
baños = []

# driver统一初始化,避免每次循环重复启动,效率更高
driver = webdriver.Chrome('C:/driver/chromedriver.exe')

for pares in par2:
    # 每次处理新页面先重置当前页所有字段为空
    current_sup_total = ""
    current_sup_cons = ""
    current_banos = ""
    
    driver.get(pares)
    time.sleep(random.uniform(1,4))
    
    # 抓取标题
    tit = driver.find_element_by_xpath('//*[@class="section-title"]').text
    titulo.append(tit)
    links.append(pares)
    
    # 抓取所有特征标签
    etiquetas = driver.find_elements_by_xpath('.//*[@class="icon-feature"]')
    for etiq in etiquetas:
        try:
            valores = etiq.text.strip().lower()
            if 'total' in valores:
                current_sup_total = valores
            elif 'construido' in valores:
                current_sup_cons = valores
            elif 'baños' in valores:
                current_banos = valores
        except:
            continue
    
    # 匹配完成后写入列表,未匹配到的字段自动保留空值
    sup_total.append(current_sup_total)
    superficie_cons.append(current_sup_cons)
    baños.append(current_banos)

# 全部爬取完成后关闭driver
driver.quit()

扩展说明

如果需要新增抓取的字段,只需要先定义对应存储列表,然后在循环内新增对应字段的初始空值,再加elif分支匹配特征关键词即可,不存在的字段会自动留空存入列表。


内容的提问来源于stack exchange,提问作者Ramiro Guzmán

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 02:48:01