使用Python Selenium爬取多网页标签时如何处理部分页面标签缺失问题
实现方案
核心调整点
- 每次进入新页面的循环时,先将当前页的字段变量初始化为空值,避免上一页的结果残留到下一页
- 简化标签取值逻辑,直接匹配特征文本对应的字段,匹配不到的字段保持默认空值
- 优化driver启动关闭逻辑,避免重复启动driver浪费资源
修正后代码
from selenium import webdriver import time import random par2 = [ 'https://www.inmuebles24.com/propiedades/bosques-de-las-lomas-departamento-a-la-venta-en-bosque-62355654.html', 'https://www.inmuebles24.com/propiedades/tu-mejor-lugar-en-tulum-hyd-60498140.html' ] links = [] titulo = [] sup_total = [] superficie_cons = [] baños = [] # driver统一初始化,避免每次循环重复启动,效率更高 driver = webdriver.Chrome('C:/driver/chromedriver.exe') for pares in par2: # 每次处理新页面先重置当前页所有字段为空 current_sup_total = "" current_sup_cons = "" current_banos = "" driver.get(pares) time.sleep(random.uniform(1,4)) # 抓取标题 tit = driver.find_element_by_xpath('//*[@class="section-title"]').text titulo.append(tit) links.append(pares) # 抓取所有特征标签 etiquetas = driver.find_elements_by_xpath('.//*[@class="icon-feature"]') for etiq in etiquetas: try: valores = etiq.text.strip().lower() if 'total' in valores: current_sup_total = valores elif 'construido' in valores: current_sup_cons = valores elif 'baños' in valores: current_banos = valores except: continue # 匹配完成后写入列表,未匹配到的字段自动保留空值 sup_total.append(current_sup_total) superficie_cons.append(current_sup_cons) baños.append(current_banos) # 全部爬取完成后关闭driver driver.quit()
扩展说明
如果需要新增抓取的字段,只需要先定义对应存储列表,然后在循环内新增对应字段的初始空值,再加elif分支匹配特征关键词即可,不存在的字段会自动留空存入列表。
内容的提问来源于stack exchange,提问作者Ramiro Guzmán
相关产品推荐
相关产品推荐

