如何用Selenium+ChromeDriver精准识别网页图片叠加文本(排除图标、Logo)
解决Selenium定位图片叠加HTML文本并排除图标的方案
核心思路:基于布局特征筛选叠加文本
图片上的HTML叠加文本通常有明确的布局特征,结合排除规则可以精准定位:
- 叠加文本元素一般为
position: absolute,且与目标图片同属一个设置了position: relative的父容器 - 图标/Logo通常带有特定类名、小尺寸、图标字体或明确的ARIA标签,可针对性排除
代码实现(Python + Selenium)
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver = webdriver.Chrome() driver.get("目标网站URL") # 等待页面加载完成,定位所有包含图片的容器 wait = WebDriverWait(driver, 15) image_containers = wait.until(EC.presence_of_all_elements_located((By.XPATH, "//*[.//img]"))) text_image_pairs = [] for container in image_containers: # 获取容器内的图片信息 try: image = container.find_element(By.TAG_NAME, "img") image_src = image.get_attribute("src") image_rect = image.rect except: continue # 跳过无有效图片的容器 # 遍历容器内所有文本候选元素,排除图标/Logo candidate_texts = container.find_elements(By.XPATH, ".//*[not(contains(@class, 'icon')) and not(contains(@class, 'logo')) and not(contains(@aria-label, 'icon')) and not(contains(@aria-label, 'logo'))]") for text_elem in candidate_texts: # 过滤空文本 text = text_elem.text.strip() if not text: continue # 判断是否为绝对定位的叠加文本 if text_elem.value_of_css_property("position") != "absolute": continue # 判断文本是否在图片范围内(避免页面其他绝对定位文本误判) text_rect = text_elem.rect if not (text_rect['x'] >= image_rect['x'] and text_rect['x'] + text_rect['width'] <= image_rect['x'] + image_rect['width'] and text_rect['y'] >= image_rect['y'] and text_rect['y'] + text_rect['height'] <= image_rect['y'] + image_rect['height']): continue # 排除图标字体(如FontAwesome、Material Icons) font_family = text_elem.value_of_css_property("font-family") if "FontAwesome" in font_family or "Material Icons" in font_family or text.startswith("\ue"): continue # 收集有效图文对 text_image_pairs.append({ "image_url": image_src, "overlay_text": text }) # 输出结果 for idx, pair in enumerate(text_image_pairs, 1): print(f"图文对 {idx}:") print(f"图片URL: {pair['image_url']}") print(f"叠加文本: {pair['overlay_text']}") print("---") driver.quit()
优化技巧
- 动态样式适配:如果网站用外部样式表设置定位,直接通过
value_of_css_property获取计算后的样式,比XPATH匹配更可靠 - 调整文本长度阈值:可根据目标网站情况修改文本长度判断(比如
len(text) > 2),进一步排除单字符图标 - Z-index验证:叠加文本的
z-index通常高于图片,可添加判断text_elem.value_of_css_property("z-index") > image.value_of_css_property("z-index")增强准确性 - 处理懒加载图片:若图片是懒加载,需先滚动到容器位置触发加载,再获取
src属性
内容的提问来源于stack exchange,提问作者Christelle Kpairi
相关产品推荐
相关产品推荐

