如何用Python Selenium收集'+'符号前后的多张图片
解决方案:区分"+"前后的图片
核心思路
通过定位h3标签内的"+"符号节点,遍历子节点时以该节点为分界,分别收集前后的图片地址,无论前后图片数量动态变化都能适配。
具体实现(分语言示例)
JavaScript 前端处理
定位"+"符号节点
从目标h3的子节点中筛选出内容为"+"的文本节点:// 先获取目标h3标签(可根据实际场景调整选择器) const targetH3 = document.querySelector('h3'); // 找到内容为"+"的文本节点 const plusSeparator = Array.from(targetH3.childNodes).find(node => node.nodeType === Node.TEXT_NODE && node.textContent.trim() === '+' );分类收集前后图片地址
遍历h3子节点,以"+"节点为分界,分别存入两个数组:const imgsBeforePlus = []; const imgsAfterPlus = []; let hasPassedPlus = false; Array.from(targetH3.childNodes).forEach(node => { // 遇到"+"节点时切换标记 if (node === plusSeparator) { hasPassedPlus = true; return; } // 收集img节点的src if (node.tagName === 'IMG') { hasPassedPlus ? imgsAfterPlus.push(node.src) : imgsBeforePlus.push(node.src); } });
Python 后端/爬虫处理
如果是用爬虫解析HTML,以BeautifulSoup为例:
from bs4 import BeautifulSoup # 假设已获取到包含目标h3的HTML内容 soup = BeautifulSoup(html_content, 'html.parser') target_h3 = soup.find('h3') imgs_before_plus = [] imgs_after_plus = [] plus_found = False for child in target_h3.children: # 匹配"+"符号(如果"+"在span等标签内,需调整判断逻辑) if child.string and child.string.strip() == '+': plus_found = True continue # 收集img的src属性 if child.name == 'img': if plus_found: imgs_after_plus.append(child['src']) else: imgs_before_plus.append(child['src'])
后续处理
拿到分类后的图片地址数组后,即可分别进行下载:
- 给"+"前的图片添加标识前缀(如
before_),"+"后的添加after_ - 按分类存储到不同目录或标记文件名,方便后续识别
内容的提问来源于stack exchange,提问作者Rodhad
相关产品推荐
相关产品推荐

