You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Selenium收集'+'符号前后的多张图片

解决方案:区分"+"前后的图片

核心思路

通过定位h3标签内的"+"符号节点,遍历子节点时以该节点为分界,分别收集前后的图片地址,无论前后图片数量动态变化都能适配。

具体实现(分语言示例)

JavaScript 前端处理

  1. 定位"+"符号节点
    从目标h3的子节点中筛选出内容为"+"的文本节点:

    // 先获取目标h3标签(可根据实际场景调整选择器)
    const targetH3 = document.querySelector('h3');
    // 找到内容为"+"的文本节点
    const plusSeparator = Array.from(targetH3.childNodes).find(node => 
      node.nodeType === Node.TEXT_NODE && node.textContent.trim() === '+'
    );
    
  2. 分类收集前后图片地址
    遍历h3子节点,以"+"节点为分界,分别存入两个数组:

    const imgsBeforePlus = [];
    const imgsAfterPlus = [];
    let hasPassedPlus = false;
    
    Array.from(targetH3.childNodes).forEach(node => {
      // 遇到"+"节点时切换标记
      if (node === plusSeparator) {
        hasPassedPlus = true;
        return;
      }
      // 收集img节点的src
      if (node.tagName === 'IMG') {
        hasPassedPlus ? imgsAfterPlus.push(node.src) : imgsBeforePlus.push(node.src);
      }
    });
    

Python 后端/爬虫处理

如果是用爬虫解析HTML,以BeautifulSoup为例:

from bs4 import BeautifulSoup

# 假设已获取到包含目标h3的HTML内容
soup = BeautifulSoup(html_content, 'html.parser')
target_h3 = soup.find('h3')

imgs_before_plus = []
imgs_after_plus = []
plus_found = False

for child in target_h3.children:
    # 匹配"+"符号(如果"+"在span等标签内,需调整判断逻辑)
    if child.string and child.string.strip() == '+':
        plus_found = True
        continue
    # 收集img的src属性
    if child.name == 'img':
        if plus_found:
            imgs_after_plus.append(child['src'])
        else:
            imgs_before_plus.append(child['src'])

后续处理

拿到分类后的图片地址数组后,即可分别进行下载:

  • 给"+"前的图片添加标识前缀(如before_),"+"后的添加after_
  • 按分类存储到不同目录或标记文件名,方便后续识别

内容的提问来源于stack exchange,提问作者Rodhad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 07:37:20