You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从div中仅提取文本和img标签并忽略其他HTML标签

问题原因

你当前编写的正则逻辑是提取所有<>标签包裹范围之外的文本,未对img标签做例外保留处理,因此所有HTML标签包括img都会被直接过滤。另外正则本身并不适合处理结构化的HTML内容,遇到标签换行、属性含特殊字符等场景很容易出现匹配错误。

推荐解决方案

优先使用DOM节点遍历的方案,稳定性远高于正则实现,逻辑是递归遍历目标div下的所有节点,仅保留文本节点内容和img标签,其余标签直接跳过,遍历其内部子节点即可。

jQuery实现代码

$(document).ready(function(){
  const targetDiv = $('#div-parent');
  const resultArr = [];

  // 递归遍历节点方法
  function processNode(node) {
    // 元素节点判断
    if (node.nodeType === 1) {
      // 遇到img标签直接保留完整内容
      if (node.tagName.toLowerCase() === 'img') {
        resultArr.push(node.outerHTML);
        return;
      }
      // 其他元素节点遍历子节点
      Array.from(node.childNodes).forEach(childNode => processNode(childNode));
    } 
    // 文本节点判断,保留非空内容
    else if (node.nodeType === 3) {
      const textContent = node.textContent.trim();
      if (textContent) {
        resultArr.push(textContent);
      }
    }
  }

  // 启动遍历
  Array.from(targetDiv[0].childNodes).forEach(child => processNode(child));
  // 输出结果,可按需拼接为字符串
  console.log(resultArr);
});

运行效果

针对你提供的HTML示例,输出结果会按顺序保留所有文本和img标签,完全符合需求:

[
  "hello world",
  "<img src=\"\" alt=\"image\"/>",
  "here we go",
  "test 1",
  "<img src=\"\" alt=\"image1\"/>",
  "list 1",
  "test 2",
  "<img src=\"\" alt=\"image2\"/>",
  "list 2",
  // 剩余内容按相同规则依次输出
]
可选正则方案(仅适合简单场景)

如果一定要用正则实现,可以用排除匹配的思路,移除所有非img的HTML标签:

const rawHtml = $('#div-parent').html();
const result = rawHtml.replace(/<(?!img\s|\/img)[^>]+>/gi, '').trim();
console.log(result);

注意:该方案存在边界风险,若标签属性内包含>、img标签写法不规范时可能出现匹配错误,仅建议临时简单场景使用。

内容的提问来源于stack exchange,提问作者Amal Ps

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 23:21:02