React项目中如何移除编辑器内容的HTML标签,提取图片URL与纯文本
React 编辑器内容提取与清理方案
核心处理逻辑
你拿到的编辑器内容是经过HTML实体转义的字符串,不要用正则直接匹配HTML内容,优先用浏览器原生DOMParser解析DOM结构,再提取对应内容,可完美覆盖各种边界异常。
步骤1:封装通用处理方法
// 第一步:解码HTML转义字符 const decodeHtml = (htmlStr) => { const textarea = document.createElement("textarea"); textarea.innerHTML = htmlStr; return textarea.value; }; // 第二步:解析内容、提取文本+图片链接、清理畸形URL const parseEditorContent = (rawContent) => { const decodedContent = decodeHtml(rawContent); const parser = new DOMParser(); const htmlDoc = parser.parseFromString(decodedContent, "text/html"); const result = { textList: [], imageUrlList: [] }; // 提取所有纯文本,自动跳过所有HTML标签 result.textList = htmlDoc.body.innerText .trim() .split(/\n+/) .filter(text => text.trim().length > 0); // 提取所有图片链接,同时清理src中的畸形锚点 const allImgs = htmlDoc.querySelectorAll("img"); allImgs.forEach(img => { // 兼容你示例中img属性写错为scr的异常情况 let rawSrc = img.getAttribute("src") || img.getAttribute("scr"); if (!rawSrc) return; // 如果src被锚点标签包裹,直接提取其中的真实HTTP链接 if (rawSrc.includes("<a") || rawSrc.includes("href")) { const urlMatch = rawSrc.match(/https?:\/\/[^\s"'>]+/); if (urlMatch) rawSrc = urlMatch[0]; } // 过滤无效链接,规则可按需调整 if (rawSrc.startsWith("http")) { result.imageUrlList.push(rawSrc); } }); return result; };
步骤2:在React组件中直接使用
function ContentRenderPage({ rawEditorContent }) { const { textList, imageUrlList } = parseEditorContent(rawEditorContent); return ( <div className="content-wrapper"> {/* 渲染图片 */} {imageUrlList.map((url, index) => ( <img key={index} src={url} alt={`content-img-${index}`} /> ))} {/* 渲染文本 */} {textList.map((text, index) => ( <p key={index}>{text}</p> ))} </div> ); }
可选扩展:保留部分标签结构
如果你不需要完全剥离所有HTML标签,只想移除figure、空p标签这类冗余元素,可以直接清理DOM后用dangerouslySetInnerHTML渲染:
const getCleanHtml = (rawContent) => { const decodedContent = decodeHtml(rawContent); const htmlDoc = new DOMParser().parseFromString(decodedContent, "text/html"); // 移除figure标签,保留内部的img元素 htmlDoc.querySelectorAll("figure").forEach(fig => { const innerImg = fig.querySelector("img"); if (innerImg) fig.parentNode.replaceChild(innerImg, fig); }); // 移除没有内容的空p标签 htmlDoc.querySelectorAll("p").forEach(p => { if (!p.innerText.trim() && !p.querySelector("img,video")) p.remove(); }); return htmlDoc.body.innerHTML; }; // 组件中使用 <div dangerouslySetInnerHTML={{ __html: getCleanHtml(rawEditorContent) }} />
内容的提问来源于stack exchange,提问作者Gabriel
相关产品推荐
相关产品推荐

