如何将含ul/li/sup/sub的HTML字符串转为带上标的字符串数组
解析含sup/sub标签的ul/li HTML字符串为带格式的文本数组
需求说明
将包含ul、li、sup(及sub)标签的HTML字符串,提取所有li标签的内容并转为数组,同时将<sup>标签内的内容替换为对应Unicode上标字符,<sub>标签内容替换为下标字符,保留格式语义。
示例输入
const htmlString = "<ul><li>Uses multiple cameras to display a digital overhead image of the area around your vehicle along with Rear Vision Camera or front views<sup>1</sup></li><li>Front and rear dynamic guidelines laid over the display image assist in parking maneuvers by showing the vehicle's path</li><li>It works at low speeds and may help you park and avoid nearby objects.</li><li>You can select additional views on your camera display</li></ul>";
期望输出
[ "Uses multiple cameras to display a digital overhead image of the area around your vehicle along with Rear Vision Camera or front views¹", "Front and rear dynamic guidelines laid over the display image assist in parking maneuvers by showing the vehicle's path", "It works at low speeds and may help you park and avoid nearby objects.", "You can select additional views on your camera display" ]
实现方案
利用浏览器原生的DOMParser解析HTML,通过DOM操作处理sup/sub标签,最终提取文本内容:
const htmlString = "<ul><li>Uses multiple cameras to display a digital overhead image of the area around your vehicle along with Rear Vision Camera or front views<sup>1</sup></li><li>Front and rear dynamic guidelines laid over the display image assist in parking maneuvers by showing the vehicle's path</li><li>It works at low speeds and may help you park and avoid nearby objects.</li><li>You can select additional views on your camera display</li></ul>"; // 解析HTML为DOM文档 const parser = new DOMParser(); const doc = parser.parseFromString(htmlString, 'text/html'); const liElements = doc.querySelectorAll('li'); // 定义数字与上标字符的映射 const supCharMap = { '1': '¹', '2': '²', '3': '³', '4': '⁴', '5': '⁵', '6': '⁶', '7': '⁷', '8': '⁸', '9': '⁹', '0': '⁰' }; // 定义数字与下标字符的映射(如需处理sub标签) const subCharMap = { '1': '₁', '2': '₂', '3': '₃', '4': '₄', '5': '₅', '6': '₆', '7': '₇', '8': '₈', '9': '₉', '0': '₀' }; // 处理每个li元素 const result = Array.from(liElements).map(li => { // 克隆节点避免修改原DOM结构 const liClone = li.cloneNode(true); // 替换所有sup标签为上标字符 liClone.querySelectorAll('sup').forEach(sup => { const supText = sup.textContent.trim(); const convertedText = supText.split('').map(char => supCharMap[char] || char).join(''); sup.replaceWith(document.createTextNode(convertedText)); }); // 替换所有sub标签为下标字符(可选,根据需求开启) liClone.querySelectorAll('sub').forEach(sub => { const subText = sub.textContent.trim(); const convertedText = subText.split('').map(char => subCharMap[char] || char).join(''); sub.replaceWith(document.createTextNode(convertedText)); }); return liClone.textContent.trim(); }); console.log(result);
方案说明
- DOMParser解析:相比正则表达式,DOM解析能更可靠地处理复杂HTML结构,避免标签嵌套导致的解析错误。
- 标签替换逻辑:通过映射表将数字转为对应Unicode上标/下标字符,支持多数字组合(如
<sup>12</sup>转为¹²)。 - 克隆节点:克隆原li节点进行修改,避免影响原始DOM文档。
内容的提问来源于stack exchange,提问作者Prakhar Mittal
相关产品推荐
相关产品推荐

