如何从WikiMedia Rest API返回的JSON中剔除.mw-parser-output内容
解决WikiMedia Mobile Rest API返回JSON中移除.mw-parser-output相关CSS的方案
你拿到接口返回的JSON后,首先定位存储页面HTML内容的字段,通常是sections数组下每个条目对应的text字段,所有和.mw-parser-output相关的样式、标签都包含在这个字段的字符串内,可通过以下两种方案处理:
方案1:正则替换(通用快速方案)
适合轻量场景,直接用正则匹配目标内容批量替换:
- 移除外层.mw-parser-output容器标签:匹配
<div class="mw-parser-output">(.*?)</div>,提取分组1的有效内容 - 移除关联style标签:匹配
<style[^>]*>.*?\.mw-parser-output.*?</style>,直接替换为空字符串 - 清理残留的mw相关类名:匹配
class="[^"]*mw-[^"]*",替换为空即可
Python示例代码:
import re import json # 假设response_json为接口返回的JSON对象 page_html = response_json["sections"][0]["text"] # 移除外层容器 page_html = re.sub(r'<div class="mw-parser-output">(.*?)</div>', r'\1', page_html, flags=re.S) # 移除关联样式标签 page_html = re.sub(r'<style[^>]*>.*?\.mw-parser-output.*?</style>', '', page_html, flags=re.S) # 清理残留mw类名 page_html = re.sub(r'class="[^"]*mw-[^"]*"', '', page_html) # 处理完成后写回原JSON response_json["sections"][0]["text"] = page_html
其他语言环境逻辑一致,正则规则无需修改,适配对应语言的字符串替换方法即可。
方案2:DOM解析处理(更稳定,避免正则误匹配)
适合对内容准确性要求高的场景,先解析为DOM树再精准操作节点:
- 将HTML字符串解析为DOM实例
- 找到所有含
mw-parser-output类名的节点,将其内部子节点全部迁移到父节点下后删除原节点 - 遍历所有style标签,删除内容包含
.mw-parser-output的标签 - 遍历所有元素节点,清除class属性中所有带
mw-前缀的类名 - 最后将DOM重新序列化为HTML字符串写回原JSON
JavaScript示例代码(Node.js环境需先引入jsdom库,浏览器环境可直接用document):
const dom = new JSDOM(page_html); const doc = dom.window.document; // 处理mw-parser-output容器 const parserOutput = doc.querySelector('.mw-parser-output'); if (parserOutput) { while (parserOutput.firstChild) { parserOutput.parentNode.insertBefore(parserOutput.firstChild, parserOutput); } parserOutput.remove(); } // 移除关联style标签 doc.querySelectorAll('style').forEach(style => { if (style.textContent.includes('.mw-parser-output')) style.remove(); }); // 清理所有mw前缀类名 doc.querySelectorAll('*[class*="mw-"]').forEach(el => { el.classList.forEach(className => { if (className.startsWith('mw-')) el.classList.remove(className); }); if (el.classList.length === 0) el.removeAttribute('class'); }); // 序列化回HTML字符串 const cleanedHtml = doc.body.innerHTML;
内容的提问来源于stack exchange,提问作者Fiaz Ali
相关产品推荐
相关产品推荐

