Node.js网页爬虫:替换HTML中的Unicode字符以获取全部链接
解决Unicode转义字符导致的链接识别问题
问题核心是页面部分标签或链接内容被转成了Unicode转义序列(\u003c对应<,\u003e对应>),而normalize()仅处理Unicode规范化(比如统一重音字符的编码形式),对这类转义序列完全无效,必须先手动替换再交给cheerio解析。
最优处理方案是在加载cheerio前,对原始HTML文本做全局替换:
import fetch from 'node-fetch' import { load } from 'cheerio' const response = await fetch('https://www.kapow.com/') let html = await response.text() // 全局替换Unicode转义的<和>为真实HTML标签字符 html = html.replace(/\\u003c/g, '<').replace(/\\u003e/g, '>') const $ = load(html) const links = $('a') .map((i, link) => link.attribs.href) .get() console.log(links)
补充说明:
- 正则的
g标志确保替换所有出现的转义字符,避免漏处理。 - 如果页面还存在其他类似转义字符(比如
\u0022对应双引号),可在替换链中追加规则,例如.replace(/\\u0022/g, '"')。
内容的提问来源于stack exchange,提问作者Tosh Velaga
相关产品推荐
相关产品推荐

