如何用Cheerio从XML的figure标签中提取img的src属性
解决Cheerio XML模式下提取RSS item中figure内img的src问题
问题背景
需要从RSS Feed的item标签内的figure标签中提取img的src属性,已使用Cheerio开启XML模式遍历item:
const $ = cheerio.load(response.data, { xmlMode: true }); $("item").each((i, item) => {
目标标签结构
目标内容的HTML结构如下:
<figure class="wp-block-image size-large"> <img decoding="async" loading="lazy" width="800" height="572" src="http://wmcmuaythai.org/wp-content/uploads/2023/04/WhatsApp-Image-2023-04-07-at-3.18.13-PM-2-800x572.jpeg" alt="" class="wp-image-43535" srcset="http://wmcmuaythai.org/wp-content/uploads/2023/04/WhatsApp-Image-2023-04-07-at-3.18.13-PM-2-800x572.jpeg 800w, http://wmcmuaythai.org/wp-content/uploads/2023/04/WhatsApp-Image-2023-04-07-at-3.18.13-PM-2-350x250.jpeg 350w, http://wmcmuaythai.org/wp-content/uploads/2023/04/WhatsApp-Image-2023-04-07-at-3.18.13-PM-2-768x549.jpeg 768w, http://wmcmuaythai.org/wp-content/uploads/2023/04/WhatsApp-Image-2023-04-07-at-3.18.13-PM-2.jpeg 1024w" sizes="(max-width: 800px) 100vw, 800px" /> </figure>
无效尝试
以下查询语句均返回undefined或不符合预期:
$(item).find("figure").find("img").attr("src") $(item).find("img").attr("src") $(item).find("figure").children().find("img").attr("src") $(item).find("figure").first().find("img").attr("src")
解决方案
核心原因
RSS中item的内容通常被包裹在CDATA块内,Cheerio开启XML模式时会将CDATA内容视为纯文本,不会解析其中的HTML结构,因此直接通过find无法定位到figure和img标签。
代码实现
先提取item中内容标签(如description或content:encoded)的文本内容,再用Cheerio以HTML模式重新解析该内容,即可正常提取img的src:
const $ = cheerio.load(response.data, { xmlMode: true }); $("item").each((i, item) => { // 提取item中的内容块,根据RSS结构选择description或content:encoded const contentText = $(item).find("description").text(); // 以HTML模式加载内容,解析其中的HTML结构 const content$ = cheerio.load(contentText); // 定位figure下的img并获取src属性 const imgSrc = content$("figure.wp-block-image img").attr("src"); // 输出或使用提取到的src console.log("图片地址:", imgSrc); });
补充说明
- 如果你的RSS使用
content:encoded标签存储富内容,将代码中的description替换为content:encoded即可。 - 若存在多个
figure标签,可使用.each()遍历或调整选择器(如figure:first)来获取目标图片。
内容的提问来源于stack exchange,提问作者Daniel Tolentino
相关产品推荐
相关产品推荐

