如何从RSS文件提取content:encoded内容并生成带格式的Google Docs
RSS内容提取转Google Docs解决方案
原有代码问题
- 直接用正则匹配XML节点稳定性极低,默认贪婪匹配规则只会提取到最后一个
<content:encoded>节点内容,无法对应匹配每个<item>的标题和内容 - 提取的内容缺少基础HTML结构声明,转换为Docs时容易出现排版丢失
- 未遍历所有
<item>节点,只能处理单条内容
前置操作要求
在Apps Script编辑器左侧「服务」菜单中点击「添加服务」,找到Drive API后确认添加,否则Drive.Files接口调用会报错。
修正后可用代码
function rssToGoogleDocs() { const rssUrl = "https://cdn.feedcontrol.net/1830/2857-8L5Ntr0N5l5Xf.xml"; // 抓取RSS内容 const rssContent = UrlFetchApp.fetch(rssUrl).getContentText("UTF-8"); // 用官方XML解析器处理内容,避免正则匹配问题 const xmlDocument = XmlService.parse(rssContent); const rootElement = xmlDocument.getRootElement(); const channelElement = rootElement.getChild("channel"); const items = channelElement.getChildren("item"); // 拼接完整HTML结构,保证排版正常 let fullHtml = `<!DOCTYPE html> <html> <head> <meta charset="UTF-8"> </head> <body>`; // 遍历所有item节点,提取标题和对应内容 items.forEach(item => { const title = item.getChildText("title"); // 处理content:encoded带命名空间的节点 const contentNs = XmlService.getNamespace("http://purl.org/rss/1.0/modules/content/"); const content = item.getChildText("encoded", contentNs); // 把标题和内容拼入HTML,标题加h1标签区分 fullHtml += `<h1>${title}</h1>${content}<hr style="margin:30px 0;border:1px solid #eee;">`; }); fullHtml += "</body></html>"; // 转换为Google Docs const htmlBlob = Utilities.newBlob(fullHtml, "text/html", "RSS导出内容.html"); const convertedDoc = Drive.Files.insert( { title: "RSS内容导出文档" }, htmlBlob, { convert: true } ); // 输出生成的DocsID,可直接拼接地址打开 Logger.log(`生成的文档ID:${convertedDoc.id}`); }
扩展说明
- 如果需要为每个
<item>生成单独的Google Docs文件,只需将文件创建逻辑移到items遍历循环内,文件名设置为对应title即可 - 生成的文档会完整保留原RSS内容里的所有HTML标签、排版、图片链接等结构,和原内容展示效果一致
- 脚本运行需要授予网络访问、Google Drive操作的相关权限,按照提示授权即可正常运行
内容的提问来源于stack exchange,提问作者Hạnh Khánh
相关产品推荐
相关产品推荐

