You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Google API根据标题提取Google文档指定章节内容?

解决方案:用Google Docs API提取指定标题下的内容

Google Drive的导出API(你当前使用的files/export)不支持直接按标题筛选内容,只能导出完整文档。要实现按需提取指定章节的需求,你需要改用Google Docs API——它能返回文档的结构化数据,让你精准定位并提取目标标题下的内容。

实现步骤

  1. 启用Google Docs API:在你的Google Cloud控制台中找到并启用Docs API,确保你的API Key拥有访问该API的权限。
  2. 确保文档公开可访问:将目标Google文档设置为「知道链接的任何人都能查看」,这样无需OAuth授权,只用API Key就能调用接口。

代码实现

下面是你需要的printGoogleDoc函数,它会调用Docs API获取结构化内容,然后提取指定标题下的文本:

async function printGoogleDoc(docID, apiKey, targetHeading) {
  try {
    // 调用Google Docs API获取文档结构化内容
    const response = await fetch(`https://docs.googleapis.com/v1/documents/${docID}?key=${apiKey}`);
    const doc = await response.json();

    let isTargetSection = false;
    let sectionContent = [];
    const targetHeadingLevel = targetHeading.split(' ')[0]; // 提取h2这类层级标识

    // 遍历文档的所有内容块
    for (const element of doc.body.content) {
      if (element.paragraph) {
        const paragraph = element.paragraph;
        // 判断当前段落是否是标题
        const isHeading = paragraph.namedStyleType && paragraph.namedStyleType.startsWith('HEADING_');
        if (isHeading) {
          // 匹配标题层级和文本
          const headingLevel = `h${paragraph.namedStyleType.split('_')[1]}`;
          const headingText = paragraph.elements.map(e => e.textRun.content).join('').trim();
          
          if (headingLevel === targetHeadingLevel && headingText === targetHeading.slice(targetHeadingLevel.length + 1).trim()) {
            // 找到目标标题,开始收集内容
            isTargetSection = true;
          } else if (isTargetSection && headingLevel === targetHeadingLevel) {
            // 遇到同层级的下一个标题,停止收集
            break;
          }
        }

        // 如果处于目标章节,收集非标题的内容文本
        if (isTargetSection && !paragraph.namedStyleType?.startsWith('HEADING_')) {
          const text = paragraph.elements.map(e => e.textRun?.content || '').join('').trim();
          if (text) sectionContent.push(text);
        }
      }
    }

    // 输出或返回目标章节内容
    console.log(sectionContent.join('\n'));
    return sectionContent.join('\n');
  } catch (error) {
    console.log("获取内容出错:", error);
    throw error;
  }
}

// 使用示例:提取"h2 sub title"下的内容
printGoogleDoc('你的文档ID', '你的API Key', 'h2 sub title');

代码说明

  • 调用Docs API的documents.get接口,返回的文档数据包含每个段落的样式类型(比如HEADING_2对应h2)和文本内容。
  • 通过遍历内容块,先定位到目标标题,然后开始收集后续的非标题文本,直到遇到同层级的下一个标题为止。
  • 适配了你的文档结构,能精准提取指定h2标题下的内容。

内容的提问来源于stack exchange,提问作者Panthera

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 07:47:24