如何用Python提取Google Docs中的评论及对应文本?
解决Google Docs API获取评论为空的问题
问题根源
你当前通过documents().get()接口获取文档数据,但该接口不会返回评论信息,必须使用专门的documents.comments.list()接口才能获取评论列表。
修正后的代码
from googleapiclient.discovery import build from googleapiclient.http import Http import nltk from nltk.tokenize import sent_tokenize # 确保下载分句模型 nltk.download('punkt') def read_comments(comments): comment_data = [] for comment in comments: # 提取评论内容 content = comment['content'] # 提取被评论文本的位置信息 comment_range = comment.get('range', {}) comment_data.append({ 'comment': content, 'start_index': comment_range.get('startIndex'), 'end_index': comment_range.get('endIndex') }) return comment_data def extract_commented_text(doc_content, start_idx, end_idx): # 遍历文档内容,提取指定索引范围内的文本 text = '' for element in doc_content: if 'paragraph' in element: for run in element['paragraph']['elements']: if 'textRun' in run: run_text = run['textRun']['content'] run_start = run['startIndex'] run_end = run_start + len(run_text) # 计算当前文本段与目标范围的交集 overlap_start = max(run_start, start_idx) overlap_end = min(run_end, end_idx) if overlap_start < overlap_end: # 截取对应部分文本 text += run_text[overlap_start - run_start : overlap_end - run_start] return text def main(): credentials = get_credentials() http = credentials.authorize(Http()) docs_service = build( 'docs', 'v1', http=http, discoveryServiceUrl=DISCOVERY_DOC) # 获取文档内容 doc = docs_service.documents().get(documentId=DOCUMENT_ID_2).execute() doc_content = doc.get('body').get('content') # 正确调用评论列表接口 comments_response = docs_service.documents().comments().list( documentId=DOCUMENT_ID_2 ).execute() comments_data = read_comments(comments_response.get('comments', [])) # 处理并输出评论及对应文本 for item in comments_data: comment_text = item['comment'] start_idx = item['start_index'] end_idx = item['end_index'] # 提取被评论的文本 commented_text = extract_commented_text(doc_content, start_idx, end_idx) if start_idx and end_idx else '无法获取对应文本' print(f"评论内容: {comment_text}") print(f"被评论文本: {commented_text}") # 分句处理 sentences = sent_tokenize(comment_text) for sentence in sentences: formatted_sentence = "{This is a PB}" + sentence + "{This is a PB}" print(formatted_sentence) print("---") if __name__ == '__main__': main()
关键说明
- 获取评论的正确接口:使用
documents.comments.list()替代documents().get(),该接口专门用于获取文档的评论集合。 - 提取被评论文本:每个评论对象包含
range字段,其中startIndex和endIndex标记了被评论内容在文档中的位置,通过遍历文档的段落元素,可以精准提取对应范围内的文本。 - 权限注意:确保你的凭证(credentials)拥有
https://www.googleapis.com/auth/documents.readonly或更高的权限,否则会无法获取评论数据。
内容的提问来源于stack exchange,提问作者Soham Deshpande
相关产品推荐
相关产品推荐

