从JIRA结构化评论数据中提取并拼接text字段内容
问题描述
我从JIRA拉取了一批如下格式的评论数据:
comment text is: [{'type': 'paragraph', 'content': [{'type': 'text', 'text': 'In conversation with the customer '}, {'type': 'mention', 'attrs': {'id': '04445152', 'text': '@Kev', 'accessLevel': ''}}, {'type': 'text', 'text': ' Text 123}]}] comment text is: [{'type': 'paragraph', 'content': [{'type': 'text', 'text': '@xyz Text abc'}]}] comment text is: [{'type': 'paragraph', 'content': [{'type': 'mention', 'attrs': {'id': '3445343', 'text': '@Hey', 'accessLevel': ''}}, {'type': 'text', 'text': ' FYI'}]}] comment text is:[{'content': [{'text': 'Output: ', 'type': 'text'}, {'type': 'hardBreak'}, {'type': 'hardBreak'}, {'text': "New Text goes here", 'type': 'text'}], 'type': 'paragraph'}]
我需要提取所有带有text键对应的数据,并且将同一条评论中的多个此类值进行拼接。预期输出如下:
In conversation with the customer @Kev Text 123 @xyz Text abc @Hey FYI Output: New Text goes here
解决方案
用Python可以快速实现这个需求,核心思路是解析每条评论的结构化数据,遍历内容节点提取所有text字段值,忽略换行类节点后拼接成完整文本。
代码示例
import ast # 模拟从JIRA拉取的原始评论数据 raw_comments = [ "[{'type': 'paragraph', 'content': [{'type': 'text', 'text': 'In conversation with the customer '}, {'type': 'mention', 'attrs': {'id': '04445152', 'text': '@Kev', 'accessLevel': ''}}, {'type': 'text', 'text': ' Text 123'}]}]", "[{'type': 'paragraph', 'content': [{'type': 'text', 'text': '@xyz Text abc'}]}]", "[{'type': 'paragraph', 'content': [{'type': 'mention', 'attrs': {'id': '3445343', 'text': '@Hey', 'accessLevel': ''}}, {'type': 'text', 'text': ' FYI'}]}]", "[{'content': [{'text': 'Output: ', 'type': 'text'}, {'type': 'hardBreak'}, {'type': 'hardBreak'}, {'text': \"New Text goes here\", 'type': 'text'}], 'type': 'paragraph'}]" ] def extract_comment_text(comment_data): # 将字符串转为Python可处理的列表/字典 data = ast.literal_eval(comment_data) text_parts = [] # 获取段落内的所有内容节点 content = data[0]['content'] for item in content: # 提取普通文本节点的text值 if item.get('type') == 'text': text_parts.append(item['text']) # 提取提及节点attrs中的text值 elif item.get('type') == 'mention' and 'attrs' in item: text_parts.append(item['attrs']['text']) # 跳过换行节点 elif item.get('type') == 'hardBreak': continue # 拼接所有文本片段 return ''.join(text_parts) # 处理每条评论并输出结果 for comment in raw_comments: print(extract_comment_text(comment)) print() # 每条评论间空一行
运行结果
In conversation with the customer @Kev Text 123 @xyz Text abc @Hey FYI Output: New Text goes here
内容的提问来源于stack exchange,提问作者Kevin Nash
相关产品推荐
相关产品推荐

