如何从protobuf文本中提取text字段值 支持可选pinned_field匹配
将protobuf转换为字符串做文本提取的思路是可行的,你现有代码的核心问题是错误调用了re.escape()转义了正则的特殊语法,导致匹配规则完全失效,调整正则规则和匹配逻辑即可实现需求。
基础需求实现(仅提取text内容)
预期输出:["5 Best Products","10 Best Products 2021"]
代码如下:
import re def extract_text_from_proto(proto_string): # 正则匹配text字段双引号内的内容 regex = r'text:\s*"([^"]+)"' result = [m.group(1) for m in re.finditer(regex, proto_string)] return result # 测试调用 proto_content = '''[pinned_field: HEADLINE_1 text: "5 Best Products" asset_performance_label: PENDING policy_summary_info { review_status: REVIEWED approval_status: APPROVED } , pinned_field: HEADLINE_1 text: "10 Best Products 2021" asset_performance_label: PENDING policy_summary_info { review_status: REVIEWED approval_status: APPROVED } # 测试无pinned_field的场景 text: "some_text_without_pinned_field" asset_performance_label: PENDING ''' print(extract_text_from_proto(proto_content)) # 输出:['5 Best Products', '10 Best Products 2021', 'some_text_without_pinned_field']
进阶需求实现(可选匹配pinned_field)
预期输出:['HEADLINE_1: 5 Best Products', 'HEADLINE_1: 10 Best Products 2021', 'some_text_without_pinned_field']
代码如下:
import re def extract_text_with_pinned(proto_string): # 正则中(?:...)?表示可选的非捕获组,优先捕获pinned_field的值,再捕获text内容 regex = r'(?:pinned_field:\s*(\w+)\s+)?text:\s*"([^"]+)"' result = [] for m in re.finditer(regex, proto_string): pinned = m.group(1) text = m.group(2) if pinned: result.append(f"{pinned}: {text}") else: result.append(text) return result # 测试调用 print(extract_text_with_pinned(proto_content)) # 输出:['HEADLINE_1: 5 Best Products', 'HEADLINE_1: 10 Best Products 2021', 'some_text_without_pinned_field']
正则规则说明
(?:pinned_field:\s*(\w+)\s+)?:可选匹配pinned_field: 字段值结构,(\w+)捕获字段值,整个片段不存在时也不影响后续text的匹配text:\s*"([^"]+)":匹配text字段,([^"]+)捕获双引号内的所有内容,自动截断到双引号结束位置,不需要额外匹配asset_performance_label做边界
内容的提问来源于stack exchange,提问作者Elad Benda
相关产品推荐
相关产品推荐

