如何用Python将模糊订阅数文本解析为结构化时间序列数据?
问题背景
我用大语言模型从文本中提取流媒体订阅数据,得到如下原始结果:
{ "raw_extractions": [ { "platform_mention": "Netflix", "year_mention": "2012", "subscriber_mention": "roughly 30 million subscribers worldwide" }, { "platform_mention": "Netflix", "year_mention": "2020", "subscriber_mention": "just under 200 million" }, { "platform_mention": "Netflix", "year_mention": "2022", "subscriber_mention": "hovered around 220 million subscribers" } ] }
我需要将其转换为可用于分析的结构化时间序列数据:
| year | platform | subscribers_min | subscribers_max | confidence |
|---|---|---|---|---|
| 2012 | Netflix | 30 | 30 | medium |
| 2020 | Netflix | 195 | 200 | medium |
| 2022 | Netflix | 220 | 220 | medium |
请问在Python中,解析“roughly 30 million”“just under 200 million”这类模糊表述为数值范围的最佳方法是什么?
解析模糊数值短语的Python方案
1. 规则匹配(轻量高效,适合固定场景)
针对常见模糊修饰词手动构建规则映射,搭配正则提取核心数值。这种方法速度快、可解释性强,适合短语类型有限的场景。
示例代码:
import re def parse_fuzzy_subscriber(text): # 提取数字(默认处理百万级单位) num_match = re.search(r'(\d+)', text) if not num_match: return None, None num = int(num_match.group(1)) min_val, max_val = num, num text_lower = text.lower() # 定义模糊修饰词对应的范围规则 if 'just under' in text_lower or 'under' in text_lower: min_val = num - 5 # 可根据业务场景调整差值 max_val = num elif 'roughly' in text_lower or 'around' in text_lower or 'hovered around' in text_lower: min_val = num - 2 max_val = num + 2 # 可扩展更多规则,比如处理'over' 'more than' 'nearly'等表述 return min_val, max_val # 测试 test_phrases = [ "roughly 30 million subscribers worldwide", "just under 200 million", "hovered around 220 million subscribers" ] for phrase in test_phrases: print(parse_fuzzy_subscriber(phrase)) # 输出: # (28, 32) 可根据需求调整为(30,30)或其他范围 # (195, 200) # (218, 222)
2. 轻量NLP工具(spaCy+规则匹配)
如果模糊短语类型更多样,可使用spaCy的Matcher组件,结合词性分析识别修饰词和数值,比纯正则更灵活。
示例代码:
import spacy from spacy.matcher import Matcher nlp = spacy.load("en_core_web_sm") matcher = Matcher(nlp.vocab) # 定义匹配模式:修饰词 + 数字 + 单位 patterns = [ [{"LOWER": {"IN": ["roughly", "around", "hovered", "about"]}}, {"LIKE_NUM": True}, {"LOWER": "million"}], [{"LOWER": {"IN": ["just", "under", "below"]}}, {"LOWER": "under", "OP": "?"}, {"LIKE_NUM": True}, {"LOWER": "million"}], [{"LOWER": {"IN": ["over", "above", "more"]}}, {"LOWER": "than", "OP": "?"}, {"LIKE_NUM": True}, {"LOWER": "million"}] ] for pattern in patterns: matcher.add("FUZZY_NUM", [pattern]) def parse_with_spacy(text): doc = nlp(text) matches = matcher(doc) if not matches: return None, None num_token = None modifier = None for match_id, start, end in matches: span = doc[start:end] # 定位数字token for token in span: if token.like_num: num_token = token # 提取前置修饰词 if start > 0: modifier = doc[start:token.i].text.lower() break if not num_token: return None, None num = int(num_token.text) min_val, max_val = num, num if modifier in ["roughly", "around", "hovered around", "about"]: min_val = num - 2 max_val = num + 2 elif modifier in ["just under", "under", "below"]: min_val = num - 5 max_val = num elif modifier in ["over", "above", "more than"]: min_val = num max_val = num + 5 return min_val, max_val # 测试 for phrase in test_phrases: print(parse_with_spacy(phrase))
3. 调用大语言模型(适合复杂模糊表述)
如果模糊短语包含复杂自然语言描述,直接调用LLM解析成结构化数值范围是最省心的方式。可通过提示词引导模型输出JSON格式结果,方便后续处理。
示例代码(以OpenAI API为例):
import openai import json openai.api_key = "your-api-key" def parse_with_llm(text): prompt = f""" 请将以下订阅用户数的模糊描述解析为数值范围(单位:百万),输出JSON格式,包含min和max字段: 输入:{text} 示例: 输入:just under 200 million 输出:{{"min": 195, "max": 200}} 输入:roughly 30 million 输出:{{"min": 30, "max": 30}} """ response = openai.ChatCompletion.create( model="gpt-3.5-turbo", messages=[{"role": "user", "content": prompt}] ) result = response.choices[0].message.content.strip() try: parsed = json.loads(result) return parsed["min"], parsed["max"] except: return None, None # 测试 for phrase in test_phrases: print(parse_with_llm(phrase))
方法选择建议
- 若模糊短语类型有限、场景固定,优先用规则匹配,速度快且无额外依赖。
- 若短语类型较多但仍属常见表述,用spaCy+规则匹配,兼顾灵活性和效率。
- 若需处理复杂自然语言模糊描述,用大语言模型解析,适配性最强但成本稍高。
内容的提问来源于stack exchange,提问作者Pawan
相关产品推荐
相关产品推荐

