You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将模糊订阅数文本解析为结构化时间序列数据?

问题背景

我用大语言模型从文本中提取流媒体订阅数据,得到如下原始结果:

{
  "raw_extractions": [
    {
      "platform_mention": "Netflix",
      "year_mention": "2012",
      "subscriber_mention": "roughly 30 million subscribers worldwide"
    },
    {
      "platform_mention": "Netflix",
      "year_mention": "2020",
      "subscriber_mention": "just under 200 million"
    },
    {
      "platform_mention": "Netflix",
      "year_mention": "2022",
      "subscriber_mention": "hovered around 220 million subscribers"
    }
  ]
}

我需要将其转换为可用于分析的结构化时间序列数据:

yearplatformsubscribers_minsubscribers_maxconfidence
2012Netflix3030medium
2020Netflix195200medium
2022Netflix220220medium

请问在Python中,解析“roughly 30 million”“just under 200 million”这类模糊表述为数值范围的最佳方法是什么?


解析模糊数值短语的Python方案

1. 规则匹配(轻量高效,适合固定场景)

针对常见模糊修饰词手动构建规则映射,搭配正则提取核心数值。这种方法速度快、可解释性强,适合短语类型有限的场景。

示例代码:

import re

def parse_fuzzy_subscriber(text):
    # 提取数字(默认处理百万级单位)
    num_match = re.search(r'(\d+)', text)
    if not num_match:
        return None, None
    
    num = int(num_match.group(1))
    min_val, max_val = num, num
    text_lower = text.lower()
    
    # 定义模糊修饰词对应的范围规则
    if 'just under' in text_lower or 'under' in text_lower:
        min_val = num - 5  # 可根据业务场景调整差值
        max_val = num
    elif 'roughly' in text_lower or 'around' in text_lower or 'hovered around' in text_lower:
        min_val = num - 2
        max_val = num + 2
    # 可扩展更多规则,比如处理'over' 'more than' 'nearly'等表述
    
    return min_val, max_val

# 测试
test_phrases = [
    "roughly 30 million subscribers worldwide",
    "just under 200 million",
    "hovered around 220 million subscribers"
]
for phrase in test_phrases:
    print(parse_fuzzy_subscriber(phrase))
# 输出:
# (28, 32) 可根据需求调整为(30,30)或其他范围
# (195, 200)
# (218, 222)

2. 轻量NLP工具(spaCy+规则匹配)

如果模糊短语类型更多样,可使用spaCy的Matcher组件,结合词性分析识别修饰词和数值,比纯正则更灵活。

示例代码:

import spacy
from spacy.matcher import Matcher

nlp = spacy.load("en_core_web_sm")
matcher = Matcher(nlp.vocab)

# 定义匹配模式:修饰词 + 数字 + 单位
patterns = [
    [{"LOWER": {"IN": ["roughly", "around", "hovered", "about"]}}, {"LIKE_NUM": True}, {"LOWER": "million"}],
    [{"LOWER": {"IN": ["just", "under", "below"]}}, {"LOWER": "under", "OP": "?"}, {"LIKE_NUM": True}, {"LOWER": "million"}],
    [{"LOWER": {"IN": ["over", "above", "more"]}}, {"LOWER": "than", "OP": "?"}, {"LIKE_NUM": True}, {"LOWER": "million"}]
]

for pattern in patterns:
    matcher.add("FUZZY_NUM", [pattern])

def parse_with_spacy(text):
    doc = nlp(text)
    matches = matcher(doc)
    if not matches:
        return None, None
    
    num_token = None
    modifier = None
    for match_id, start, end in matches:
        span = doc[start:end]
        # 定位数字token
        for token in span:
            if token.like_num:
                num_token = token
                # 提取前置修饰词
                if start > 0:
                    modifier = doc[start:token.i].text.lower()
                break
    if not num_token:
        return None, None
    
    num = int(num_token.text)
    min_val, max_val = num, num
    
    if modifier in ["roughly", "around", "hovered around", "about"]:
        min_val = num - 2
        max_val = num + 2
    elif modifier in ["just under", "under", "below"]:
        min_val = num - 5
        max_val = num
    elif modifier in ["over", "above", "more than"]:
        min_val = num
        max_val = num + 5
    
    return min_val, max_val

# 测试
for phrase in test_phrases:
    print(parse_with_spacy(phrase))

3. 调用大语言模型(适合复杂模糊表述)

如果模糊短语包含复杂自然语言描述,直接调用LLM解析成结构化数值范围是最省心的方式。可通过提示词引导模型输出JSON格式结果,方便后续处理。

示例代码(以OpenAI API为例):

import openai
import json

openai.api_key = "your-api-key"

def parse_with_llm(text):
    prompt = f"""
    请将以下订阅用户数的模糊描述解析为数值范围(单位:百万),输出JSON格式,包含min和max字段:
    输入:{text}
    示例:
    输入:just under 200 million
    输出:{{"min": 195, "max": 200}}
    输入:roughly 30 million
    输出:{{"min": 30, "max": 30}}
    """
    response = openai.ChatCompletion.create(
        model="gpt-3.5-turbo",
        messages=[{"role": "user", "content": prompt}]
    )
    result = response.choices[0].message.content.strip()
    
    try:
        parsed = json.loads(result)
        return parsed["min"], parsed["max"]
    except:
        return None, None

# 测试
for phrase in test_phrases:
    print(parse_with_llm(phrase))

方法选择建议

  • 若模糊短语类型有限、场景固定,优先用规则匹配,速度快且无额外依赖。
  • 若短语类型较多但仍属常见表述,用spaCy+规则匹配,兼顾灵活性和效率。
  • 若需处理复杂自然语言模糊描述,用大语言模型解析,适配性最强但成本稍高。

内容的提问来源于stack exchange,提问作者Pawan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.01 22:23:12