You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Spacy提取数值属性时转字典丢失数据的技术问题

问题分析与解决方案

问题原因

你遇到的7 men丢失问题,核心原因是字典的键具有唯一性:原代码提取的属性列表中包含两个men(分别对应数值7和3),当你将attributes和values通过dict(zip(...))转换为字典时,后出现的men:3会直接覆盖先出现的men:7,导致前者丢失。

另外原代码中的i > 0判断并非问题根源,但可以移除以让代码更健壮(避免遗漏句子开头的数字)。


修复方案

根据需求,提供两种可行的处理方式:

方案1:保留所有数值-属性对(用列表存储)

如果需要完整保留所有提取到的数值与属性组合,直接用列表存储元组即可,避免字典键重复的限制:

import spacy

nlp = spacy.blank('en')

sentence = "From a group of 7 men and 6 women, five persons are to be selected to form a committee so that at least 3 men are there on the committee. In how many ways can it be done?"

doc = nlp(sentence)

# 提取所有数值-属性对
num_attr_pairs = []
for i, token in enumerate(doc):
    if token.is_digit or token.like_num:
        # 防止索引越界(避免数字在句子末尾的情况)
        if i + 1 < len(doc):
            num_attr_pairs.append((token.text, doc[i+1].text))

# 输出所有结果
print("完整提取的数值-属性对:")
for num, attr in num_attr_pairs:
    print(f"{attr} {num}")

运行输出:

完整提取的数值-属性对:
men 7
women 6
persons five
men 3

方案2:按属性分组存储(用带列表的字典)

如果希望按属性分类,将同一属性的所有数值放在一起,可以用defaultdict实现:

import spacy
from collections import defaultdict

nlp = spacy.blank('en')

sentence = "From a group of 7 men and 6 women, five persons are to be selected to form a committee so that at least 3 men are there on the committee. In how many ways can it be done?"

doc = nlp(sentence)

# 按属性分组存储数值
num_attr_dict = defaultdict(list)
for i, token in enumerate(doc):
    if token.is_digit or token.like_num:
        if i + 1 < len(doc):
            num_attr_dict[doc[i+1].text].append(token.text)

# 转换为普通字典(可选)
num_attr_dict = dict(num_attr_dict)

print("按属性分组的结果:")
for attr, nums in num_attr_dict.items():
    print(f"{attr}: {nums}")

运行输出:

按属性分组的结果:
men: ['7', '3']
women: ['6']
persons: ['five']

内容的提问来源于stack exchange,提问作者user6541792

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 16:51:16