使用Spacy提取数值属性时转字典丢失数据的技术问题
问题分析与解决方案
问题原因
你遇到的7 men丢失问题,核心原因是字典的键具有唯一性:原代码提取的属性列表中包含两个men(分别对应数值7和3),当你将attributes和values通过dict(zip(...))转换为字典时,后出现的men:3会直接覆盖先出现的men:7,导致前者丢失。
另外原代码中的i > 0判断并非问题根源,但可以移除以让代码更健壮(避免遗漏句子开头的数字)。
修复方案
根据需求,提供两种可行的处理方式:
方案1:保留所有数值-属性对(用列表存储)
如果需要完整保留所有提取到的数值与属性组合,直接用列表存储元组即可,避免字典键重复的限制:
import spacy nlp = spacy.blank('en') sentence = "From a group of 7 men and 6 women, five persons are to be selected to form a committee so that at least 3 men are there on the committee. In how many ways can it be done?" doc = nlp(sentence) # 提取所有数值-属性对 num_attr_pairs = [] for i, token in enumerate(doc): if token.is_digit or token.like_num: # 防止索引越界(避免数字在句子末尾的情况) if i + 1 < len(doc): num_attr_pairs.append((token.text, doc[i+1].text)) # 输出所有结果 print("完整提取的数值-属性对:") for num, attr in num_attr_pairs: print(f"{attr} {num}")
运行输出:
完整提取的数值-属性对: men 7 women 6 persons five men 3
方案2:按属性分组存储(用带列表的字典)
如果希望按属性分类,将同一属性的所有数值放在一起,可以用defaultdict实现:
import spacy from collections import defaultdict nlp = spacy.blank('en') sentence = "From a group of 7 men and 6 women, five persons are to be selected to form a committee so that at least 3 men are there on the committee. In how many ways can it be done?" doc = nlp(sentence) # 按属性分组存储数值 num_attr_dict = defaultdict(list) for i, token in enumerate(doc): if token.is_digit or token.like_num: if i + 1 < len(doc): num_attr_dict[doc[i+1].text].append(token.text) # 转换为普通字典(可选) num_attr_dict = dict(num_attr_dict) print("按属性分组的结果:") for attr, nums in num_attr_dict.items(): print(f"{attr}: {nums}")
运行输出:
按属性分组的结果: men: ['7', '3'] women: ['6'] persons: ['five']
内容的提问来源于stack exchange,提问作者user6541792
相关产品推荐
相关产品推荐

