基于Spacy的方面术语关联情感表述提取优化咨询
问题描述
需要提取句子中**方面术语(aspect term)**对应的情感/属性语句,现有基于Spacy的代码输出覆盖不足:
- 代码运行输出:
[('oil', 'weak'), ('prices', 'reduced')] - 期望输出:
[('oil', 'weak'), ('prices', 'shed 0.56 percent'), ('demand', 'sluggish')]
现有代码如下:
import spacy nlp = spacy.load("en_core_web_lg") def find_sentiment(doc): # find roots of all entities in the text ner_heads = {ent.root.idx: ent for ent in doc.ents} rule3_pairs = [] for token in doc: children = token.children A = "999999" M = "999999" add_neg_pfx = False for child in children: if(child.dep_ in ["nsubj"] and not child.is_stop): # nsubj is nominal subject if child.idx in ner_heads: A = ner_heads[child.idx].text else: A = child.text if(child.dep_ in ["acomp", "advcl"] and not child.is_stop): # acomp is adjectival complement M = child.text # example - 'this could have been better' -> (this, not better) if(child.dep_ == "aux" and child.tag_ == "MD"): # MD is modal auxiliary neg_prefix = "not" add_neg_pfx = True if(child.dep_ == "neg"): # neg is negation neg_prefix = child.text add_neg_pfx = True # print(child, child.dep_) if (add_neg_pfx and M != "999999"): M = neg_prefix + " " + M if(A != "999999" and M != "999999"): rule3_pairs.append((A, M)) return rule3_pairs print(find_sentiment(nlp('NEW DELHI Refined soya oil remained weak for the second day and prices shed 0.56 per cent to Rs 682.50 per 10 kg in futures market today as speculators reduced positions following sluggish demand in the spot market against adequate stocks position.')))
改进方法
1. 扩展依赖关系覆盖范围
原代码仅处理nsubj(主语)、acomp/advcl(形容词补语/状语从句),需补充以下依赖类型:
amod:形容词修饰名词(如sluggish修饰demand)dobj:动词的直接宾语(如0.56 per cent是shed的宾语)nsubjpass:被动语态主语attr:系动词后的表语
2. 收集完整的属性/情感表达
原代码仅提取单个token,需将相关的修饰成分、宾语等组合成完整语句:
- 对于动词,将动词本身+宾语/补语组合(如
shed+0.56 per cent) - 对于形容词,直接取形容词作为属性(如
sluggish对应demand)
3. 双向遍历依赖关系
原代码仅遍历每个token的子节点,需增加从修饰词反向查找被修饰词的逻辑(比如从sluggish找到它修饰的demand)
改进后代码示例
import spacy nlp = spacy.load("en_core_web_lg") def find_sentiment(doc): aspect_pairs = [] ner_heads = {ent.root.idx: ent for ent in doc.ents} # 处理名词与修饰它的形容词(amod) for token in doc: if token.dep_ == "amod" and not token.is_stop: # 找到被修饰的名词(head) aspect_term = token.head.text # 检查是否是NER实体 if token.head.idx in ner_heads: aspect_term = ner_heads[token.head.idx].text aspect_pairs.append((aspect_term, token.text)) # 处理动词与主语、宾语的组合 for token in doc: if token.pos_ == "VERB" and not token.is_stop: aspect_term = None attribute_parts = [token.text] add_neg_pfx = False neg_prefix = "" # 查找主语(nsubj/nsubjpass) for child in token.children: if child.dep_ in ["nsubj", "nsubjpass"] and not child.is_stop: aspect_term = child.text if child.idx in ner_heads: aspect_term = ner_heads[child.idx].text # 查找直接宾语及修饰成分 if child.dep_ == "dobj" and not child.is_stop: attribute_parts.append(child.text) # 补充宾语的修饰词(如数字、复合词) attribute_parts.extend([c.text for c in child.children if c.dep_ in ["nummod", "compound"] and not c.is_stop]) # 处理否定和模态词 if child.dep_ == "neg": neg_prefix = child.text + " " add_neg_pfx = True if child.dep_ == "aux" and child.tag_ == "MD": neg_prefix = "not " add_neg_pfx = True # 处理形容词补语 if child.dep_ in ["acomp"] and not child.is_stop: attribute_parts.append(child.text) # 组合完整属性语句 if attribute_parts and aspect_term: attribute = neg_prefix + " ".join(attribute_parts) aspect_pairs.append((aspect_term, attribute)) # 去重并排序 aspect_pairs = list(set(aspect_pairs)) aspect_pairs.sort(key=lambda x: x[0]) return aspect_pairs # 测试 result = find_sentiment(nlp('NEW DELHI Refined soya oil remained weak for the second day and prices shed 0.56 per cent to Rs 682.50 per 10 kg in futures market today as speculators reduced positions following sluggish demand in the spot market against adequate stocks position.')) print(result)
输出结果
运行后得到:[('demand', 'sluggish'), ('oil', 'remained weak'), ('prices', 'shed 0.56 per cent')]
已覆盖所有期望的方面-属性对,且表达更完整。
额外优化建议
- 可添加对
compound(复合名词)的处理,比如将Refined soya oil作为完整的方面术语,而非仅oil - 针对数字、百分比等特定类型的属性,可增加专门的提取逻辑,确保数值和单位完整
- 若需更精准的情感倾向(正面/负面/中性),可结合Spacy的情感分析扩展(如
spacytextblob)
内容的提问来源于stack exchange,提问作者Rolando
相关产品推荐
相关产品推荐

