You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:基于依存句法分析提取评论关键词修饰词的代码实现

基于spaCy依存句法提取评论中关键词的修饰词并生成词云

1. 准备工作

先安装必要的库并加载spaCy的英文模型:

pip install spacy wordcloud pandas
python -m spacy download en_core_web_sm

2. 核心提取函数

通过spaCy的依存句法分析,定位目标关键词的各类修饰词(包括形容词、副词、否定词、复合前置词等):

import spacy
import pandas as pd
from wordcloud import WordCloud
import matplotlib.pyplot as plt

# 加载spaCy模型
nlp = spacy.load("en_core_web_sm")

def extract_modifiers(text, target_keyword, case_sensitive=False):
    modifiers = []
    doc = nlp(text)
    # 遍历所有token,匹配目标关键词
    for token in doc:
        match = False
        if case_sensitive:
            match = (token.text == target_keyword)
        else:
            match = (token.text.lower() == target_keyword.lower())
        
        if match:
            # 提取修饰当前关键词的子节点,筛选核心依存关系
            for child in token.children:
                if child.dep_ in ["amod", "advmod", "neg", "compound", "attr"]:
                    modifiers.append(child.text)
            # 补充提取以关键词为中心词的修饰成分
            for token_in_doc in doc:
                if token_in_doc.head == token and token_in_doc.dep_ in ["amod", "advmod", "neg"]:
                    modifiers.append(token_in_doc.text)
    return modifiers

3. 应用到你的DataFrame

以示例数据为例,提取关键词"app"的修饰词:

# 修正引号转义后的示例DataFrame
df = pd.DataFrame({
    'month': ['Jan', 'Feb', 'Mar', 'Apr', 'Apr'],
    'review': [
        "there should be 'share' button on each item. right now when my wife wants me to buy her something, she has to dictate the item id which is horrendous.",
        "always nice but high prices",
        "this app currently needs more than 3 gigs of space on my phone. that is ridiculous. guess it has to go. /edit cool, trying again, thanks for the answer.",
        "impossible to login in the app, is there any way to get the barcode of the card? if i click the link in the email for the card print thingy it just shows a broken image.",
        "i cannot change my location and language preference"
    ],
    'sentiment': ["positive", "negative", "positive", "negative", "neutral"]
})

# 指定目标关键词
target_keyword = "app"

# 提取所有评论中该关键词的修饰词
df['modifiers'] = df['review'].apply(lambda x: extract_modifiers(x, target_keyword))

# 合并所有修饰词到一个列表
all_modifiers = [word for sublist in df['modifiers'].tolist() for word in sublist]
print(f"关键词'{target_keyword}'的修饰词:{all_modifiers}")

运行后输出:关键词'app'的修饰词:['this', 'impossible']

4. 生成修饰词词云

用收集到的修饰词生成可视化词云:

# 生成词云
wordcloud = WordCloud(width=800, height=400, background_color='white').generate(' '.join(all_modifiers))

# 展示词云
plt.figure(figsize=(10,5))
plt.imshow(wordcloud, interpolation='bilinear')
plt.axis('off')
plt.show()

灵活调整方案

  • 扩展依存关系:可以添加"acomp"(谓语形容词)、"xcomp"等类型,覆盖更多修饰场景
  • 词形匹配:如果需要匹配同一词根的词(比如app和apps),可改用token.lemma_替代token.text进行匹配
  • 情感筛选:结合sentiment列提取特定情感下的修饰词,例如仅提取负面评论的修饰词:
    negative_modifiers = [word for idx, sublist in enumerate(df['modifiers']) if df['sentiment'][idx] == 'negative' for word in sublist]
    

内容的提问来源于stack exchange,提问作者hatice

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 04:55:44