You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Stanza(斯坦福NLP)对txt逐句做情感分析及报错解决

报错原因

你遇到的AttributeError: 'Document' object has no attribute 'sentiment'由两个核心问题导致:

  1. 初始化Stanza Pipeline时没有显式加入sentiment处理器,Stanza默认加载的处理器仅包含分词、词性标注、依存句法分析等基础任务,不会加载情感分析模型,自然不会生成情感属性。
  2. Stanza的情感分析结果挂载在句子(Sentence)层级,而非顶层的Document对象上,就算加载了情感模型,直接访问doc.sentiment也会报属性不存在的错误。
修复步骤
  • 初始化stanza.Pipeline时,通过processors参数显式指定加载情感分析模块
  • 对每行文本生成doc对象后,遍历doc.sentences获取每个独立句子的情感结果
  • 用with上下文管理器读写文件,避免文件句柄泄漏
  • 增加空行过滤、行数自动截断逻辑,避免测试时行数不足导致索引越界
  • 补充结果写入文件的逻辑,满足输出要求
修正后可运行代码
import stanza

def sentiment_analysis(input_file, output_file, nlp, test_lines=None):
    # 读取输入文件,自动过滤空行
    with open(input_file, 'r', encoding='utf-8') as f:
        lines = [line.strip() for line in f.read().splitlines() if line.strip()]
    
    # 测试模式下仅处理指定数量的行
    if test_lines is not None:
        lines = lines[:test_lines]
    
    # 逐行分析并写入结果
    with open(output_file, 'w', encoding='utf-8') as f_out:
        for line_idx, line_text in enumerate(lines):
            doc = nlp(line_text)
            # 若一行仅包含一个句子,直接取第一个句子的情感结果;一行多句可遍历所有句子取值
            sent_cls = doc.sentences[0].sentiment
            # 映射分类标签:0=负面,1=中性,2=正面
            sent_label = ["负面", "中性", "正面"][sent_cls]
            # 按格式写入文件,可根据需求调整输出格式
            output_content = f"行号:{line_idx+1}\n原句:{line_text}\n情感结果:{sent_label}\n\n"
            f_out.write(output_content)
            # 控制台打印验证逻辑
            print(f"处理第{line_idx+1}行,情感结果:{sent_label}")

if __name__ == "__main__":
    test_line_count = 10  # 测试时处理前10行,正式全量处理时设为None
    filename = input("Enter the name (with format) of the text you want to filter:\n")
    lang = input("In what language is the text typed? ('ca' for catalan, 'es' for spanish, 'en' for english...)\n")
    stanza.download(lang, verbose=False)
    # 关键:初始化管道时必须加入sentiment处理器
    nlp_pipeline = stanza.Pipeline(lang=lang, processors='tokenize,sentiment', verbose=False)
    sentiment_analysis(filename, "sentiment_result.txt", nlp_pipeline, test_lines=test_line_count)
补充说明
  • Stanza内置情感分析模型返回的是分类结果,取值为0(负面)、1(中性)、2(正面),如果需要0-1区间的连续情感得分,可以根据业务需求对分类结果做映射,或更换支持连续得分的情感模型
  • 不需要手动拆分句子,Stanza自带的分句模型会自动识别文本中的独立句子,一行包含多句时遍历doc.sentences即可拿到每个句子的单独分析结果
  • 首次运行时stanza.download会自动下载对应语言的模型包,后续运行不会重复下载

内容的提问来源于stack exchange,提问作者SVP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 16:27:40