You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

NLP入门求助:如何使用spaCy获取词性标注(POS)

使用spaCy实现词性标注(POS)

前置准备

首先安装spaCy及英文预训练模型:

  • 安装spaCy:pip install spacy
  • 下载英文轻量模型:python -m spacy download en_core_web_sm

实现代码

import spacy

# 加载预训练模型
nlp = spacy.load("en_core_web_sm")

# 输入待处理列表
processed_lst = [['The', 'wild', 'is', 'dangerous'], ['The', 'rockstar', 'is', 'wild']]

final_lst = []

# 遍历每个子列表进行处理
for token_group in processed_lst:
    # 将token列表拼接为spaCy可处理的字符串
    text = " ".join(token_group)
    doc = nlp(text)
    # 生成(token, POS标签)元组列表
    pos_result = [(token.text, token.pos_) for token in doc]
    final_lst.append(pos_result)

# 查看结果
print(final_lst)

说明

  • spaCy的nlp对象仅接收字符串输入,因此需要先将每个子列表的token拼接为空格分隔的文本。
  • token.pos_返回标准化的通用词性标签,与示例中的DET、NOUN、AUX、ADJ完全对应。
  • 运行代码后将得到与期望输出一致的final_lst。

内容的提问来源于stack exchange,提问作者The Humble Coder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 08:48:21