You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python分类文本文件名词及提取商业文章业务属性词汇

如何用Python提取并分类文本中的名词(含商业业务属性词汇)

嘿,我来帮你搞定这个问题!你需要从文本(尤其是商业文章)里提取名词,还要识别出像‘Retail Banking’‘Courier Service’这类能定义业务类型的词汇对吧?咱们一步步来解决:

一、基础名词提取与分类

首先,我们可以用NLTK的词性标注功能,把文本中的词汇按名词类型分类(普通名词、专有名词,单数/复数)。这里需要先处理文本分词、过滤无关词汇,再做标注分类。

二、针对商业业务属性词汇的优化

商业文章里的业务属性词汇大多是复合名词搭配(比如“Steel Plant”),单纯提取单个名词很难识别这类词汇。这时候可以用NLTK的搭配提取工具,找出高频的名词+名词组合,再通过评分筛选出有意义的业务词汇。

下面是完整的可运行代码,我修正了你原代码里的一些小问题(比如Python3中无需decode('utf8')、补全了未完成的逻辑):

import nltk
from nltk.collocations import BigramCollocationFinder
from nltk.metrics import BigramAssocMeasures
from nltk.corpus import stopwords
import csv

# 首次运行需要下载NLTK的必要资源
nltk.download('punkt')
nltk.download('averaged_perceptron_tagger')
nltk.download('stopwords')

# 读取目标文本文件
with open('bbb_2.txt', 'r', encoding='utf-8') as text_file:
    raw_text = text_file.read().lower()  # 转为小写统一处理

# 分词:把文本拆分为单个词汇
tokens = nltk.wordpunct_tokenize(raw_text)

# 过滤停用词(比如the、and这类无意义词汇)和非字母字符
stop_words = set(stopwords.words('english'))
filtered_tokens = [token for token in tokens if token.isalpha() and token not in stop_words]

# 词性标注:给每个词汇打上词性标签(比如NN=单数普通名词,NNP=单数专有名词)
pos_tagged_words = nltk.pos_tag(filtered_tokens)

# 1. 按类型分类名词
noun_categories = {
    "单数普通名词": [word for word, tag in pos_tagged_words if tag == 'NN'],
    "复数普通名词": [word for word, tag in pos_tagged_words if tag == 'NNS'],
    "单数专有名词": [word for word, tag in pos_tagged_words if tag == 'NNP'],
    "复数专有名词": [word for word, tag in pos_tagged_words if tag == 'NNPS']
}

# 打印分类结果(只显示前10个示例避免输出过长)
print("=== 分类后的名词 ===")
for category, nouns in noun_categories.items():
    print(f"{category}: {', '.join(nouns[:10])}{'...' if len(nouns) >10 else ''}")

# 2. 提取商业业务属性的复合名词搭配
# 定义过滤规则:只保留名词+名词的组合
def is_valid_noun_pair(pair):
    tag1 = nltk.pos_tag([pair[0]])[0][1]
    tag2 = nltk.pos_tag([pair[1]])[0][1]
    return tag1.startswith('NN') and tag2.startswith('NN')

# 初始化二元搭配查找器
bigram_finder = BigramCollocationFinder.from_words(filtered_tokens)
# 过滤出现次数少于2的搭配(避免偶然出现的组合)
bigram_finder.apply_freq_filter(2)
# 只保留符合名词+名词规则的搭配
bigram_finder.apply_ngram_filter(is_valid_noun_pair)

# 用PMI(点互信息)评分排序,PMI越高说明组合的关联性越强
top_business_terms = bigram_finder.score_ngrams(BigramAssocMeasures.pmi)

# 打印识别出的业务词汇
print("\n=== 潜在业务属性词汇(按关联性排序) ===")
for term_pair, score in top_business_terms[:10]:
    print(f"{' '.join(term_pair)} | PMI评分: {score:.2f}")

# 可选:把结果保存到CSV文件方便后续分析
with open('business_terms_output.csv', 'w', newline='', encoding='utf-8') as csv_file:
    writer = csv.writer(csv_file)
    writer.writerow(["业务词汇", "PMI评分"])
    for term_pair, score in top_business_terms[:20]:
        writer.writerow([' '.join(term_pair), round(score, 2)])

关键说明:

  • 停用词过滤:去掉无意义的虚词,让结果更聚焦于有实际意义的业务词汇。
  • 词性分类:通过NLTK的词性标签,精准区分不同类型的名词。
  • 搭配提取:用二元搭配+PMI评分,能有效识别像“Retail Banking”这类紧密关联的业务复合词,比单个名词提取更贴合你的需求。
  • 可扩展性:如果你的业务有特定领域词汇,可以自定义一个行业词典,搭配词频统计进一步提升识别准确率。

内容的提问来源于stack exchange,提问作者SKY

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:03:54