如何用Python分类文本文件名词及提取商业文章业务属性词汇
如何用Python提取并分类文本中的名词(含商业业务属性词汇)
嘿,我来帮你搞定这个问题!你需要从文本(尤其是商业文章)里提取名词,还要识别出像‘Retail Banking’‘Courier Service’这类能定义业务类型的词汇对吧?咱们一步步来解决:
一、基础名词提取与分类
首先,我们可以用NLTK的词性标注功能,把文本中的词汇按名词类型分类(普通名词、专有名词,单数/复数)。这里需要先处理文本分词、过滤无关词汇,再做标注分类。
二、针对商业业务属性词汇的优化
商业文章里的业务属性词汇大多是复合名词搭配(比如“Steel Plant”),单纯提取单个名词很难识别这类词汇。这时候可以用NLTK的搭配提取工具,找出高频的名词+名词组合,再通过评分筛选出有意义的业务词汇。
下面是完整的可运行代码,我修正了你原代码里的一些小问题(比如Python3中无需decode('utf8')、补全了未完成的逻辑):
import nltk from nltk.collocations import BigramCollocationFinder from nltk.metrics import BigramAssocMeasures from nltk.corpus import stopwords import csv # 首次运行需要下载NLTK的必要资源 nltk.download('punkt') nltk.download('averaged_perceptron_tagger') nltk.download('stopwords') # 读取目标文本文件 with open('bbb_2.txt', 'r', encoding='utf-8') as text_file: raw_text = text_file.read().lower() # 转为小写统一处理 # 分词:把文本拆分为单个词汇 tokens = nltk.wordpunct_tokenize(raw_text) # 过滤停用词(比如the、and这类无意义词汇)和非字母字符 stop_words = set(stopwords.words('english')) filtered_tokens = [token for token in tokens if token.isalpha() and token not in stop_words] # 词性标注:给每个词汇打上词性标签(比如NN=单数普通名词,NNP=单数专有名词) pos_tagged_words = nltk.pos_tag(filtered_tokens) # 1. 按类型分类名词 noun_categories = { "单数普通名词": [word for word, tag in pos_tagged_words if tag == 'NN'], "复数普通名词": [word for word, tag in pos_tagged_words if tag == 'NNS'], "单数专有名词": [word for word, tag in pos_tagged_words if tag == 'NNP'], "复数专有名词": [word for word, tag in pos_tagged_words if tag == 'NNPS'] } # 打印分类结果(只显示前10个示例避免输出过长) print("=== 分类后的名词 ===") for category, nouns in noun_categories.items(): print(f"{category}: {', '.join(nouns[:10])}{'...' if len(nouns) >10 else ''}") # 2. 提取商业业务属性的复合名词搭配 # 定义过滤规则:只保留名词+名词的组合 def is_valid_noun_pair(pair): tag1 = nltk.pos_tag([pair[0]])[0][1] tag2 = nltk.pos_tag([pair[1]])[0][1] return tag1.startswith('NN') and tag2.startswith('NN') # 初始化二元搭配查找器 bigram_finder = BigramCollocationFinder.from_words(filtered_tokens) # 过滤出现次数少于2的搭配(避免偶然出现的组合) bigram_finder.apply_freq_filter(2) # 只保留符合名词+名词规则的搭配 bigram_finder.apply_ngram_filter(is_valid_noun_pair) # 用PMI(点互信息)评分排序,PMI越高说明组合的关联性越强 top_business_terms = bigram_finder.score_ngrams(BigramAssocMeasures.pmi) # 打印识别出的业务词汇 print("\n=== 潜在业务属性词汇(按关联性排序) ===") for term_pair, score in top_business_terms[:10]: print(f"{' '.join(term_pair)} | PMI评分: {score:.2f}") # 可选:把结果保存到CSV文件方便后续分析 with open('business_terms_output.csv', 'w', newline='', encoding='utf-8') as csv_file: writer = csv.writer(csv_file) writer.writerow(["业务词汇", "PMI评分"]) for term_pair, score in top_business_terms[:20]: writer.writerow([' '.join(term_pair), round(score, 2)])
关键说明:
- 停用词过滤:去掉无意义的虚词,让结果更聚焦于有实际意义的业务词汇。
- 词性分类:通过NLTK的词性标签,精准区分不同类型的名词。
- 搭配提取:用二元搭配+PMI评分,能有效识别像“Retail Banking”这类紧密关联的业务复合词,比单个名词提取更贴合你的需求。
- 可扩展性:如果你的业务有特定领域词汇,可以自定义一个行业词典,搭配词频统计进一步提升识别准确率。
内容的提问来源于stack exchange,提问作者SKY
相关产品推荐
相关产品推荐

