如何用Python从混合大小写数据集中提取专有名词及对应语句?
提取专有名词及对应语句的Python实现
1. 准备依赖
作为NLP新手,推荐用spaCy这个易用的库来做命名实体识别(NER),它能准确识别大部分专有名词,同时处理大小写混合的场景。先安装必要工具:
- 安装spaCy:
pip install spacy - 下载预训练模型:
- 英文数据集:
python -m spacy download en_core_web_sm - 中文数据集:
python -m spacy download zh_core_web_sm
- 英文数据集:
2. 核心实现代码
下面是完整的示例代码,包含数据处理、专有名词提取和结果整理成两列的逻辑:
import spacy import pandas as pd # 加载预训练模型(中文替换为zh_core_web_sm) nlp = spacy.load("en_core_web_sm") # 替换成你的实际句子数据集 sample_sentences = [ "apple was founded by Steve Jobs in Cupertino.", "The eiffel tower is located in Paris, France.", "Microsoft released windows 11 last year." ] # 存储提取结果的列表 extracted_data = [] for sentence in sample_sentences: # 用spaCy处理句子 doc = nlp(sentence) # 提取所有命名实体(即专有名词),可通过ent.label_筛选特定类型(如PERSON/ORG/GPE) proper_nouns = [ent.text for ent in doc.ents] # 将句子和对应专有名词存入列表 extracted_data.append({ "语句": sentence, "专有名词": ", ".join(proper_nouns) if proper_nouns else "无" }) # 转换为DataFrame,生成两列结构 result_df = pd.DataFrame(extracted_data) print(result_df)
3. 针对大小写混合的补充优化
如果遇到spaCy未识别的特殊专有名词(比如小众品牌、领域术语),可以补充一个基于大小写的规则辅助提取:
def extract_proper_nouns(sentence): doc = nlp(sentence) # NER识别的专有名词 ner_nouns = [ent.text for ent in doc.ents] # 提取句子中除开头外首字母大写的单词(补充NER遗漏的情况) words = sentence.split() case_based_nouns = [] for idx, word in enumerate(words): if idx != 0 and word.istitle() and word not in ner_nouns: case_based_nouns.append(word) # 合并去重 all_proper_nouns = list(set(ner_nouns + case_based_nouns)) return ", ".join(all_proper_nouns) if all_proper_nouns else "无" # 重新处理数据集 extracted_data = [] for sentence in sample_sentences: extracted_data.append({ "语句": sentence, "专有名词": extract_proper_nouns(sentence) }) result_df = pd.DataFrame(extracted_data) print(result_df)
4. 保存结果到文件
如果需要把结果导出到本地文件:
# 保存为CSV(支持中文,用utf-8-sig编码) result_df.to_csv("专有名词提取结果.csv", index=False, encoding="utf-8-sig") # 保存为Excel result_df.to_excel("专有名词提取结果.xlsx", index=False)
内容的提问来源于stack exchange,提问作者Prat1
相关产品推荐
相关产品推荐

