You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从混合大小写数据集中提取专有名词及对应语句?

提取专有名词及对应语句的Python实现

1. 准备依赖

作为NLP新手,推荐用spaCy这个易用的库来做命名实体识别(NER),它能准确识别大部分专有名词,同时处理大小写混合的场景。先安装必要工具:

  • 安装spaCy:pip install spacy
  • 下载预训练模型:
    • 英文数据集:python -m spacy download en_core_web_sm
    • 中文数据集:python -m spacy download zh_core_web_sm

2. 核心实现代码

下面是完整的示例代码,包含数据处理、专有名词提取和结果整理成两列的逻辑:

import spacy
import pandas as pd

# 加载预训练模型(中文替换为zh_core_web_sm)
nlp = spacy.load("en_core_web_sm")

# 替换成你的实际句子数据集
sample_sentences = [
    "apple was founded by Steve Jobs in Cupertino.",
    "The eiffel tower is located in Paris, France.",
    "Microsoft released windows 11 last year."
]

# 存储提取结果的列表
extracted_data = []

for sentence in sample_sentences:
    # 用spaCy处理句子
    doc = nlp(sentence)
    # 提取所有命名实体(即专有名词),可通过ent.label_筛选特定类型(如PERSON/ORG/GPE)
    proper_nouns = [ent.text for ent in doc.ents]
    # 将句子和对应专有名词存入列表
    extracted_data.append({
        "语句": sentence,
        "专有名词": ", ".join(proper_nouns) if proper_nouns else "无"
    })

# 转换为DataFrame,生成两列结构
result_df = pd.DataFrame(extracted_data)
print(result_df)

3. 针对大小写混合的补充优化

如果遇到spaCy未识别的特殊专有名词(比如小众品牌、领域术语),可以补充一个基于大小写的规则辅助提取:

def extract_proper_nouns(sentence):
    doc = nlp(sentence)
    # NER识别的专有名词
    ner_nouns = [ent.text for ent in doc.ents]
    # 提取句子中除开头外首字母大写的单词(补充NER遗漏的情况)
    words = sentence.split()
    case_based_nouns = []
    for idx, word in enumerate(words):
        if idx != 0 and word.istitle() and word not in ner_nouns:
            case_based_nouns.append(word)
    # 合并去重
    all_proper_nouns = list(set(ner_nouns + case_based_nouns))
    return ", ".join(all_proper_nouns) if all_proper_nouns else "无"

# 重新处理数据集
extracted_data = []
for sentence in sample_sentences:
    extracted_data.append({
        "语句": sentence,
        "专有名词": extract_proper_nouns(sentence)
    })

result_df = pd.DataFrame(extracted_data)
print(result_df)

4. 保存结果到文件

如果需要把结果导出到本地文件:

# 保存为CSV(支持中文,用utf-8-sig编码)
result_df.to_csv("专有名词提取结果.csv", index=False, encoding="utf-8-sig")

# 保存为Excel
result_df.to_excel("专有名词提取结果.xlsx", index=False)

内容的提问来源于stack exchange,提问作者Prat1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 04:05:26