如何用Python NLP按语义拆分法语文本或提取主题标签?
法语文本语义拆分与主题标签提取方案
一、语义单元拆分(按主题拆分文本)
普通分句工具(如NLTK)仅基于标点拆分,无法识别语义边界。针对法语文本,推荐以下两种精准方案:
1. spaCy法语模型结合语义聚类
spaCy的法语预训练模型(fr_core_news_md/fr_core_news_lg)能更好理解法语文法与上下文,先分句再通过句子向量聚类实现语义分组:
import spacy import pandas as pd from sklearn.cluster import KMeans # 加载spaCy法语模型 nlp = spacy.load("fr_core_news_md") def split_by_semantics(text, num_topics=3): # 分句并过滤空内容 doc = nlp(text) sentences = [sent.text.strip() for sent in doc.sents if sent.text.strip()] if not sentences: return [] # 提取句子向量用于聚类 sent_vectors = [sent.vector for sent in doc.sents if sent.text.strip()] # K均值聚类划分语义组 kmeans = KMeans(n_clusters=num_topics, random_state=42) clusters = kmeans.fit_predict(sent_vectors) # 按聚类结果拼接语义单元 semantic_units = {} for idx, cluster in enumerate(clusters): semantic_units.setdefault(cluster, []).append(sentences[idx]) return [" ".join(units) for units in semantic_units.values()] # 应用到DataFrame df = pd.DataFrame({"text": ["Votre texte mixte sur sport, économie et éducation..."]}) df["semantic_units"] = df["text"].apply(split_by_semantics)
2. 基于Transformer的主题导向拆分
用法语预训练模型(CamemBERT)做零样本主题分类,将同主题句子归为一类:
from transformers import pipeline # 初始化零样本分类管道 classifier = pipeline("zero-shot-classification", model="camembert-base") def split_by_topic(text, candidate_topics=["sport", "économie", "éducation"]): doc = nlp(text) sentences = [sent.text.strip() for sent in doc.sents if sent.text.strip()] if not sentences: return [] # 对每个句子做主题匹配 topic_groups = {topic: [] for topic in candidate_topics} for sent in sentences: result = classifier(sent, candidate_topics) top_topic = result["labels"][0] topic_groups[top_topic].append(sent) # 过滤空分组并返回 return [" ".join(group) for group in topic_groups.values() if group] # 应用到DataFrame df["topic_based_units"] = df["text"].apply(split_by_topic)
二、主题标签提取
如果仅需提取主题标签而非拆分文本,推荐以下高效方案:
1. PyTextRank核心短语提取
PyTextRank支持法语,能直接从文本中提取核心短语作为主题标签:
import pytextrank # 给spaCy添加PyTextRank管道 nlp.add_pipe("textrank") def extract_key_tags(text, num_tags=3): doc = nlp(text) # 提取排名靠前的核心短语 key_phrases = [phrase.text for phrase in doc._.phrases[:num_tags]] # 可按需映射为英文标签 tag_map = {"sport": "sports", "économie": "economic", "éducation": "school"} return [tag_map.get(tag.lower(), tag) for tag in key_phrases] # 应用到DataFrame df["topic_tags"] = df["text"].apply(extract_key_tags)
2. LDA主题建模(批量文本场景)
针对大规模法语文本,用LDA挖掘潜在主题并提取标签:
import gensim from gensim.corpora import Dictionary from gensim.models import LdaModel # 文本预处理:分词、去停用词、词形还原 def preprocess(text): doc = nlp(text) return [token.lemma_ for token in doc if not token.is_stop and not token.is_punct and token.is_alpha] # 预处理所有文本 texts = df["text"].apply(preprocess).tolist() # 创建词典与语料库 dictionary = Dictionary(texts) corpus = [dictionary.doc2bow(text) for text in texts] # 训练LDA模型 lda_model = LdaModel(corpus=corpus, id2word=dictionary, num_topics=3, random_state=42) # 提取单条文本的主题标签 def get_lda_tags(text, num_tags=2): bow = dictionary.doc2bow(preprocess(text)) top_topic = max(lda_model.get_document_topics(bow), key=lambda x: x[1])[0] topic_words = lda_model.show_topic(top_topic, num_words=num_tags) return [word for word, prob in topic_words] df["lda_tags"] = df["text"].apply(get_lda_tags)
3. 零样本标签匹配(已知候选主题)
如果已有明确的候选标签,直接用零样本分类匹配:
def extract_zero_shot_tags(text, candidate_tags=["sports", "economic", "school"]): fr_candidates = ["sport", "économie", "éducation"] result = classifier(text, fr_candidates) tag_map = dict(zip(fr_candidates, candidate_tags)) # 取置信度Top2的标签 return [tag_map[tag] for tag in result["labels"][:2]] df["zero_shot_tags"] = df["text"].apply(extract_zero_shot_tags)
依赖安装
执行以下命令安装所需工具:
pip install spacy pandas scikit-learn transformers pytextrank gensim python -m spacy download fr_core_news_md
内容的提问来源于stack exchange,提问作者Paradisum
相关产品推荐
相关产品推荐

