You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python NLP按语义拆分法语文本或提取主题标签?

法语文本语义拆分与主题标签提取方案

一、语义单元拆分(按主题拆分文本)

普通分句工具(如NLTK)仅基于标点拆分,无法识别语义边界。针对法语文本,推荐以下两种精准方案:

1. spaCy法语模型结合语义聚类

spaCy的法语预训练模型(fr_core_news_md/fr_core_news_lg)能更好理解法语文法与上下文,先分句再通过句子向量聚类实现语义分组:

import spacy
import pandas as pd
from sklearn.cluster import KMeans

# 加载spaCy法语模型
nlp = spacy.load("fr_core_news_md")

def split_by_semantics(text, num_topics=3):
    # 分句并过滤空内容
    doc = nlp(text)
    sentences = [sent.text.strip() for sent in doc.sents if sent.text.strip()]
    if not sentences:
        return []
    
    # 提取句子向量用于聚类
    sent_vectors = [sent.vector for sent in doc.sents if sent.text.strip()]
    
    # K均值聚类划分语义组
    kmeans = KMeans(n_clusters=num_topics, random_state=42)
    clusters = kmeans.fit_predict(sent_vectors)
    
    # 按聚类结果拼接语义单元
    semantic_units = {}
    for idx, cluster in enumerate(clusters):
        semantic_units.setdefault(cluster, []).append(sentences[idx])
    
    return [" ".join(units) for units in semantic_units.values()]

# 应用到DataFrame
df = pd.DataFrame({"text": ["Votre texte mixte sur sport, économie et éducation..."]})
df["semantic_units"] = df["text"].apply(split_by_semantics)

2. 基于Transformer的主题导向拆分

用法语预训练模型(CamemBERT)做零样本主题分类,将同主题句子归为一类:

from transformers import pipeline

# 初始化零样本分类管道
classifier = pipeline("zero-shot-classification", model="camembert-base")

def split_by_topic(text, candidate_topics=["sport", "économie", "éducation"]):
    doc = nlp(text)
    sentences = [sent.text.strip() for sent in doc.sents if sent.text.strip()]
    if not sentences:
        return []
    
    # 对每个句子做主题匹配
    topic_groups = {topic: [] for topic in candidate_topics}
    for sent in sentences:
        result = classifier(sent, candidate_topics)
        top_topic = result["labels"][0]
        topic_groups[top_topic].append(sent)
    
    # 过滤空分组并返回
    return [" ".join(group) for group in topic_groups.values() if group]

# 应用到DataFrame
df["topic_based_units"] = df["text"].apply(split_by_topic)

二、主题标签提取

如果仅需提取主题标签而非拆分文本,推荐以下高效方案:

1. PyTextRank核心短语提取

PyTextRank支持法语,能直接从文本中提取核心短语作为主题标签:

import pytextrank

# 给spaCy添加PyTextRank管道
nlp.add_pipe("textrank")

def extract_key_tags(text, num_tags=3):
    doc = nlp(text)
    # 提取排名靠前的核心短语
    key_phrases = [phrase.text for phrase in doc._.phrases[:num_tags]]
    # 可按需映射为英文标签
    tag_map = {"sport": "sports", "économie": "economic", "éducation": "school"}
    return [tag_map.get(tag.lower(), tag) for tag in key_phrases]

# 应用到DataFrame
df["topic_tags"] = df["text"].apply(extract_key_tags)

2. LDA主题建模(批量文本场景)

针对大规模法语文本,用LDA挖掘潜在主题并提取标签:

import gensim
from gensim.corpora import Dictionary
from gensim.models import LdaModel

# 文本预处理:分词、去停用词、词形还原
def preprocess(text):
    doc = nlp(text)
    return [token.lemma_ for token in doc if not token.is_stop and not token.is_punct and token.is_alpha]

# 预处理所有文本
texts = df["text"].apply(preprocess).tolist()

# 创建词典与语料库
dictionary = Dictionary(texts)
corpus = [dictionary.doc2bow(text) for text in texts]

# 训练LDA模型
lda_model = LdaModel(corpus=corpus, id2word=dictionary, num_topics=3, random_state=42)

# 提取单条文本的主题标签
def get_lda_tags(text, num_tags=2):
    bow = dictionary.doc2bow(preprocess(text))
    top_topic = max(lda_model.get_document_topics(bow), key=lambda x: x[1])[0]
    topic_words = lda_model.show_topic(top_topic, num_words=num_tags)
    return [word for word, prob in topic_words]

df["lda_tags"] = df["text"].apply(get_lda_tags)

3. 零样本标签匹配(已知候选主题)

如果已有明确的候选标签,直接用零样本分类匹配:

def extract_zero_shot_tags(text, candidate_tags=["sports", "economic", "school"]):
    fr_candidates = ["sport", "économie", "éducation"]
    result = classifier(text, fr_candidates)
    tag_map = dict(zip(fr_candidates, candidate_tags))
    # 取置信度Top2的标签
    return [tag_map[tag] for tag in result["labels"][:2]]

df["zero_shot_tags"] = df["text"].apply(extract_zero_shot_tags)

依赖安装

执行以下命令安装所需工具:

pip install spacy pandas scikit-learn transformers pytextrank gensim
python -m spacy download fr_core_news_md

内容的提问来源于stack exchange,提问作者Paradisum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 01:35:16