You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python的小样本自然语言识别:短字符串中法语内容识别

需求说明

需要使用Python从整体为英语、夹杂法语内容的短字符串(单条长度1到50词左右)列表中识别出法语文本,可接受假阴性(法语被误识别为英语)结果,需优先避免假阳性(英语被误识别为法语)。

输入示例
year of the snake, legendary 'dragon horse', thunder, damsel-fly, larvae of mosquito, 
treillage, libellule, mythical water creature, petites chevrettes, de papillon hideux, 
the horse-fly, 5th earthly branch, dragon, mythical creature, 
a shore plant whose leaves dry a bright orange, dragon horse, god of rain, year of the dragon, 
orthopteran, crocodile, dont le duvet des ailes s'en va en poussière, insecte, dragonfly, 
dracontomelon vitiense, dragon king, petit filet pour une espèce de papillon, sorte d'insecte
实现方案

推荐使用成熟的轻量语言检测库实现,通过调整置信度阈值满足低假阳性的要求:

  • 优先选择langdetect库,部署简单、对短文本适配性好,可自定义判定阈值过滤低置信度的结果,完全符合业务需求
    • 安装命令:pip install langdetect
    • 实现代码示例:
from langdetect import detect_langs

# 处理输入字符串,拆分出独立文本条目
input_str = """year of the snake, legendary 'dragon horse', thunder, damsel-fly, larvae of mosquito, 
treillage, libellule, mythical water creature, petites chevrettes, de papillon hideux, 
the horse-fly, 5th earthly branch, dragon, mythical creature, 
a shore plant whose leaves dry a bright orange, dragon horse, god of rain, year of the dragon, 
orthopteran, crocodile, dont le duvet des ailes s'en va en poussière, insecte, dragonfly, 
dracontomelon vitiense, dragon king, petit filet pour une espèce de papillon, sorte d'insecte"""
text_list = [t.strip() for t in input_str.split(',') if t.strip()]

# 法语判定置信度阈值,阈值越高假阳性概率越低,可根据实际效果调整
FRENCH_THRESHOLD = 0.95
french_texts = []

for text in text_list:
    try:
        detect_result = detect_langs(text)
        for res in detect_result:
            # 仅当检测为法语的置信度达标时,才判定为法语文本
            if res.lang == 'fr' and res.prob >= FRENCH_THRESHOLD:
                french_texts.append(text)
                break
    except:
        # 过短无法识别的文本直接归为英语,避免误判
        continue

print("识别出的法语文本:", french_texts)
  • 若需要更高的检测准确率,可以替换为谷歌的pycld3库,检测速度和短文本识别准确率更优,API逻辑与上述示例基本一致,同样支持置信度阈值配置。
  • 针对1-2个单词的极短文本,可额外引入常用法语单词列表做二次校验,只有文本中超过半数的单词属于法语词库时才判定为法语,可进一步降低假阳性概率。

内容的提问来源于stack exchange,提问作者Rowan Jacobs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 21:24:02