You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Spacy提取城市遇阻:无法加载法语模型及报错求助

问题修复方案

一、法语Spacy模型加载失败的解决

  1. 安装命令错误:你写的pip install spacy download fr是无效命令,正确的法语模型安装步骤是:
    pip install spacy
    python -m spacy download fr_core_news_sm
    
  2. 自定义lemmatizer未返回实例:你的create_french_lemmatizer函数注释了return语句,导致添加管道时没有有效组件,取消注释即可:
    @Language.factory('french_lemmatizer')
    def create_french_lemmatizer(nlp, name):
        return LefffLemmatizer()
    

二、TypeError: decoding str is not supported的解决

这个错误来自wikipedia.summary的调用方式错误:str(wikipedia.summary(text),"html.parser")里str()的第二个参数不是解码格式,且wikipedia.summary默认返回纯文本,无需额外转str。修改为:

summary = wikipedia.summary(text)

如果要处理维基百科的异常情况(比如条目不存在、歧义),可以加try-except:

for text in gpe:
    try:
        summary = wikipedia.summary(text)
        if 'city' in summary.lower():  # 转小写避免大小写匹配问题
            cities.append(text)
        elif 'country' in summary.lower():
            countries.append(text)
        else:
            other_places.append(text)
    except wikipedia.exceptions.DisambiguationError:
        other_places.append(text)
    except wikipedia.exceptions.PageError:
        other_places.append(text)

三、其他代码问题修复

  1. 变量未初始化:gpe和loc在使用前没有定义,要在循环前先初始化:
    gpe = []
    loc = []
    
  2. 模型覆盖问题:你先加载法语模型,后又加载英文模型en_core_web_sm,会覆盖之前的nlp对象。如果处理的是法语文本(比如《悲惨世界》原文),全程用法语模型才能正确识别实体。
  3. 文件读取编码:打开文件时指定编码避免乱码:
    with open("/content/drive/My Drive/Miserables/miserable.txt", 'r', encoding='utf-8') as f:
        myString = f.read()
    

完整修复后的代码示例

# 先在终端运行安装命令
# pip install spacy spacy_lefff wikipedia
# python -m spacy download fr_core_news_sm

import spacy
from spacy_lefff import LefffLemmatizer
from spacy.language import Language
import os
from google.colab import drive
import wikipedia

# 初始化法语lemmatizer组件
@Language.factory('french_lemmatizer')
def create_french_lemmatizer(nlp, name):
    return LefffLemmatizer()

# 加载法语模型(处理法语文本用这个)
nlp = spacy.load('fr_core_news_sm')
nlp.add_pipe('french_lemmatizer', name='lefff')

# 挂载Google Drive
drive.mount('/content/drive/', force_remount=True)
if not os.path.exists('/content/drive/My Drive/Miserables'):
    os.makedirs('/content/drive/My Drive/Miserables')

# 读取目标文本文件
file_path = '/content/drive/My Drive/Miserables/miserable.txt'
with open(file_path, 'r', encoding='utf-8') as f:
    text_content = f.read()

# 处理文本提取地理实体
doc = nlp(text_content)
gpe = []
loc = []
for ent in doc.ents:
    if ent.label_ == 'GPE':
        gpe.append(ent.text)
    elif ent.label_ == 'LOC':
        loc.append(ent.text)

# 分类提取到的地点
cities = []
countries = []
other_places = []

for place in gpe:
    try:
        summary = wikipedia.summary(place)
        summary_lower = summary.lower()
        # 同时考虑法语和英文关键词,适配维基百科的多语言摘要
        if 'ville' in summary_lower or 'city' in summary_lower:
            cities.append(place)
        elif 'pays' in summary_lower or 'country' in summary_lower:
            countries.append(place)
        else:
            other_places.append(place)
    except (wikipedia.exceptions.DisambiguationError, wikipedia.exceptions.PageError):
        other_places.append(place)

# 将LOC类别的地点归入其他位置
other_places.extend(loc)

# 输出结果
print("提取到的城市:", cities)
print("提取到的国家:", countries)
print("其他地点:", other_places)

内容的提问来源于stack exchange,提问作者NoobWithPython

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 01:40:26