使用Spacy提取城市遇阻:无法加载法语模型及报错求助
问题修复方案
一、法语Spacy模型加载失败的解决
- 安装命令错误:你写的
pip install spacy download fr是无效命令,正确的法语模型安装步骤是:pip install spacy python -m spacy download fr_core_news_sm - 自定义lemmatizer未返回实例:你的
create_french_lemmatizer函数注释了return语句,导致添加管道时没有有效组件,取消注释即可:@Language.factory('french_lemmatizer') def create_french_lemmatizer(nlp, name): return LefffLemmatizer()
二、TypeError: decoding str is not supported的解决
这个错误来自wikipedia.summary的调用方式错误:str(wikipedia.summary(text),"html.parser")里str()的第二个参数不是解码格式,且wikipedia.summary默认返回纯文本,无需额外转str。修改为:
summary = wikipedia.summary(text)
如果要处理维基百科的异常情况(比如条目不存在、歧义),可以加try-except:
for text in gpe: try: summary = wikipedia.summary(text) if 'city' in summary.lower(): # 转小写避免大小写匹配问题 cities.append(text) elif 'country' in summary.lower(): countries.append(text) else: other_places.append(text) except wikipedia.exceptions.DisambiguationError: other_places.append(text) except wikipedia.exceptions.PageError: other_places.append(text)
三、其他代码问题修复
- 变量未初始化:
gpe和loc在使用前没有定义,要在循环前先初始化:gpe = [] loc = [] - 模型覆盖问题:你先加载法语模型,后又加载英文模型
en_core_web_sm,会覆盖之前的nlp对象。如果处理的是法语文本(比如《悲惨世界》原文),全程用法语模型才能正确识别实体。 - 文件读取编码:打开文件时指定编码避免乱码:
with open("/content/drive/My Drive/Miserables/miserable.txt", 'r', encoding='utf-8') as f: myString = f.read()
完整修复后的代码示例
# 先在终端运行安装命令 # pip install spacy spacy_lefff wikipedia # python -m spacy download fr_core_news_sm import spacy from spacy_lefff import LefffLemmatizer from spacy.language import Language import os from google.colab import drive import wikipedia # 初始化法语lemmatizer组件 @Language.factory('french_lemmatizer') def create_french_lemmatizer(nlp, name): return LefffLemmatizer() # 加载法语模型(处理法语文本用这个) nlp = spacy.load('fr_core_news_sm') nlp.add_pipe('french_lemmatizer', name='lefff') # 挂载Google Drive drive.mount('/content/drive/', force_remount=True) if not os.path.exists('/content/drive/My Drive/Miserables'): os.makedirs('/content/drive/My Drive/Miserables') # 读取目标文本文件 file_path = '/content/drive/My Drive/Miserables/miserable.txt' with open(file_path, 'r', encoding='utf-8') as f: text_content = f.read() # 处理文本提取地理实体 doc = nlp(text_content) gpe = [] loc = [] for ent in doc.ents: if ent.label_ == 'GPE': gpe.append(ent.text) elif ent.label_ == 'LOC': loc.append(ent.text) # 分类提取到的地点 cities = [] countries = [] other_places = [] for place in gpe: try: summary = wikipedia.summary(place) summary_lower = summary.lower() # 同时考虑法语和英文关键词,适配维基百科的多语言摘要 if 'ville' in summary_lower or 'city' in summary_lower: cities.append(place) elif 'pays' in summary_lower or 'country' in summary_lower: countries.append(place) else: other_places.append(place) except (wikipedia.exceptions.DisambiguationError, wikipedia.exceptions.PageError): other_places.append(place) # 将LOC类别的地点归入其他位置 other_places.extend(loc) # 输出结果 print("提取到的城市:", cities) print("提取到的国家:", countries) print("其他地点:", other_places)
内容的提问来源于stack exchange,提问作者NoobWithPython
相关产品推荐
相关产品推荐

