使用SpaCy提取Excel德语文本形容词失败问题求助
问题分析与修正方案
核心问题1:行号获取逻辑错误
原代码中iter_rows使用了values_only=True,导致row[0]是单元格的文本值而非Cell对象,调用row[0].row会触发AttributeError,根本无法正确写入列E的数据,这是多数行结果为空的主要原因。
核心问题2:形容词标签判断存在遗漏
德语形容词在SpaCy的标签体系中存在两种标注逻辑:UD标签统一为ADJ,但STTS标签会细分为ADJA(定语形容词,如schönes Haus)和ADJD(谓语形容词,如Das Haus ist schön)。仅判断pos_ == 'ADJ'会遗漏部分场景下的形容词;另外de_core_news_sm是小型模型,对复杂长文本的标注精度有限,可能出现冠词误标为形容词的情况。
修正后的代码
import openpyxl import spacy source_file_path = r'C:\Users\USERNAME\Desktop\test.xlsx' # 加载Excel文件 workbook = openpyxl.load_workbook(source_file_path) sheet = workbook.active # 加载德语模型 nlp = spacy.load('de_core_news_sm') def extract_adjectives(text): # 预处理:转换为字符串并清理空字符、换行符 clean_text = str(text).strip() if not clean_text: return "" doc = nlp(clean_text) # 同时匹配UD标签和STTS标签,覆盖更多形容词场景 adjectives = [token.text for token in doc if token.pos_ == 'ADJ' or token.tag_ in ('ADJA', 'ADJD')] # 去重避免重复提取同一形容词 unique_adjectives = list(dict.fromkeys(adjectives)) return ', '.join(unique_adjectives) # 遍历A列,用enumerate获取行号(start=1对应Excel第一行) for row_num, row in enumerate(sheet.iter_rows(min_col=1, max_col=1), start=1): cell_value = row[0].value if cell_value: adjectives = extract_adjectives(cell_value) # 写入E列(第5列) sheet.cell(row=row_num, column=5, value=adjectives) workbook.save(source_file_path) print(f'形容词已提取至文件 {source_file_path} 的E列。')
调试建议
如果仍有提取错误,可在extract_adjectives函数中加入调试代码,查看模型对每个词汇的标注结果:
def extract_adjectives(text): clean_text = str(text).strip() if not clean_text: return "" doc = nlp(clean_text) # 打印每个词的文本、POS标签、STTS标签 for token in doc: print(f"词: {token.text}, POS: {token.pos_}, TAG: {token.tag_}") adjectives = [token.text for token in doc if token.pos_ == 'ADJ' or token.tag_ in ('ADJA', 'ADJD')] unique_adjectives = list(dict.fromkeys(adjectives)) return ', '.join(unique_adjectives)
通过调试信息可以明确模型对特定词汇的标注情况,针对性调整判断逻辑。
内容的提问来源于stack exchange,提问作者DO5DKL
相关产品推荐
相关产品推荐

