使用spaCy提取含指定连词的句子时遭遇TypeError错误求助
问题修复:提取含指定连词的句子
错误原因
你代码里的if any(conjunction in sentence for conjunction in conjunctions)这行触发了类型错误:spaCy的sentence是Span对象,使用in操作符时,它会检查Token对象是否存在于Span中,而你传入的是字符串(比如'and'),类型不匹配,因此抛出TypeError。
另外还有个隐藏问题:coord_sents = []放在循环内部,每次遍历句子都会重置为空列表,最终只会返回最后一个符合条件的句子,必须把列表初始化移到循环外面。
修正方案
方案1:检查Token文本(简单直接)
遍历句子中的每个Token,将Token文本转为小写后匹配连词列表:
import spacy lang_model = spacy.load("en_core_web_sm") text_to_look = "A woman is looking at books in a library. She's looking to buy one, but she hasn't got any money. She really wanted to book, so she asks another customer to lend her money. The man accepts. They get along really well, so they both exchange phone numbers and go their separate ways." def get_coordinate_sents(file_to_examine): conjunctions = ['and', 'but', 'for', 'nor', 'or', 'yet', 'so'] text = lang_model(file_to_examine) coord_sents = [] # 移到循环外,避免每次清空 for sentence in text.sents: # 遍历句子内的Token,检查文本是否在连词列表中 if any(token.text.lower() in conjunctions for token in sentence): coord_sents.append(sentence.text) # 要Span对象就直接append(sentence) return coord_sents wanted_sents = get_coordinate_sents(text_to_look) print(wanted_sents)
方案2:利用词性标注(更精准)
spaCy中并列连词的词性标签为CC,直接检查句子内是否存在该标签的Token,可避免误判(比如单词band包含and字符串的情况):
import spacy lang_model = spacy.load("en_core_web_sm") text_to_look = "A woman is looking at books in a library. She's looking to buy one, but she hasn't got any money. She really wanted to book, so she asks another customer to lend her money. The man accepts. They get along really well, so they both exchange phone numbers and go their separate ways." def get_coordinate_sents(file_to_examine): text = lang_model(file_to_examine) coord_sents = [] for sentence in text.sents: # 检查句子中是否有并列连词(词性标签CC) if any(token.pos_ == "CC" for token in sentence): coord_sents.append(sentence.text) return coord_sents wanted_sents = get_coordinate_sents(text_to_look) print(wanted_sents)
内容的提问来源于stack exchange,提问作者user17169994
相关产品推荐
相关产品推荐

