循环中temp_list内容更新但len(temp_list)值不更新的问题排查
拼写错误生成DataFrame的行数异常问题修复
问题背景
要创建包含word和misspelling两列的DataFrame data,通过Peter Norvig的拼写生成函数处理单词列表['a', 'is', 'the'],预期总行数为390(三个单词生成的错误拼写数量之和),但实际仅得到234行。排查发现循环中temp_list的内容会更新,但长度始终停留在首次处理'a'时的长度,尝试用temp_list.drop(columns = ['misspelling'])重置无效。
相关代码如下:
拼写生成函数
def generate(word): letters = 'abcdefghijklmnopqrstuvwxyz' splits = [(word[:i], word[i:]) for i in range(len(word) +1)] deletes = [L + R[1:] for L, R in splits if R] transposes = [L + R[1] + R[0] + R[2:] for L, R in splits if len(R)>1] replaces = [L + c + R[1:] for L, R in splits if R for c in letters] inserts = [L + c + R for L, R in splits for c in letters] return set(deletes + transposes + replaces + inserts)
初始DataFrame定义
import pandas as pd wl = ['a', 'is', 'the'] word_list = pd.DataFrame(wl, columns = ['word']) data = pd.DataFrame(columns = ['word', 'misspelling']) temp_list = pd.DataFrame(columns = ['misspelling'])
原循环代码
y = 0 for a in range(len(word_list)): temp_list['misspelling'] = pd.DataFrame(generate(word_list.at[a,'word'])) data = pd.concat([data,temp_list], ignore_index = True) print(len(temp_list)) #检查每次循环中temp_list的长度 for x in range(len(temp_list)): data.at[y,'word'] = word_list.at[a,'word'] y = y + 1 y = data.index[-1] + 1 temp_list.drop(columns = ['misspelling'])
问题根源
temp_list赋值逻辑错误:直接给temp_list['misspelling']赋值新DataFrame时,Pandas会按索引匹配数据。如果新数据的行数和原temp_list不一致,多余的行会被丢弃,不足的行填充NaN。第一次处理'a'后,temp_list的行数固定,后续处理更长的错误拼写列表时,超出原行数的部分都会被丢掉。drop方法未生效:temp_list.drop(columns = ['misspelling'])默认返回新的DataFrame,不会修改原temp_list,必须加上inplace=True或者重新赋值给temp_list才会生效,但即使生效,也不如直接重建temp_list简便。
修复方案
方案1:修改原循环,每次重建temp_list
不需要复用temp_list,每次循环直接生成新的临时DataFrame,避免索引匹配问题:
import pandas as pd def generate(word): letters = 'abcdefghijklmnopqrstuvwxyz' splits = [(word[:i], word[i:]) for i in range(len(word) +1)] deletes = [L + R[1:] for L, R in splits if R] transposes = [L + R[1] + R[0] + R[2:] for L, R in splits if len(R)>1] replaces = [L + c + R[1:] for L, R in splits if R for c in letters] inserts = [L + c + R for L, R in splits for c in letters] return set(deletes + transposes + replaces + inserts) wl = ['a', 'is', 'the'] word_list = pd.DataFrame(wl, columns = ['word']) data = pd.DataFrame(columns = ['word', 'misspelling']) for _, row in word_list.iterrows(): word = row['word'] # 直接生成当前单词的错误拼写DataFrame misspellings = pd.DataFrame(generate(word), columns=['misspelling']) # 给当前所有错误拼写添加对应的原单词 misspellings['word'] = word # 合并到结果DataFrame data = pd.concat([data, misspellings], ignore_index=True) print(len(data)) # 输出390,符合预期
方案2:用Pandas原生方法简化实现(推荐)
避免手动循环,用apply生成错误拼写列表,再用explode展开,代码更简洁高效:
import pandas as pd def generate(word): letters = 'abcdefghijklmnopqrstuvwxyz' splits = [(word[:i], word[i:]) for i in range(len(word) +1)] deletes = [L + R[1:] for L, R in splits if R] transposes = [L + R[1] + R[0] + R[2:] for L, R in splits if len(R)>1] replaces = [L + c + R[1:] for L, R in splits if R for c in letters] inserts = [L + c + R for L, R in splits for c in letters] return list(set(deletes + transposes + replaces + inserts)) # 返回列表而非集合 wl = ['a', 'is', 'the'] word_list = pd.DataFrame(wl, columns = ['word']) # 生成错误拼写列表列,再展开 data = word_list.assign(misspelling=word_list['word'].apply(generate)).explode('misspelling').reset_index(drop=True) print(len(data)) # 输出390,符合预期
内容的提问来源于stack exchange,提问作者michi
相关产品推荐
相关产品推荐

