如何在Pandas DataFrame单元格插入音节列表及报错排查
问题:将每个TXT文件的单词音节数列表插入DataFrame
我试过某Stack Overflow方案但无效,请勿重复该方案。目前需要统计output目录下每个txt文件的单个单词音节数,并将每个文件对应的音节数列表存入DataFrame中。
我的实现代码
directory = r"..\output" result = [] i = 0 for filename in os.listdir(directory): if filename.endswith('.txt'): filepath = os.path.join(directory, filename) with open(filepath, 'rb') as f: encoding = chardet.detect(f.read())['encoding'] with open(filepath, 'r', encoding=encoding) as f: text = f.read() words = text.split() for word in words: result.append(count_syllables(word)) results.at[i,'SYLLABLE PER WORD'] = result i += 1
运行报错信息
--------------------------------------------------------------------------- KeyError Traceback (most recent call last) ~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexes\base.py in get_loc(self, key, method, tolerance) 3801 try: -> 3802 return self._engine.get_loc(casted_key) 3803 except KeyError as err: ~\AppData\Roaming\Python\Python39\site-packages\pandas\_libs\index.pyx in pandas._libs.index.IndexEngine.get_loc() ~\AppData\Roaming\Python\Python39\site-packages\pandas\_libs\index.pyx in pandas._libs.index.IndexEngine.get_loc() pandas\_libs\hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item() pandas\_libs\hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item() KeyError: 'SYLLABLE PER WORD' The above exception was the direct cause of the following exception: KeyError Traceback (most recent call last) ~\AppData\Roaming\Python\Python39\site-packages\pandas\core\frame.py in _set_value(self, index, col, value, takeable) 4209 else: -> 4210 icol = self.columns.get_loc(col) 4211 iindex = self.index.get_loc(index) ~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexes\base.py in get_loc(self, key, method, tolerance) 3803 except KeyError as err: -> 3804 raise KeyError(key) from err 3805 except TypeError: KeyError: 'SYLLABLE PER WORD' During handling of the above exception, another exception occurred: ValueError Traceback (most recent call last) ~\AppData\Local\Temp\ipykernel_19000\1445037766.py in <module> 12 for word in words: 13 result.append(count_syllables(word)) ---> 14 results.at[i,'SYLLABLE PER WORD'] = result 15 i += 1 ~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexing.py in __setitem__(self, key, value) 2440 return 2441 -> 2442 return super().__setitem__(key, value) 2443 2444 ~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexing.py in __setitem__(self, key, value) 2395 raise ValueError("Not enough indexers for scalar access (setting)!") 2396 -> 2397 self.obj._set_value(*key, value=value, takeable=self._takeable) 2398 2399 ~\AppData\Roaming\Python\Python39\site-packages\pandas\core\frame.py in _set_value(self, index, col, value, takeable) 4222 self.iloc[index, col] = value 4223 else: -> 4224 self.loc[index, col] = value 4225 self._item_cache.pop(col, None) 4226 ~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexing.py in __setitem__(self, key, value) 816 817 iloc = self if self.name == "iloc" else self.obj.iloc -> 818 iloc._setitem_with_indexer(indexer, value, self.name) 819 820 def _validate_key(self, key, axis: int): ~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexing.py in _setitem_with_indexer(self, indexer, value, name) 1748 indexer, self.obj.axes 1749 ) -> 1750 self._setitem_with_indexer(new_indexer, value, name) 1751 1752 return ~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexing.py in _setitem_with_indexer(self, indexer, value, name) 1793 if take_split_path: 1794 # We have to operate column-wise -> 1795 self._setitem_with_indexer_split_path(indexer, value, name) 1796 else: 1797 self._setitem_single_block(indexer, value, name) ~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexing.py in _setitem_with_indexer_split_path(self, indexer, value, name) 1848 return self._setitem_with_indexer((pi, info_axis[0]), value[0]) 1849 -> 1850 raise ValueError( 1851 "Must have equal len keys and value " 1852 "when setting with an iterable" ValueError: Must have equal len keys and value when setting with an iterable
问题根源
resultsDataFrame未提前初始化,也没有创建SYLLABLE PER WORD列,直接用at赋值会触发KeyError。result列表未在每个文件处理后清空,导致每次循环累加所有历史文件的音节数,数据完全混乱。- 用
at给单元格赋值列表时,Pandas会默认将列表拆分为多个值匹配行/列,触发长度不匹配的ValueError。
修复后的代码
import os import chardet import pandas as pd def count_syllables(word): # 替换为你实际的音节统计逻辑,以下为示例实现 vowels = "aeiouyAEIOUY" syllable_count = 0 prev_char_vowel = False for char in word: if char in vowels: if not prev_char_vowel: syllable_count += 1 prev_char_vowel = True else: prev_char_vowel = False return syllable_count if syllable_count > 0 else 1 directory = r"..\output" file_data = [] for filename in os.listdir(directory): if not filename.endswith('.txt'): continue # 跳过非txt文件 filepath = os.path.join(directory, filename) # 检测文件编码 with open(filepath, 'rb') as f: encoding = chardet.detect(f.read())['encoding'] # 统计当前文件的单词音节数 current_syllables = [] with open(filepath, 'r', encoding=encoding) as f: text = f.read() for word in text.split(): current_syllables.append(count_syllables(word)) # 收集当前文件的数据 file_data.append({ '文件名': filename, '单词音节数列表': current_syllables }) # 生成最终DataFrame results = pd.DataFrame(file_data) print(results)
修复说明
- 用字典列表批量收集每个文件的数据,最后统一生成DataFrame,避免逐行赋值的各种问题。
- 每个文件处理时单独初始化音节数列表,确保每个文件的数据独立。
- 无需提前创建列,字典列表会自动生成对应的列名,符合Pandas的最佳实践。
内容的提问来源于stack exchange,提问作者Karthik Bhandary
相关产品推荐
相关产品推荐

