You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame单元格插入音节列表及报错排查

问题:将每个TXT文件的单词音节数列表插入DataFrame

我试过某Stack Overflow方案但无效,请勿重复该方案。目前需要统计output目录下每个txt文件的单个单词音节数,并将每个文件对应的音节数列表存入DataFrame中。

我的实现代码

directory = r"..\output" 
result = []
i = 0
for filename in os.listdir(directory):
    if filename.endswith('.txt'):
        filepath = os.path.join(directory, filename)
    with open(filepath, 'rb') as f:
            encoding = chardet.detect(f.read())['encoding']
    with open(filepath, 'r', encoding=encoding) as f:
            text = f.read()
            words = text.split()
            for word in words:
                result.append(count_syllables(word))
    results.at[i,'SYLLABLE PER WORD'] = result
    i += 1

运行报错信息

---------------------------------------------------------------------------
KeyError                                  Traceback (most recent call last)
~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexes\base.py in get_loc(self, key, method, tolerance)
   3801             try:
-> 3802                 return self._engine.get_loc(casted_key)
   3803             except KeyError as err:

~\AppData\Roaming\Python\Python39\site-packages\pandas\_libs\index.pyx in pandas._libs.index.IndexEngine.get_loc()

~\AppData\Roaming\Python\Python39\site-packages\pandas\_libs\index.pyx in pandas._libs.index.IndexEngine.get_loc()

pandas\_libs\hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()

pandas\_libs\hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()

KeyError: 'SYLLABLE PER WORD'

The above exception was the direct cause of the following exception:

KeyError                                  Traceback (most recent call last)
~\AppData\Roaming\Python\Python39\site-packages\pandas\core\frame.py in _set_value(self, index, col, value, takeable)
   4209             else:
-> 4210                 icol = self.columns.get_loc(col)
   4211                 iindex = self.index.get_loc(index)

~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexes\base.py in get_loc(self, key, method, tolerance)
   3803             except KeyError as err:
-> 3804                 raise KeyError(key) from err
   3805             except TypeError:

KeyError: 'SYLLABLE PER WORD'

During handling of the above exception, another exception occurred:

ValueError                                Traceback (most recent call last)
~\AppData\Local\Temp\ipykernel_19000\1445037766.py in <module>
     12             for word in words:
     13                 result.append(count_syllables(word))
---&gt; 14     results.at[i,'SYLLABLE PER WORD'] = result
     15     i += 1

~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexing.py in __setitem__(self, key, value)
   2440             return
   2441 
-> 2442         return super().__setitem__(key, value)
   2443 
   2444 

~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexing.py in __setitem__(self, key, value)
   2395             raise ValueError("Not enough indexers for scalar access (setting)!")
   2396 
-> 2397         self.obj._set_value(*key, value=value, takeable=self._takeable)
   2398 
   2399 

~\AppData\Roaming\Python\Python39\site-packages\pandas\core\frame.py in _set_value(self, index, col, value, takeable)
   4222                 self.iloc[index, col] = value
   4223             else:
-> 4224                 self.loc[index, col] = value
   4225             self._item_cache.pop(col, None)
   4226 

~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexing.py in __setitem__(self, key, value)
    816 
    817         iloc = self if self.name == "iloc" else self.obj.iloc
-> 818         iloc._setitem_with_indexer(indexer, value, self.name)
    819 
    820     def _validate_key(self, key, axis: int):

~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexing.py in _setitem_with_indexer(self, indexer, value, name)
   1748                             indexer, self.obj.axes
   1749                         )
-> 1750                         self._setitem_with_indexer(new_indexer, value, name)
   1751 
   1752                         return

~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexing.py in _setitem_with_indexer(self, indexer, value, name)
   1793         if take_split_path:
   1794             # We have to operate column-wise
-> 1795             self._setitem_with_indexer_split_path(indexer, value, name)
   1796         else:
   1797             self._setitem_single_block(indexer, value, name)

~\AppData\Roaming\Python\Python39\site-packages\pandas\core\indexing.py in _setitem_with_indexer_split_path(self, indexer, value, name)
   1848                     return self._setitem_with_indexer((pi, info_axis[0]), value[0])
   1849 
-> 1850                 raise ValueError(
   1851                     "Must have equal len keys and value "
   1852                     "when setting with an iterable"

ValueError: Must have equal len keys and value when setting with an iterable

问题根源

  1. results DataFrame未提前初始化,也没有创建SYLLABLE PER WORD列,直接用at赋值会触发KeyError。
  2. result列表未在每个文件处理后清空,导致每次循环累加所有历史文件的音节数,数据完全混乱。
  3. 用at给单元格赋值列表时,Pandas会默认将列表拆分为多个值匹配行/列,触发长度不匹配的ValueError。

修复后的代码

import os
import chardet
import pandas as pd

def count_syllables(word):
    # 替换为你实际的音节统计逻辑,以下为示例实现
    vowels = "aeiouyAEIOUY"
    syllable_count = 0
    prev_char_vowel = False
    for char in word:
        if char in vowels:
            if not prev_char_vowel:
                syllable_count += 1
            prev_char_vowel = True
        else:
            prev_char_vowel = False
    return syllable_count if syllable_count > 0 else 1

directory = r"..\output" 
file_data = []

for filename in os.listdir(directory):
    if not filename.endswith('.txt'):
        continue  # 跳过非txt文件
    filepath = os.path.join(directory, filename)
    
    # 检测文件编码
    with open(filepath, 'rb') as f:
        encoding = chardet.detect(f.read())['encoding']
    
    # 统计当前文件的单词音节数
    current_syllables = []
    with open(filepath, 'r', encoding=encoding) as f:
        text = f.read()
        for word in text.split():
            current_syllables.append(count_syllables(word))
    
    # 收集当前文件的数据
    file_data.append({
        '文件名': filename,
        '单词音节数列表': current_syllables
    })

# 生成最终DataFrame
results = pd.DataFrame(file_data)
print(results)

修复说明

  • 用字典列表批量收集每个文件的数据,最后统一生成DataFrame,避免逐行赋值的各种问题。
  • 每个文件处理时单独初始化音节数列表,确保每个文件的数据独立。
  • 无需提前创建列,字典列表会自动生成对应的列名,符合Pandas的最佳实践。

内容的提问来源于stack exchange,提问作者Karthik Bhandary

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 03:07:08