使用nx.from_pandas_edgelist处理CSV数据时触发KeyError问题
解决NetworkX from_pandas_edgelist抛出KeyError: 'weight'的问题
问题背景
有UTF-8编码的CSV文件connections.csv,内容如下:
source, target, weight 0, 1, 100 1, 2, 50
在JupyterLab中调用nx.from_pandas_edgelist(foo, 'source', 'target', 'weight')时抛出KeyError: 'weight',最终触发NetworkXError,但另一个结构看起来完全一致的DataFramebar却能正常处理。
完整错误信息
--------------------------------------------------------------------------- KeyError Traceback (most recent call last) File /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3802, in Index.get_loc(self, key, method, tolerance) 3801 try: -> 3802 return self._engine.get_loc(casted_key) 3803 except KeyError as err: File /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:138, in pandas._libs.index.IndexEngine.get_loc() File /opt/conda/lib/python3.10/site-packages/pandas/_libs/index.pyx:165, in pandas._libs.index.IndexEngine.get_loc() File pandas/_libs/hashtable_class_helper.pxi:5745, in pandas._libs.hashtable.PyObjectHashTable.get_item() File pandas/_libs/hashtable_class_helper.pxi:5753, in pandas._libs.hashtable.PyObjectHashTable.get_item() KeyError: 'weight' The above exception was the direct cause of the following exception: KeyError Traceback (most recent call last) File /opt/conda/lib/python3.10/site-packages/networkx/convert_matrix.py:435, in from_pandas_edgelist(df, source, target, edge_attr, create_using, edge_key) 434 try: -> 435 attribute_data = zip(*[df[col] for col in attr_col_headings]) 436 except (KeyError, TypeError) as err: File /opt/conda/lib/python3.10/site-packages/networkx/convert_matrix.py:435, in <listcomp>(.0) 434 try: -> 435 attribute_data = zip(*[df[col] for col in attr_col_headings]) 436 except (KeyError, TypeError) as err: File /opt/conda/lib/python3.10/site-packages/pandas/core/frame.py:3807, in DataFrame.__getitem__(self, key) 3806 return self._getitem_multilevel(key) -> 3807 indexer = self.columns.get_loc(key) 3808 if is_integer(indexer): File /opt/conda/lib/python3.10/site-packages/pandas/core/indexes/base.py:3804, in Index.get_loc(self, key, method, tolerance) 3803 except KeyError as err: -> 3804 raise KeyError(key) from err 3805 except TypeError: 3806 # If we have a listlike key, _check_indexing_error will raise 3807 # InvalidIndexError. Otherwise we fall through and re-raise 3808 # the TypeError. KeyError: 'weight' The above exception was the direct cause of the following exception: NetworkXError Traceback (most recent call last) Cell In[9], line 1 ----> 1 F = nx.from_pandas_edgelist(foo, 'source', 'target', 'weight') File /opt/conda/lib/python3.10/site-packages/networkx/convert_matrix.py:438, in from_pandas_edgelist(df, source, target, edge_attr, create_using, edge_key) 436 except (KeyError, TypeError) as err: 437 msg = f"Invalid edge_attr argument: {edge_attr}" -> 438 raise nx.NetworkXError(msg) from err 440 if g.is_multigraph(): 441 # => append the edge keys from the df to the bundled data 442 if edge_key is not None: NetworkXError: Invalid edge_attr argument: weight
解决方法
1. 排查列名隐藏差异
看起来一致的DataFrame,实际列名可能存在空格、制表符或其他不可见字符,导致'weight'无法匹配。执行以下代码检查列名的原始形式:
# 打印所有列名的原始表示,查看是否有隐藏字符 print([repr(col) for col in foo.columns])
如果输出类似['source', ' target', ' weight'],说明列名包含前导空格。
2. 清理列名
如果发现列名有多余字符,执行以下代码清理:
# 去除列名前后的空白字符 foo.columns = foo.columns.str.strip() # 或者针对特定列重命名 foo = foo.rename(columns={' weight': 'weight', ' target': 'target'})
3. 优化CSV读取方式
如果是读取CSV时引入的问题,下次读取时添加skipinitialspace=True参数,自动忽略列名和字段的前后空格:
import pandas as pd foo = pd.read_csv('connections.csv', skipinitialspace=True)
4. 验证修复结果
清理后再次检查列名,确认无误后调用NetworkX函数:
print(foo.columns) # 应输出: Index(['source', 'target', 'weight'], dtype='object') import networkx as nx F = nx.from_pandas_edgelist(foo, 'source', 'target', 'weight')
内容的提问来源于stack exchange,提问作者jerik
相关产品推荐
相关产品推荐

