使用pandas的df.iloc选择单列触发ValueError的原因解析
为什么
df.iloc[0:1, 'A']会触发ValueError? 咱们先从你给出的例子说起:你用日期索引创建了DataFrame后,df.loc['20130101':'20130102', 'A']能正常返回结果,但换成df.iloc[0:1, 'A']就直接抛出了ValueError,核心原因其实是**iloc和loc的设计逻辑完全不同**:
先明确两个索引器的核心定位
loc是标签索引器:它完全靠行/列的标签名称定位数据,所以你传日期字符串(行标签)和列名'A'(列标签)完全符合它的规则。iloc是位置索引器:它只认行/列的整数位置(从0开始计数),不接受任何标签类参数——比如列名、日期字符串这类都不属于它的合法输入范围。
你的错误用法哪里出问题了?
你在iloc的第二个参数传了'A',这是列的标签名称,不是它对应的整数位置,直接违反了iloc的使用规则,所以它抛出的错误信息已经说得很直白:
Location based indexing can only have [integer, integer slice (START point is INCLUDED, END point is EXCLUDED), listlike of integers, boolean array] types
翻译过来就是:基于位置的索引只能用整数、整数切片、整数列表/数组或者布尔数组类型的参数。
正确的iloc用法
如果想用iloc实现和df.loc['20130101':'20130102', 'A']一样的效果,你需要把列名'A'换成它的整数位置(这里'A'是第0列):
# 注意iloc是左闭右开切片,要取前两行得写0:2 df.iloc[0:2, 0]
如果不确定列的位置,也可以用df.columns.get_loc('A')动态获取列的位置索引:
col_pos = df.columns.get_loc('A') df.iloc[0:2, col_pos]
回顾你的完整代码示例
创建DataFrame的代码
import pandas as pd import numpy as np dates = pd.date_range('20130101', periods=6) df = pd.DataFrame(np.random.randn(6, 4), index=dates, columns=list('ABCD'))
正确的loc调用及输出
df.loc['20130101':'20130102', 'A']
输出:
2013-01-01 0.469112 2013-01-02 1.212112 Freq: D, Name: A, dtype: float64
错误的iloc调用触发的完整报错
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) /opt/conda/lib/python3.6/site-packages/pandas/core/indexing.py in _has_valid_tuple(self, key) 221 try: ---> 222 self._validate_key(k, i) 223 except ValueError: /opt/conda/lib/python3.6/site-packages/pandas/core/indexing.py in _validate_key(self, key, axis) 1970 raise ValueError("Can only index by location with " -> 1971 "a [{types}]".format(types=self._valid_types)) 1972 ValueError: Can only index by location with a [integer, integer slice (START point is INCLUDED, END point is EXCLUDED), listlike of integers, boolean array] During handling of the above exception, another exception occurred: ValueError Traceback (most recent call last) <ipython-input-73-c05fbef69c91> in <module>() ----> 1 df.iloc[0:1, 'A'] /opt/conda/lib/python3.6/site-packages/pandas/core/indexing.py in __getitem__(self, key) 1470 except (KeyError, IndexError): 1471 pass -> 1472 return self._getitem_tuple(key) 1473 else: 1474 # we by definition only have the 0th axis /opt/conda/lib/python3.6/site-packages/pandas/core/indexing.py in _getitem_tuple(self, tup) 2011 def _getitem_tuple(self, tup): 2012 -> 2013 self._has_valid_tuple(tup) 2014 try: 2015 return self._getitem_lowerdim(tup) /opt/conda/lib/python3.6/site-packages/pandas/core/indexing.py in _has_valid_tuple(self, key) 224 raise ValueError("Location based indexing can only have " 225 "[{types}] types" --> 226 .format(types=self._valid_types)) 227 228 def _is_nested_tuple_indexer(self, tup): ValueError: Location based indexing can only have [integer, integer slice (START point is INCLUDED, END point is EXCLUDED), listlike of integers, boolean array] types
内容的提问来源于stack exchange,提问作者goedi
相关产品推荐
相关产品推荐

