如何为df.itertuples的输出正确添加Python类型注解?
如何为pandas itertuples返回的对象添加正确的类型注解?
我的代码及输出
我编写了如下Python脚本:
import pandas as pd from typing import Any def info(i: Any) -> None: print(f'{type(i)=:}') print(f'{i=:}') print(f'{i.Index=:}') print(f'{i.x=:}') print(f'{i.y=:}') if __name__ == "__main__": df = pd.DataFrame([[1,'a'], [2, 'b']], columns=['x', 'y']) for i in df.itertuples(): info(i)
运行后输出为:
type(i)=<class 'pandas.core.frame.Pandas'> i=Pandas(Index=0, x=1, y='a') i.Index=0 i.x=1 i.y=a type(i)=<class 'pandas.core.frame.Pandas'> i=Pandas(Index=1, x=2, y='b') i.Index=1 i.x=2 i.y=b
尝试的错误
我希望避免使用Any类型,按照pandas的类型标注方式(Python3.8)尝试用Tuple[Any, ...]注解:
def info(i: Tuple[Any, ...]) -> None:
但mypy报错:
toy.py:8:11: error: "Tuple[Any, ...]" has no attribute "Index"; maybe "index"? [attr-defined] toy.py:9:11: error: "Tuple[Any, ...]" has no attribute "x" [attr-defined] toy.py:10:11: error: "Tuple[Any, ...]" has no attribute "y" [attr-defined] Found 3 errors in 1 file (checked 1 source file)
问题
请问正确的类型注解方式是什么?
解决方案
方案1:使用Protocol(推荐,兼顾兼容性与类型检查)
itertuples()返回的是自定义命名元组实例,我们可以通过Protocol定义一个包含所需属性的接口,让类型检查工具识别这些字段。由于Python3.8的typing模块未内置Protocol,需要先安装typing_extensions:
pip install typing_extensions
修改后的代码:
import pandas as pd from typing_extensions import Protocol class DataRow(Protocol): Index: int x: int y: str def info(i: DataRow) -> None: print(f'{type(i)=:}') print(f'{i=:}') print(f'{i.Index=:}') print(f'{i.x=:}') print(f'{i.y=:}') if __name__ == "__main__": df = pd.DataFrame([[1,'a'], [2, 'b']], columns=['x', 'y']) for i in df.itertuples(): info(i)
此方案不依赖pandas内部实现,兼容性更强,mypy可以正常识别所有属性。
方案2:直接引用pandas内部类型
从输出可知,itertuples()返回的是pandas.core.frame.Pandas类的实例,可直接用该类型注解:
import pandas as pd from pandas.core.frame import Pandas def info(i: Pandas) -> None: print(f'{type(i)=:}') print(f'{i=:}') print(f'{i.Index=:}') print(f'{i.x=:}') print(f'{i.y=:}') if __name__ == "__main__": df = pd.DataFrame([[1,'a'], [2, 'b']], columns=['x', 'y']) for i in df.itertuples(): info(i)
注意:Pandas是pandas内部实现类,版本迭代中可能发生变更,长期兼容性不如Protocol方案。
方案3:自定义NamedTuple
如果数据结构固定,可定义与返回结构匹配的NamedTuple,类型检查工具会因结构兼容通过检查:
import pandas as pd from typing import NamedTuple class DataRow(NamedTuple): Index: int x: int y: str def info(i: DataRow) -> None: print(f'{type(i)=:}') print(f'{i=:}') print(f'{i.Index=:}') print(f'{i.x=:}') print(f'{i.y=:}') if __name__ == "__main__": df = pd.DataFrame([[1,'a'], [2, 'b']], columns=['x', 'y']) for i in df.itertuples(): info(i)
内容的提问来源于stack exchange,提问作者zyxue
相关产品推荐
相关产品推荐

