如何使用两个长度不等的列表创建符合匹配要求的pandas DataFrame
错误原因
你当前代码的逻辑是生成两个列表的笛卡尔积,总长度为len(column1)*len(column2)=27,和你需要的逐段匹配逻辑不符:你需要的是把column1按长度等分为和column2元素数量相同的段,每一段对应column2的一个元素,总长度保持len(column1)=9。
解决方案
方法1:使用numpy实现(和你原有写法兼容性最高)
import pandas as pd import numpy as np column1 = [30, 40, 50, 60,90,20,30,20,30] column2 = ['bat', 'ball','tent'] # 计算每个column2元素需要重复的次数 repeat_times = len(column1) // len(column2) # column1直接复用原列表,column2每个元素重复对应次数即可 df = pd.DataFrame({ "table1": column1, "table2": np.repeat(column2, repeat_times) }) print(df)
方法2:纯Python实现,无需依赖numpy
import pandas as pd column1 = [30, 40, 50, 60,90,20,30,20,30] column2 = ['bat', 'ball','tent'] repeat_times = len(column1) // len(column2) # 用列表推导式生成匹配的table2列表 table2 = [item for item in column2 for _ in range(repeat_times)] df = pd.DataFrame({ "table1": column1, "table2": table2 }) print(df)
两种方法运行后都可以得到你需要的预期输出:
table1 table2 0 30 bat 1 40 bat 2 50 bat 3 60 ball 4 90 ball 5 20 ball 6 30 tent 7 20 tent 8 30 tent
注意事项
使用前建议先校验len(column1) % len(column2) == 0,保证column1的长度可以被column2的长度整除,避免出现匹配错位的问题。
内容的提问来源于stack exchange,提问作者gm tom
相关产品推荐
相关产品推荐

