如何用PySpark或Pandas生成满足beds>=baths>=cars的列组合
实现方法
可以用itertools.product生成三个列表的所有笛卡尔积组合,再筛选符合beds >= baths >= cars条件的结果,最后转为Pandas DataFrame。具体步骤如下:
代码实现
import pandas as pd from itertools import product # 给定的原始列表 beds = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0] baths = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0] cars = [0, 1, 2, 3, 4, 5, 6, 7, 8] # 生成所有可能的组合,并筛选符合条件的 valid_combinations = [ (bed, bath, car) for bed, bath, car in product(beds, baths, cars) if bed >= bath and bath >= car ] # 转换为DataFrame并设置列名 df = pd.DataFrame(valid_combinations, columns=['nbed', 'nbath', 'ncar']) # 查看前20行(和示例结果一致) print(df.head(20))
另一种Pandas原生方法
如果不想用itertools,也可以通过多次交叉合并+条件筛选实现:
import pandas as pd beds = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0] baths = [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0] cars = [0, 1, 2, 3, 4, 5, 6, 7, 8] # 创建各列的DataFrame df_beds = pd.DataFrame({'nbed': beds}) df_baths = pd.DataFrame({'nbath': baths}) df_cars = pd.DataFrame({'ncar': cars}) # 交叉合并所有可能的组合 df_all = df_beds.merge(df_baths, how='cross').merge(df_cars, how='cross') # 筛选符合条件的行 df_valid = df_all.query('nbed >= nbath and nbath >= ncar') # 重置索引(可选) df_valid = df_valid.reset_index(drop=True) print(df_valid.head(20))
结果说明
两种方法都会生成符合要求的数据集,其中:
- 当
nbed=1.0时,nbath只能取1.0,ncar可取0、1(满足1.0>=1.0>=car) - 当
nbed=2.0时,nbath可取1.0、2.0,对应ncar分别可取0-1、0-2 - 以此类推,完全匹配你给出的示例结果。
内容的提问来源于stack exchange,提问作者Phikho
相关产品推荐
相关产品推荐

