如何在嵌套DataFrame中使用LabelEncoder编码麦芽和酒花名称
对嵌套DataFrame中的分类列进行LabelEncoder编码的实现方案
场景说明
已加载啤酒配方数据集,并将主DataFrame中的fermentables和hops列转换为嵌套DataFrame结构,需要对所有行的嵌套DataFrame内的Malt(麦芽名称)和hop(酒花名称)列统一使用LabelEncoder进行编码。
实现步骤及代码
导入所需库
补充导入机器学习编码工具及pandas库:import json import pandas as pd from sklearn.preprocessing import LabelEncoder加载并预处理数据
沿用你提供的数据集加载与嵌套结构转换代码:filename = 'recipes_full copy.json' with open(filename, 'r') as f: try: json_data = json.load(f) print("The JSON file is valid") except ValueError as e: print("The JSON file is invalid:", e) df = pd.DataFrame(json_data.values()) # 转换为嵌套DataFrame df['fermentables'] = df['fermentables'].apply(pd.DataFrame, columns=["kg","Malt","ppg", "°L Degree Lintner", "bill"]) df['hops'] = df['hops'].apply(pd.DataFrame, columns=["grams", "hop","hoptype", " % AA", "Type", "Time", "IBU", "Percentage"])全局收集分类唯一值
为保证同一名称在所有嵌套DataFrame中编码一致,先收集所有目标列的唯一值:# 收集所有麦芽名称的唯一值 all_malts = pd.concat(df['fermentables'].tolist())['Malt'].unique() # 收集所有酒花名称的唯一值 all_hops = pd.concat(df['hops'].tolist())['hop'].unique()初始化并拟合编码器
# 分别初始化编码器 malt_encoder = LabelEncoder() hop_encoder = LabelEncoder() # 用全局唯一值拟合编码器 malt_encoder.fit(all_malts) hop_encoder.fit(all_hops)编写编码函数并应用
# 定义麦芽列编码函数 def encode_malt(nested_df): nested_df['Malt_encoded'] = malt_encoder.transform(nested_df['Malt']) return nested_df # 定义酒花列编码函数 def encode_hop(nested_df): nested_df['hop_encoded'] = hop_encoder.transform(nested_df['hop']) return nested_df # 将编码逻辑应用到所有嵌套DataFrame df['fermentables'] = df['fermentables'].apply(encode_malt) df['hops'] = df['hops'].apply(encode_hop)
关键说明
- 必须先全局收集分类值再拟合编码器,避免不同嵌套DataFrame中同一名称出现编码不一致的问题。
- 编码后的列会以
Malt_encoded和hop_encoded的名称添加到对应嵌套DataFrame中,保留原分类列方便后续追溯。
内容的提问来源于stack exchange,提问作者Luis Valencia
相关产品推荐
相关产品推荐

