如何将多轮样本统计数据转为DataFrame并实现层级索引查询?
Convert Nested Sample Stats Dictionary to Pandas DataFrame with Multi-Index
Hey there! Let's walk through converting your nested sample stats dictionary into a Pandas DataFrame that supports all the indexing and query operations you're looking for.
First, let's start with your original data structure and transform it into a multi-indexed DataFrame—this will make it easy to slice and dice your stats exactly how you want.
Step 1: Convert the Dictionary to a Multi-Index DataFrame
Here's the code to do this cleanly:
import pandas as pd # Your original data d = { "sample1": [ {"stat1": 'a', "stat2": 98}, # 1st run for sample1 {"stat1": 'z', "stat2": 13}, # 2nd run for sample1 ], "sample2": [ {"stat1": 'y', "stat2": 1089}, # 1st run for sample2 {"stat1": 'a', "stat2": 1015}, # 2nd run for sample2 ], } # Flatten the data and build a list of records records = [] for sample_name, run_stats in d.items(): for run_idx, stats in enumerate(run_stats): records.append({ "sample": sample_name, "run": run_idx, # Uses 0-based index (so 0 = 1st run, 3 = 4th run) **stats }) # Create DataFrame and set multi-index (sample + run) df = pd.DataFrame(records).set_index(["sample", "run"])
After running this, your DataFrame will look like this:
stat1 stat2 sample run sample1 0 a 98 1 z 13 sample2 0 y 1089 1 a 1015
Step 2: Verify Your Desired Queries
Now let's check that all the indexing operations you wanted work as expected:
- Get all stats for sample2:
df.loc["sample2"]returns:stat1 stat2 run 0 y 1089 1 a 1015 - Get 4th run data for sample1: (Note: Since we use 0-based indexing, the 4th run corresponds to
run=3.) Use a tuple to target the multi-index:df.loc[("sample1", 3)](this will work once you add that run to your dataset). - Get all stat1 values across all samples/runs:
df["stat1"]returns a Series with the multi-index preserved:sample run sample1 0 a 1 z sample2 0 y 1 a Name: stat1, dtype: object - Get stat2 values for sample1:
df.loc["sample1"]["stat2"](or the more efficientdf.loc["sample1", "stat2"]) returns:run 0 98 1 13 Name: stat2, dtype: int64
Step 3: Run Your Desired Statistical Queries
Now that your data is in a DataFrame, those statistical operations are straightforward:
- Average stat2 for sample1:
df.loc["sample1", "stat2"].mean()→ returns55.5for your current data. - Most common stat1 value across all samples:
df["stat1"].value_counts().idxmax()→ returns'a'(since it appears twice in your data).
内容的提问来源于stack exchange,提问作者dabadaba
相关产品推荐
相关产品推荐

