You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中高效提取字典字符串中的患者信息并生成访问计数列?

Great question! Your current solution gets the job done, but it’s not the most efficient for large datasets—those repeated apply(axis=1) calls and patient-wise loops can get slow quickly. Let’s refactor this into a more idiomatic, performant Pandas workflow that avoids unnecessary iteration and uses safer parsing.

Key Improvements We’ll Make:

  • Use ast.literal_eval instead of raw eval (safer for untrusted data)
  • Leverage Pandas’ vectorized operations (explode, crosstab) instead of逐行 iteration
  • Avoid looping through each patient to count visits

Solution 1: Using explode + crosstab (Most Efficient)

This approach fully leverages Pandas’ optimized backend and minimizes manual iteration:

import pandas as pd
import ast

# Original DataFrame setup
d1 = str([{"Alice": "1/2"}, {"Bob": "3/11"}, {"Charlotte": "5/2"}])
d2 = str([{"Bob": "2/7"}, {"Daniel": "4/5"}])
d3 = str([{"Emily": "12/5"}])
d4 = str([{"Bob": "4/11"}, {"Victor": "2/2"}, {"Bob": "5/6"}])
df = pd.DataFrame([d1, d2, d3, d4], columns=['visits'])

# Step 1: Parse stringified list of dicts into actual Python objects
df['visits_parsed'] = df['visits'].apply(ast.literal_eval)

# Step 2: Explode each dict in the list into its own row
exploded_df = df.explode('visits_parsed')

# Step 3: Extract patient names from each dict (each dict has one key)
exploded_df['patient'] = exploded_df['visits_parsed'].apply(lambda d: next(iter(d.keys())))

# Step 4: Count visits per patient per original row (using crosstab)
patient_counts = pd.crosstab(exploded_df.index, exploded_df['patient']).astype(int)

# Step 5: Combine original visits column with counts
df_desired = pd.concat([df['visits'], patient_counts], axis=1)

Solution 2: Using value_counts per Row (Simpler for Smaller Data)

If you prefer a more concise approach for smaller datasets, this works well too:

import pandas as pd
import ast

# Original DataFrame setup (same as above)
df = pd.DataFrame([d1, d2, d3, d4], columns=['visits'])

# Parse and extract patient lists (including duplicates)
df['patients'] = df['visits'].apply(
    lambda x: [next(iter(d.keys())) for d in ast.literal_eval(x)]
)

# Count occurrences per patient per row, then reshape to columns
patient_counts = df['patients'].apply(pd.Series.value_counts).fillna(0).astype(int)

# Combine with original data
df_desired = pd.concat([df['visits'], patient_counts], axis=1)

Why This Is Better Than Your Original Approach:

  1. Safety: ast.literal_eval only evaluates literal Python objects (lists, dicts, strings) instead of arbitrary code, eliminating security risks from raw eval.
  2. Performance: Both methods avoid looping through each patient to run apply(axis=1). The crosstab method, in particular, uses Pandas’ optimized C-backed operations, which are orders of magnitude faster than manual row-wise iteration for large datasets.
  3. Readability: The logic is more straightforward—we parse the data, reshape it, and count in clear, modular steps.

Notes:

  • This assumes each dict in the visits list has exactly one key (patient name) and one value (date), which matches your problem description. If dicts could have multiple keys, you’d adjust the patient extraction step to flatten all keys instead of taking the first one.

内容的提问来源于stack exchange,提问作者AJY

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 11:09:06