You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

跨多列统计医院出院记录中20个特定操作代码的频率

Cross-Column Frequency Count for Target Procedure Codes

Hey there! Let's break down how to get that frequency list for your 20 specific procedure codes across all those columns in your hospital discharge dataset. I'll use pandas (the go-to tool for tabular data in Python) since it makes this task super straightforward.

Step 1: Set Up Your Environment & Load Data

First, make sure you have pandas installed (if not, run pip install pandas in your terminal). Then load your dataset:

import pandas as pd

# Replace with your actual file path (could be CSV, Excel, etc.)
df = pd.read_csv("discharge_records.csv")
# If using Excel: df = pd.read_excel("discharge_records.xlsx")

Step 2: Define Your Target Codes & Relevant Columns

List out the 20 procedure codes you care about, and specify which columns to check (main code + all the "其他代码" columns):

# Replace these with your actual 20 target codes
target_codes = ["PROC123", "PROC456", ..., "PROC789"]

# Grab all columns that contain procedure codes
# Adjust the string matches to fit your actual column names
code_columns = [col for col in df.columns if col in ["主代码"] or col.startswith("其他代码")]
# Alternatively, list them explicitly if you know exact names:
# code_columns = ["主代码", "其他代码1", "其他代码2", ..., "其他代码24"]

Step 3: Calculate Cross-Column Frequencies

We'll stack all the code columns into a single list of codes, then count how often each target code appears:

# Combine all code columns into one long series (ignores empty values automatically)
all_codes = df[code_columns].stack()

# Filter to only your target codes and count occurrences
frequency_counts = all_codes[all_codes.isin(target_codes)].value_counts().reset_index()
frequency_counts.columns = ["操作代码", "出现频次"]

# Optional: Include target codes that didn't appear (with 0 count)
frequency_counts = frequency_counts.merge(
    pd.DataFrame({"操作代码": target_codes}),
    on="操作代码",
    how="right"
).fillna(0).sort_values("出现频次", ascending=False)

Step 4: View or Save the Results

You can print the results directly or save them to a file for later use:

# Print the frequency list
print(frequency_counts)

# Save to a CSV file (easy to share or analyze further)
frequency_counts.to_csv("procedure_code_frequencies.csv", index=False, encoding="utf-8-sig")

Quick Tips to Avoid Issues

  • Data Type Check: Ensure your code columns are stored as strings (not numbers) to avoid mismatches. Use df[code_columns] = df[code_columns].astype(str) if needed.
  • Missing Values: Empty cells in code columns are ignored by stack(), so they won't skew your counts.
  • Large Datasets: Pandas handles big datasets efficiently, but if your file is massive, you can process it in chunks with pd.read_csv(chunksize=10000) to save memory.

内容的提问来源于stack exchange,提问作者W. Bensken

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:54:14