跨多列统计医院出院记录中20个特定操作代码的频率
Hey there! Let's break down how to get that frequency list for your 20 specific procedure codes across all those columns in your hospital discharge dataset. I'll use pandas (the go-to tool for tabular data in Python) since it makes this task super straightforward.
Step 1: Set Up Your Environment & Load Data
First, make sure you have pandas installed (if not, run pip install pandas in your terminal). Then load your dataset:
import pandas as pd # Replace with your actual file path (could be CSV, Excel, etc.) df = pd.read_csv("discharge_records.csv") # If using Excel: df = pd.read_excel("discharge_records.xlsx")
Step 2: Define Your Target Codes & Relevant Columns
List out the 20 procedure codes you care about, and specify which columns to check (main code + all the "其他代码" columns):
# Replace these with your actual 20 target codes target_codes = ["PROC123", "PROC456", ..., "PROC789"] # Grab all columns that contain procedure codes # Adjust the string matches to fit your actual column names code_columns = [col for col in df.columns if col in ["主代码"] or col.startswith("其他代码")] # Alternatively, list them explicitly if you know exact names: # code_columns = ["主代码", "其他代码1", "其他代码2", ..., "其他代码24"]
Step 3: Calculate Cross-Column Frequencies
We'll stack all the code columns into a single list of codes, then count how often each target code appears:
# Combine all code columns into one long series (ignores empty values automatically) all_codes = df[code_columns].stack() # Filter to only your target codes and count occurrences frequency_counts = all_codes[all_codes.isin(target_codes)].value_counts().reset_index() frequency_counts.columns = ["操作代码", "出现频次"] # Optional: Include target codes that didn't appear (with 0 count) frequency_counts = frequency_counts.merge( pd.DataFrame({"操作代码": target_codes}), on="操作代码", how="right" ).fillna(0).sort_values("出现频次", ascending=False)
Step 4: View or Save the Results
You can print the results directly or save them to a file for later use:
# Print the frequency list print(frequency_counts) # Save to a CSV file (easy to share or analyze further) frequency_counts.to_csv("procedure_code_frequencies.csv", index=False, encoding="utf-8-sig")
Quick Tips to Avoid Issues
- Data Type Check: Ensure your code columns are stored as strings (not numbers) to avoid mismatches. Use
df[code_columns] = df[code_columns].astype(str)if needed. - Missing Values: Empty cells in code columns are ignored by
stack(), so they won't skew your counts. - Large Datasets: Pandas handles big datasets efficiently, but if your file is massive, you can process it in chunks with
pd.read_csv(chunksize=10000)to save memory.
内容的提问来源于stack exchange,提问作者W. Bensken

