如何通过函数批量生成带对应Exam列的多科目考试DataFrame?
Exam Column & Batch Process Subject DataFrames Let's break down your problem step by step and fix it, plus set up a scalable way to handle all 10 subjects at once.
Why You're Getting NaN in the Exam Column
Looking at your subject_df function, the issue is with the order of operations and how you initialize the DataFrame:
def subject_df(subject, subject_dl_df): df = pd.DataFrame() # Empty DataFrame with 0 rows df['Exam'] = subject # Assigns scalar to 0 rows, so this column stays empty df['Last Name'] = subject_dl_df['Last Name'] # Now df has rows from subject_dl_df, but Exam column remains NaN # ... more logic return df
When you first create an empty df and assign subject to the Exam column, there are no rows to hold that value. Later, when you add columns from subject_dl_df, pandas doesn't backfill the Exam column with the scalar you assigned earlier—it leaves it as NaN because the initial assignment was to an empty dataset.
Fixed Single Subject Function
Instead of starting with an empty DataFrame, we can build from the source data and add the Exam column cleanly. Here's the corrected function:
import pandas as pd def subject_df(subject, subject_dl_df): # Start with the relevant columns from the raw DataFrame df = subject_dl_df[['Student ID', 'Final Score']].copy() # Add the Exam column with the subject code (repeated for all rows) df['Exam'] = subject # Reorder columns to match your desired output structure df = df[['Exam', 'Student ID', 'Final Score']] # If you need additional columns like Last Name, include them in the initial copy # df = subject_dl_df[['Student ID', 'Final Score', 'Last Name']].copy() return df
Testing this with your sample data:
SXRX_df = subject_df('SXRX', SXRX_dl_df) print(SXRX_df)
Will output exactly what you expect:
Exam Student ID Final Score 0 SXRX 000000000 99 1 SXRX 000000001 100 2 SXRX 000000011 99
No more NaNs!
Batch Process All 10 Subjects
To handle all subjects efficiently, we can use a dictionary to map subject codes to their raw DataFrames, then loop through and combine everything into one final DataFrame.
Step 1: Organize your subject data in a dictionary
Map each subject code to its corresponding raw DataFrame:
# Dictionary: key = subject code, value = raw DataFrame subject_data = { 'SXRX': SXRX_dl_df, 'EXRC': EXRC_dl_df, # Add the other 8 subjects here (e.g., 'MATH': MATH_dl_df, 'PHYS': PHYS_dl_df) }
Step 2: Loop through and process all subjects
# Initialize a list to store processed DataFrames processed_dfs = [] for subject_code, raw_df in subject_data.items(): processed_df = subject_df(subject_code, raw_df) processed_dfs.append(processed_df) # Combine all processed DataFrames into one final dataset final_exam_df = pd.concat(processed_dfs, ignore_index=True)
Step 3: Verify the final output
print(final_exam_df)
You'll get a combined DataFrame with all subjects, each row correctly labeled with its Exam code:
Exam Student ID Final Score 0 SXRX 000000000 99 1 SXRX 000000001 100 2 SXRX 000000011 99 3 EXRC 111111110 100 4 EXRC 111111100 99 # ... rows from the remaining 8 subjects
Bonus: Auto-Extract Subject Codes from Filenames
If you're loading raw DataFrames from files (e.g., CSV/Excel), you can skip manual subject code entry by extracting codes from filenames. For example, if your files are named SXRX_dl_df.csv, EXRC_dl_df.csv:
import os import glob # Get all relevant files (adjust path/pattern to match your files) file_paths = glob.glob('*_dl_df.csv') subject_data = {} for file_path in file_paths: # Extract subject code from filename (e.g., 'SXRX_dl_df.csv' -> 'SXRX') subject_code = os.path.basename(file_path).split('_')[0] # Load the raw DataFrame raw_df = pd.read_csv(file_path) # Add to dictionary subject_data[subject_code] = raw_df # Process as before with the loop above
This makes scaling to 10+ subjects totally painless.
内容的提问来源于stack exchange,提问作者Christina Zhou

