遍历工作目录子文件夹提取CSV指定数据的实现问题
data.csv in Subfolders & Build Target DataFrame Hey there! You’re already halfway there since you can list the folders and files—let’s lock in the nested logic to pull the data you need and assemble that DataFrame. Here’s a straightforward, robust approach:
Step 1: Import Required Tools
First, grab the libraries we’ll need for file navigation and DataFrame building:
import os import pandas as pd
Step 2: Prep a Storage List
We’ll use a list to collect each folder’s data (name, 2nd value, last value) before converting it to a DataFrame—this keeps things clean as we loop:
extracted_data = []
Step 3: Loop Through Subfolders & Extract Values
We’ll iterate through every subfolder in your current working directory, locate data.csv, pull the required values, and handle edge cases (like missing files or short CSVs) to avoid crashes:
# Get all subfolders in the current working directory subfolders = [folder for folder in os.listdir('.') if os.path.isdir(folder)] for folder_name in subfolders: # Build the full path to the data.csv file csv_file_path = os.path.join(folder_name, 'data.csv') # Skip if the file doesn't exist if not os.path.exists(csv_file_path): print(f"Warning: No data.csv found in {folder_name}—skipping.") continue try: # Read the CSV (use header=None if your file has no column headers) csv_content = pd.read_csv(csv_file_path, header=None) # Grab the 2nd value (Python uses 0-indexing, so index 1) second_value = csv_content.iloc[1, 0] if len(csv_content) >= 2 else "N/A" # Grab the last value last_value = csv_content.iloc[-1, 0] if len(csv_content) >= 1 else "N/A" # Add the folder's data to our list extracted_data.append({ 'Folder Name': folder_name, '2nd value': second_value, 'Last value': last_value }) except Exception as e: print(f"Error processing {csv_file_path}: {str(e)}")
Step 4: Assemble the Final DataFrame
Turn our collected list into the DataFrame format you want:
final_df = pd.DataFrame(extracted_data) # Reorder columns to match your requested layout final_df = final_df[['Folder Name', '2nd value', 'Last value']] # Check the result or save it print(final_df) # final_df.to_csv('folder_data_summary.csv', index=False)
Quick Adjustments for Your CSV Structure:
- If your
data.csvhas column headers, removeheader=Nonefrompd.read_csv()and adjust the index logic (e.g., usecsv_content['your_column_name'].iloc[1]instead ofcsv_content.iloc[1,0]). - If you need values from a specific column (not the first one), replace the
0iniloc[1,0]with the column index you need.
内容的提问来源于stack exchange,提问作者user1969717

