基于条件循环向Pandas数据框插入新行及CSV修改需求
Here's a practical, pandas-based solution to modify your CSV files by inserting a 2 between the specified patterns in the 'Markers' column. This builds on the shifted-column pattern detection idea you mentioned, while ensuring clean row insertion without index issues.
Step 1: Flag Target Patterns
First, we'll create helper columns to identify exactly where your target patterns occur by comparing each marker to the next one:
import pandas as pd def add_pattern_flags(df): # Shift the Markers column to compare consecutive values df['Next_Marker'] = df['Markers'].shift(-1) # Flag each pattern you need to target df['TwoThrees'] = (df['Markers'] == 3) & (df['Next_Marker'] == 3) df['TwoFours'] = (df['Markers'] == 4) & (df['Next_Marker'] == 4) df['FiveThree'] = (df['Markers'] == 5) & (df['Next_Marker'] == 3) df['FourThree'] = (df['Markers'] == 4) & (df['Next_Marker'] == 3) # Combine all flags into a single "needs insertion" column df['Insert_After'] = df[['TwoThrees', 'TwoFours', 'FiveThree', 'FourThree']].any(axis=1) return df
Step 2: Insert the 2 Between Matched Patterns
Next, we'll iterate through the DataFrame and build a new list of rows, inserting the 2 wherever our flag is true. This avoids index shifting issues that come with direct row insertion:
def insert_two_in_markers(df): df_with_flags = add_pattern_flags(df) new_rows = [] for idx, row in df_with_flags.iterrows(): # Add the current row to our output list new_rows.append(row.to_dict()) # Insert a 2 row if the pattern matches and we're not at the last row if row['Insert_After'] and idx < len(df_with_flags) - 1: # Create a new row with Markers=2 (copy other columns from current row) insert_row = row.to_dict() insert_row['Markers'] = 2 # Clean up helper columns from the inserted row for col in ['Next_Marker', 'TwoThrees', 'TwoFours', 'FiveThree', 'FourThree', 'Insert_After']: insert_row.pop(col, None) new_rows.append(insert_row) # Convert back to DataFrame and remove helper columns from original rows modified_df = pd.DataFrame(new_rows).drop(columns=['Next_Marker', 'TwoThrees', 'TwoFours', 'FiveThree', 'FourThree', 'Insert_After'], errors='ignore') return modified_df
Step 3: Process All Your CSV Files
To apply this to every CSV in your directory and collect results in PVTdfs:
import glob # Replace with your CSV directory path csv_directory = "path/to/your/csv/files/*.csv" PVTdfs = [] for file in glob.glob(csv_directory): original_df = pd.read_csv(file) modified_df = insert_two_in_markers(original_df) PVTdfs.append(modified_df) # Optional: Save the modified CSV with a new name modified_df.to_csv(f"modified_{file.split('/')[-1]}", index=False)
Quick Notes:
- Other Columns: The code copies values from the preceding row for all columns except 'Markers' in the inserted row. If you need NaN or specific values here, adjust the
insert_rowcreation step. - Edge Cases: This handles consecutive patterns and the last row correctly (no insertion after it).
- Cleanup: Helper pattern flags are removed from the final output, so your CSV stays tidy.
Let me know if you need to tweak this for your specific CSV structure!
内容的提问来源于stack exchange,提问作者Heather

