CSV列值替换与重复前三段IP的第四段修改需求
Solution for CSV Column Replacement and IP Address Modification
Hey there! Let's break down how to solve your CSV manipulation task—you need to replace target values in a specific column and modify IP addresses based on the frequency of their first three octets. I'll show you two practical approaches to get this done.
Approach 1: Using Pandas (Fast & Convenient)
Pandas is ideal for this kind of tabular data processing, as it simplifies column-wise operations and value counting. Here's a step-by-step implementation:
import pandas as pd # 1. Load your CSV file (replace with your actual file path) df = pd.read_csv("your_input.csv") # -------------------------- # Step 1: Replace target values in a specified column # -------------------------- # Configure these variables to match your needs target_column = "your_target_column_name" old_value = "value_to_replace" new_value = "replacement_value" # Perform the replacement df[target_column] = df[target_column].replace(old_value, new_value) # -------------------------- # Step 2: Modify IP addresses based on prefix frequency # -------------------------- ip_column = "your_ip_column_name" # Extract the first three octets (prefix) of each IP df["ip_prefix"] = df[ip_column].apply(lambda ip: ".".join(ip.split(".")[:3])) # Count how many times each prefix appears prefix_counts = df["ip_prefix"].value_counts() # Define a function to adjust the IP def adjust_ip(ip): prefix = ".".join(ip.split(".")[:3]) # If prefix repeats, set fourth octet to 255; keep original otherwise return f"{prefix}.255" if prefix_counts[prefix] > 1 else ip # Apply the adjustment to the IP column df[ip_column] = df[ip_column].apply(adjust_ip) # Remove the temporary prefix column df = df.drop("ip_prefix", axis=1) # Save the processed CSV (replace with your desired output path) df.to_csv("your_output.csv", index=False)
Key Notes:
- Replace the placeholder names (like
your_input.csv,your_target_column_name) with your actual file and column details. - The
replace()method handles exact matches by default—if you need partial matches, you can usestr.replace()instead. - The
value_counts()method quickly tallies how often each IP prefix appears, making it easy to identify duplicates.
Approach 2: Using Python's Built-in CSV Module (No Dependencies)
If you can't install pandas or prefer using standard libraries, here's how to do it with the native csv module:
import csv # Configure your parameters input_file = "your_input.csv" output_file = "your_output.csv" target_col_index = 1 # Index of the column to replace values (0-based) old_value = "value_to_replace" new_value = "replacement_value" ip_col_index = 2 # Index of the IP column (0-based) # First pass: Read all rows, perform initial replacement, and count IP prefixes rows = [] ip_prefix_counts = {} with open(input_file, "r", newline="") as infile: reader = csv.reader(infile) header = next(reader) rows.append(header) for row in reader: # Replace target value in the specified column if row[target_col_index] == old_value: row[target_col_index] = new_value # Extract and count IP prefixes ip_address = row[ip_col_index] prefix = ".".join(ip_address.split(".")[:3]) ip_prefix_counts[prefix] = ip_prefix_counts.get(prefix, 0) + 1 rows.append(row) # Second pass: Modify IP addresses based on prefix counts with open(output_file, "w", newline="") as outfile: writer = csv.writer(outfile) writer.writerow(rows[0]) for row in rows[1:]: ip_address = row[ip_col_index] prefix = ".".join(ip_address.split(".")[:3]) # Update IP if prefix repeats if ip_prefix_counts[prefix] > 1: row[ip_col_index] = f"{prefix}.255" writer.writerow(row)
Key Notes:
- We use two passes: first to count prefix frequencies, then to modify the IPs (since we need to know the total count before making changes).
- Adjust the column indices (
target_col_index,ip_col_index) to match your CSV's structure (remember, indices start at 0).
Let me know if you need help adapting this to your specific CSV structure or edge cases!
内容的提问来源于stack exchange,提问作者Mohamed Rached Ben Cheikh
相关产品推荐
相关产品推荐

