如何用Python从含大量行列的CSV文件中提取指定列数据?
Got it, let's fix this for you! The issue with your current code is that you're looping through every column in each row and printing only when you hit 'age'—so you're getting one value per line instead of grouping the target columns together for each row. Here are two solid approaches to extract multiple columns at once, perfect for your large CSV (5000+ columns is no problem with these methods since they process rows one at a time):
Method 1: Use csv.DictReader (Intuitive, Column Name-Based)
This method leverages the dictionary-like access to rows, making it easy to grab exactly the columns you want by name.
Example Code (Print or Write to New CSV)
import csv # Define the columns you want to extract target_columns = ['height', 'age'] ### Option 1: Print the results to console with open('your_file.csv', 'r') as file: reader = csv.DictReader(file, delimiter=',') # Print the header first print(','.join(target_columns)) for row in reader: # Extract values for your target columns in order extracted_values = [row[col] for col in target_columns] # Print the full row of extracted data print(','.join(extracted_values)) ### Option 2: Write results to a new CSV file with open('your_file.csv', 'r') as infile, open('output.csv', 'w', newline='') as outfile: reader = csv.DictReader(infile, delimiter=',') writer = csv.writer(outfile) # Write the target columns as the header of the new file writer.writerow(target_columns) for row in reader: extracted_values = [row[col] for col in target_columns] writer.writerow(extracted_values)
Method 2: Use csv.reader (Index-Based, More Efficient for Massive Columns)
If you're dealing with extremely large numbers of columns, accessing values by index can be slightly faster. This method first maps your target column names to their positions in the header, then pulls values by those indices.
Example Code
import csv target_columns = ['height', 'age'] with open('your_file.csv', 'r') as infile, open('output.csv', 'w', newline='') as outfile: reader = csv.reader(infile, delimiter=',') writer = csv.writer(outfile) # Read the header row to find column positions header = next(reader) # Get the indices of your target columns col_indices = [header.index(col) for col in target_columns] # Write the header to the output file writer.writerow(target_columns) # Loop through each row and extract values by index for row in reader: extracted_values = [row[idx] for idx in col_indices] writer.writerow(extracted_values)
Key Notes to Avoid Issues
- Case Sensitivity: Make sure your
target_columnsmatch the exact case of the header in your CSV (e.g., if the header isHeightinstead ofheight, adjust your list accordingly). If you want to ignore case, you can normalize the header and target columns (e.g., convert both to lowercase). - Missing Values: If some rows don't have data for a target column, the code will return an empty string. You can add logic to replace these with a default value (like
'N/A') if needed. - Memory Efficiency: Both methods process rows one at a time, so you won't run into memory issues even with 1000+ rows and 5000+ columns.
内容的提问来源于stack exchange,提问作者whowho1

