You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从含大量行列的CSV文件中提取指定列数据?

Solution for Extracting Multiple Target Columns from Large CSV

Got it, let's fix this for you! The issue with your current code is that you're looping through every column in each row and printing only when you hit 'age'—so you're getting one value per line instead of grouping the target columns together for each row. Here are two solid approaches to extract multiple columns at once, perfect for your large CSV (5000+ columns is no problem with these methods since they process rows one at a time):

Method 1: Use csv.DictReader (Intuitive, Column Name-Based)

This method leverages the dictionary-like access to rows, making it easy to grab exactly the columns you want by name.

Example Code (Print or Write to New CSV)

import csv

# Define the columns you want to extract
target_columns = ['height', 'age']

### Option 1: Print the results to console
with open('your_file.csv', 'r') as file:
    reader = csv.DictReader(file, delimiter=',')
    # Print the header first
    print(','.join(target_columns))
    for row in reader:
        # Extract values for your target columns in order
        extracted_values = [row[col] for col in target_columns]
        # Print the full row of extracted data
        print(','.join(extracted_values))

### Option 2: Write results to a new CSV file
with open('your_file.csv', 'r') as infile, open('output.csv', 'w', newline='') as outfile:
    reader = csv.DictReader(infile, delimiter=',')
    writer = csv.writer(outfile)
    # Write the target columns as the header of the new file
    writer.writerow(target_columns)
    for row in reader:
        extracted_values = [row[col] for col in target_columns]
        writer.writerow(extracted_values)

Method 2: Use csv.reader (Index-Based, More Efficient for Massive Columns)

If you're dealing with extremely large numbers of columns, accessing values by index can be slightly faster. This method first maps your target column names to their positions in the header, then pulls values by those indices.

Example Code

import csv

target_columns = ['height', 'age']

with open('your_file.csv', 'r') as infile, open('output.csv', 'w', newline='') as outfile:
    reader = csv.reader(infile, delimiter=',')
    writer = csv.writer(outfile)
    
    # Read the header row to find column positions
    header = next(reader)
    # Get the indices of your target columns
    col_indices = [header.index(col) for col in target_columns]
    
    # Write the header to the output file
    writer.writerow(target_columns)
    
    # Loop through each row and extract values by index
    for row in reader:
        extracted_values = [row[idx] for idx in col_indices]
        writer.writerow(extracted_values)

Key Notes to Avoid Issues

  • Case Sensitivity: Make sure your target_columns match the exact case of the header in your CSV (e.g., if the header is Height instead of height, adjust your list accordingly). If you want to ignore case, you can normalize the header and target columns (e.g., convert both to lowercase).
  • Missing Values: If some rows don't have data for a target column, the code will return an empty string. You can add logic to replace these with a default value (like 'N/A') if needed.
  • Memory Efficiency: Both methods process rows one at a time, so you won't run into memory issues even with 1000+ rows and 5000+ columns.

内容的提问来源于stack exchange,提问作者whowho1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 15:57:59