如何在Python Pandas的read_csv中直接删除无表头的列?
Yes, absolutely! You can streamline this process by using pandas' usecols parameter directly in read_csv()—this lets you load only the columns you need upfront, instead of reading all columns and then dropping the unwanted ones. This is not only cleaner but also more efficient (especially for large CSV files, since you avoid loading unnecessary data into memory).
Here are two straightforward approaches to do this:
1. Load only the columns you want to keep
Instead of defining which columns to delete, list the ones you actually need. This is great if your desired column list is short:
import pandas as pd # Define the columns you want to retain desired_columns = ['station', 'date', 'observation', 'value'] # Read the CSV, set your custom headers, and load only the first 4 columns (matching your desired list) df = pd.read_csv('filename', names=desired_columns, usecols=range(4), header=None)
Note: usecols=range(4) targets the first 4 columns in your CSV. If your desired columns are in different positions, adjust the indices accordingly (e.g., usecols=[0,2,3,5] for specific positions).
If you want to reuse your full column name list, you can also reference the desired names directly:
import pandas as pd full_columns = ['station', 'date', 'observation', 'value', 'other_1', 'other_2', 'other_3', 'other_4'] desired_columns = ['station', 'date', 'observation', 'value'] df = pd.read_csv('filename', names=full_columns, usecols=desired_columns, header=None)
2. Exclude the columns you don't want
If you prefer to keep your existing del_columns_name list, use a lambda function with usecols to filter those out:
import pandas as pd columns_name = ['station', 'date', 'observation', 'value', 'other_1', 'other_2', 'other_3', 'other_4'] del_columns_name = ['other_1', 'other_2', 'other_3', 'other_4'] # Load all columns except the ones in del_columns_name df = pd.read_csv('filename', names=columns_name, usecols=lambda col: col not in del_columns_name, header=None)
A quick note on your original code: When using df.drop(), remember to either assign the result back to df (df = df.drop(...)) or use inplace=True (though inplace=True is generally not recommended for readability). But using usecols is definitely the better approach here.
内容的提问来源于stack exchange,提问作者matcha latte

