Pandas代码未实现过滤仍打印全部列头,求技术解决建议
Hey there! First off, welcome to posting here—no worries about any small oversights, we’ve all been there 😊 Let’s break down what’s happening and fix your issue.
为什么会打印所有列头?
Looking at your code, this line is the main culprit:
print(list(data))
When you call list(data) on a Pandas DataFrame, it returns a list of all the column names by default. So that’s exactly why you’re seeing every column header printed out. If that’s not what you want, just comment or delete this line.
如何实现“过滤并打印目标数据”?
It looks like your code cuts off at the sort_values part, and you haven’t added the actual filtering logic yet. Let’s cover two common scenarios you might be aiming for:
1. 只保留特定列(过滤列)
If you want to print only certain columns (not all), pick the column names you need and slice the DataFrame:
import pandas as pd # Read the CSV (already returns a DataFrame, no need to re-wrap it) data = pd.read_csv("/Users/andrewschaper/Desktop/EQR_Data/EQR_Transactions_1.csv", low_memory=False) print("Total rows: {0}".format(len(data))) # Define your target columns target_columns = ['Filing_Quarter', 'product_name', 'time_zone'] # Add your actual target columns here filtered_data = data[target_columns] # Sort if needed, then print the filtered data filtered_data_sorted = filtered_data.sort_values(by=['Filing_Quarter','product_name','time_zone']) print(filtered_data_sorted)
2. 筛选满足条件的行(过滤行)
If you want to print only rows that meet a specific condition (e.g., rows where Filing_Quarter is "2024Q1"), use boolean indexing:
import pandas as pd data = pd.read_csv("/Users/andrewschaper/Desktop/EQR_Data/EQR_Transactions_1.csv", low_memory=False) print("Total rows: {0}".format(len(data))) # Define your filter condition filter_condition = data['Filing_Quarter'] == '2024Q1' # Replace with your actual condition filtered_data = data[filter_condition] # Sort and print filtered_data_sorted = filtered_data.sort_values(by=['Filing_Quarter','product_name','time_zone']) print(filtered_data_sorted)
额外小提示
You don’t need this line:
df = pd.DataFrame(data)
Because pd.read_csv() already returns a Pandas DataFrame—re-wrapping it is redundant and doesn’t do anything useful.
If you share more details about exactly what "target data" you’re trying to filter (e.g., specific column values, column names), I can refine this even more!
内容的提问来源于stack exchange,提问作者SOCO

