Jupyter中处理DataFrame:排查Weight列<90值并导出CSV出错求助
Hey there! Let's work through this problem step by step. I've run into similar headaches with DataFrame row handling in Jupyter before, so I'll break down the correct approach, common pitfalls to avoid, and full working code examples.
First: Ditch Slow Row Traversal (Unless You Really Need It)
Pandas is optimized for vectorized operations—looping through rows with iterrows() should be your last resort (it's slow for large datasets). Let's start with the efficient way to filter your Weight column first.
Step 1: Filter Rows Where Weight < 90
Use boolean indexing to directly get all matching rows, including their indices and column values:
import pandas as pd # Replace with your actual DataFrame df = pd.DataFrame({ "Name": ["Alice", "Bob", "Charlie", "Diana"], "Weight": [87, 93, 82, 91], "Height": [165, 180, 170, 175] }) # Filter rows where Weight is less than 90 low_weight_records = df[df["Weight"] < 90] # Check the result (this will show indices and all column values) print(low_weight_records)
Step 2: Store Indices and Weight Values in Variables
If you need to isolate the indices and corresponding Weight values, you can extract them directly from the filtered DataFrame:
# Get list of indices for low-weight records low_weight_indices = low_weight_records.index.tolist() # Get list of matching Weight values low_weight_values = low_weight_records["Weight"].tolist() # Or store as a dictionary to keep index-value pairs linked low_weight_dict = dict(zip(low_weight_indices, low_weight_values))
Step 3: Export to CSV
Exporting the filtered records to CSV is straightforward. Use to_csv()—add index=False if you don't want the Pandas index to show up in the output file:
# Export the full filtered DataFrame (all columns) low_weight_records.to_csv("low_weight_full_records.csv", index=False) # Or export only the Index and Weight columns low_weight_records[["Weight"]].to_csv("low_weight_only.csv") # Index is included by default # If you want explicit "Index" column: pd.DataFrame({ "Index": low_weight_indices, "Weight": low_weight_values }).to_csv("low_weight_index_value.csv", index=False)
If You Must Traverse Rows (For Complex Logic)
If you have custom logic that requires looping through each row (e.g., more complex checks than just Weight < 90), use iterrows() correctly:
low_weight_list = [] for idx, row in df.iterrows(): # Access index with `idx`, column values with row["Column Name"] current_weight = row["Weight"] if current_weight < 90: # Store index and value as a dictionary low_weight_list.append({"Index": idx, "Weight": current_weight}) # Convert the list to a DataFrame for easy export low_weight_df = pd.DataFrame(low_weight_list) low_weight_df.to_csv("low_weight_iterrows.csv", index=False)
Common Causes of Unexpected Output
If your original code was failing, check these common issues:
- Non-numeric values in
Weightcolumn: IfWeighthas strings (e.g., "N/A"), the<comparison will throw an error. Fix withdf["Weight"] = pd.to_numeric(df["Weight"], errors="coerce")to convert to numbers (and turn invalid values intoNaN). - Incorrect row traversal: Using
df.rows(doesn't exist) or forgetting to unpackidxandrowwithiterrows(). - CSV export including extra indices: Forgetting
index=Falseleads to an extra unnamed column in your CSV. - Storing values incorrectly: If you stored indices and values as separate lists without linking them, your export might misalign data. Always use a DataFrame or dictionary to keep pairs together.
内容的提问来源于stack exchange,提问作者user9264558

