Python:如何从for循环获取的DataFrame结果中提取最高价行?
Hey there! Let's break down how to grab the row with the highest price for each product in your loop. Your current code is already fetching all rows for a given product—we just need to add a few steps to filter for the maximum price entry.
Method 1: Using idxmax() with Your Existing Index Setup
Since you've set ProductName as the index, here's how to adjust your loop to get the highest price row:
import pandas as pd import numpy as np df = pd.read_csv('/Users/caleb/ic/US_FINAL.csv', names=['ProductName', 'Year', 'Production', 'Price']) df.set_index("ProductName", inplace=True) products=['FortuneCookie'] for product in products: product_data = df.loc[product] # Handle cases where the product has only one row (returns a Series) if isinstance(product_data, pd.Series): print(f"Highest price row for {product}:\n{product_data}") else: # Get the index of the row with the highest Price max_price_idx = product_data['Price'].idxmax() # Extract that row max_price_row = product_data.loc[max_price_idx] print(f"Highest price row for {product}:\n{max_price_row}")
How this works:
product_data['Price'].idxmax()returns the index of the row where the Price value is the highest.- We use that index with
.loc[]to pull the full row data. - The type check handles edge cases where a product only has one entry (so
df.loc[product]returns a Series instead of a DataFrame).
Method 2: Boolean Filter (No Index Modification)
If you prefer not to alter the original DataFrame's index, this approach uses boolean filtering to isolate product-specific data first:
import pandas as pd df = pd.read_csv('/Users/caleb/ic/US_FINAL.csv', names=['ProductName', 'Year', 'Production', 'Price']) products=['FortuneCookie'] for product in products: # Filter rows for the current product product_df = df[df['ProductName'] == product] # Skip if no data exists for the product if product_df.empty: print(f"No records found for {product}") continue # Get the highest price row max_price_row = product_df.loc[product_df['Price'].idxmax()] print(f"Highest price row for {product}:\n{max_price_row}")
Handling Ties (Multiple Rows with Same Highest Price)
If multiple rows have the same maximum price and you want to retrieve all of them, replace the idxmax() step with a boolean filter for the maximum price value:
import pandas as pd df = pd.read_csv('/Users/caleb/ic/US_FINAL.csv', names=['ProductName', 'Year', 'Production', 'Price']) products=['FortuneCookie'] for product in products: product_df = df[df['ProductName'] == product] if product_df.empty: print(f"No records found for {product}") continue # Get the maximum price value highest_price = product_df['Price'].max() # Filter all rows where Price equals the maximum max_price_rows = product_df[product_df['Price'] == highest_price] print(f"All rows with highest price for {product}:\n{max_price_rows}")
This will return every row where the product's price matches the highest value in its dataset.
内容的提问来源于stack exchange,提问作者BPDESILVA

