求助:基于Pandas检测Excel中多种函数循环数量的Python代码
Hey there! Let's tackle those remaining loop detection tasks step by step. First, let's clarify the exact definitions of each loop type based on your description (I'll make assumptions where needed—feel free to adjust if my interpretation is off):
Clarified Loop Definitions
- (x>x): As you noted, this detects pairs where elements in the x column are equal (i.e.,
x_i = x_jfori ≠ j) - (x>y, y>x): Cross-row mutual inequalities: exists rows
iandj(i≠j) wherex_i > y_jANDy_i > x_j - (x>y,x>y1): Single x value greater than multiple y values: exists a row
iwherex_i > y_jANDx_i > y_kfor distinct rowsj ≠ k - (x>y,y>x1): Chained inequalities: exists rows
i,j,k(all distinct) wherex_i > y_jANDy_j > x_k
Prerequisites
We'll use Python with pandas to read and process your Excel data (install it first if you haven't: pip install pandas openpyxl).
Step 1: Load the Data
First, read your Excel file into a DataFrame:
import pandas as pd # Replace 'your_file.xlsx' with your actual file path df = pd.read_excel('your_file.xlsx', usecols=['x', 'y', 'z'])
Step 2: Detect (x>y, y>x) Loops
To find pairs of rows where the mutual x-y inequalities hold:
# Extract values for faster comparison x_vals = df['x'].values y_vals = df['y'].values # Initialize a set to store unique pairs (avoids duplicates like (i,j) and (j,i)) mutual_pairs = set() for i in range(len(df)): for j in range(len(df)): if i != j and x_vals[i] > y_vals[j] and y_vals[i] > x_vals[j]: # Store pairs as sorted tuples to eliminate duplicate entries pair = tuple(sorted((i, j))) mutual_pairs.add(pair) print(f"Number of (x>y, y>x) loops: {len(mutual_pairs)}") print("Affected row pairs:", mutual_pairs)
Step 3: Detect (x>y,x>y1) Loops
Find rows where the x value is greater than at least two distinct y values:
# For each x value, count how many y values it's greater than x_greater_counts = df['x'].apply(lambda x_val: sum(y_val < x_val for y_val in df['y'])) # Filter rows where the count is 2 or more x_multiple_y = df[x_greater_counts >= 2] print(f"Number of (x>y,x>y1) loops: {len(x_multiple_y)}") print("Rows with x greater than multiple y values:\n", x_multiple_y)
Step 4: Detect (x>y,y>x1) Loops
Find chained inequalities where an x from row i is greater than y from row j, which in turn is greater than x from row k:
# Create a list to store valid chained triplets chained_loops = [] # Iterate all possible triplet combinations for i in range(len(df)): for j in range(len(df)): if i == j: continue if df['x'].iloc[i] > df['y'].iloc[j]: # Find k values where y_j is greater than x_k for k in range(len(df)): if k != i and k != j and df['y'].iloc[j] > df['x'].iloc[k]: chained_loops.append((i, j, k)) # Remove duplicate triplets (if order doesn't matter for your use case) unique_chains = list(set(chained_loops)) print(f"Number of (x>y,y>x1) loops: {len(unique_chains)}") print("Chained row triplets:", unique_chains)
Notes
- If your definitions of the loops differ (e.g., if (x>y,x>y1) refers to
x_i > y_iANDx_i > y_jfori≠j), feel free to tweak the logic! The code is modular and easy to adjust. - For large datasets, nested loops might be slow—we can optimize using vectorized operations with pandas/numpy if you need better performance.
内容的提问来源于stack exchange,提问作者Amit

