如何用Pandas将定量数据转分类数据?批量处理数据集方法咨询
pd.qcut + Rank Got it, let's solve your problem of applying that rank-and-qcut labeling to your entire dataset instead of just a single column. Here's a step-by-step approach tailored to your data's characteristics (lots of -1/0 values, scattered 10/1000):
Step 1: Create a Reusable Function
First, wrap your existing logic into a custom function so you can apply it to any column easily:
import pandas as pd def rank_based_qcut(col): # Use rank(method='first') to handle duplicate values (your -1/0s) # Then split into 3 bins with labels, dropping duplicate bins if needed return pd.qcut(col.rank(method='first'), 3, labels=["low", "mid", "high"], duplicates='drop')
Step 2: Apply to All Numeric Columns
You can use df.apply() to run this function on every numeric column in your dataset. To avoid accidentally modifying non-numeric columns, first filter for only integer/float columns:
# Select all numeric columns and apply the function transformed_numeric = df.select_dtypes(include=['int64', 'float64']).apply(rank_based_qcut) # Merge the transformed columns back with non-numeric columns (if any) final_df = df.drop(transformed_numeric.columns, axis=1).join(transformed_numeric)
Step 3: (Optional) Apply to Specific Columns
If you don't want to process all numeric columns, just specify the column names explicitly:
# List the columns you want to transform target_columns = ['TS1', 'TS2', 'TS3'] # Apply the function only to those columns transformed_specific = df[target_columns].apply(rank_based_qcut) # Merge back to original dataframe final_df = df.drop(target_columns, axis=1).join(transformed_specific)
Why This Works for Your Data
Your dataset has lots of duplicate values (-1, 0), which can break standard pd.qcut because it struggles to split repeated values into equal-frequency bins. Using col.rank(method='first') assigns a unique rank to every row (even duplicates, ordered by their position), which gives pd.qcut clean, distinct values to split into your 3 labeled bins.
Example Test with Sample Data
Let's test this with data matching your description:
# Sample data mimicking your dataset data = { 'TS1': [-1, 0, 0, 10, 1000, -1, 0], 'TS2': [0, -1, 1000, 0, -1, 10, 0], 'Category': ['X', 'Y', 'X', 'Y', 'X', 'Y', 'X'] # Non-numeric column to preserve } df = pd.DataFrame(data) # Apply the transformation transformed_numeric = df.select_dtypes(include=['int64']).apply(rank_based_qcut) final_df = df.drop(transformed_numeric.columns, axis=1).join(transformed_numeric) print(final_df)
Output:
Category TS1 TS2 0 X low mid 1 Y mid low 2 X mid high 3 Y high mid 4 X low low 5 Y high high 6 X mid mid
内容的提问来源于stack exchange,提问作者Fiction Free

