使用re.match()清洗Pandas DataFrame时全部置为NaN的问题求助
Hey there, let's break down why your function is turning every value to NaN and fix it step by step!
1. 核心问题:正则表达式完全不符合需求
Your regex ^\d{1}(\.)\d{2}$ is designed to match numbers with one integer digit, a decimal point, and two decimal places (like 1.23 or 5.67), but your goal is to keep only values that are exactly 1 or 2. This mismatch means none of your data entries are passing the regex check, so everything gets set to NaN.
The correct regex to match exactly "1" or "2" is ^[12]$:
^anchors the match to the start of the string[12]matches either the character "1" or "2"$anchors the match to the end of the string (ensures there's no extra characters before/after)
2. 优化Pandas操作方式
Double-looping through every cell in a DataFrame is inefficient and can cause unexpected behavior (like SettingWithCopyWarning). Let's use Pandas' vectorized operations instead, which are faster and more reliable.
Fixed Code
import pandas as pd import re import numpy as np data = [['Bob',10,1],['Bob',2,2],['Clarke',13,1]] my_df = pd.DataFrame(data,columns=['Name','Age','Sex']) def my_fun(df): # Iterate through each column and apply the check for col in df.columns: df[col] = df[col].apply( lambda x: x if re.match(r'^[12]$', str(x)) else np.nan ) return df # Test the function cleaned_df = my_fun(my_df) print(cleaned_df)
Expected Output
Name Age Sex 0 NaN NaN 1.0 1 NaN 2.0 2.0 2 NaN NaN 1.0
3. Even Better: Skip Regex Entirely
Since you're only checking for exact values 1 or 2, you don't need regex at all. Using Pandas' isin() method is faster and simpler:
def my_fun_v2(df): for col in df.columns: # Keep values that are 1 or 2, set others to NaN df[col] = df[col].where(df[col].isin([1, 2]), np.nan) return df
This achieves the same result without converting values to strings, making it more efficient for numeric columns.
内容的提问来源于stack exchange,提问作者Robbie

