如何用通配符表达式替换DataFrame中带M的数值为NA?
Hey there! Let's sort out this DataFrame issue you're facing. The problem with your current df[df == "*M"] approach is that the * doesn't act as a wildcard here—it's treated as a literal character, so it only looks for cells that are exactly "*M" (which none of your data has). Here are a few simple, effective ways to replace those 'M'-suffixed values with NA:
Step 1: First, let's recreate your input DataFrame
import pandas as pd # Your input data data = { "Depth": [0.001, 0.002, 0.003], "Density": ["0.3", "0.31M", "0.33"] } df = pd.DataFrame(data)
Method 1: Target rows with str.endswith()
This approach directly identifies cells ending with 'M' and replaces them with NA, then converts the column to numeric type (since your Density values might be stored as strings initially):
# Replace values ending with 'M' with NA df.loc[df['Density'].str.endswith('M'), 'Density'] = pd.NA # Convert Density to numeric type (optional but recommended for future calculations) df['Density'] = pd.to_numeric(df['Density'])
Method 2: One-liner with replace() and pd.to_numeric()
Use a regular expression to strip the trailing 'M', then convert the column to numeric—any invalid values (like empty strings left after stripping) will be coerced to NA automatically:
df['Density'] = pd.to_numeric(df['Density'].replace(r'M$', '', regex=True), errors='coerce')
The regex M$ matches the 'M' only when it's at the end of the string, so it won't affect any other characters.
Method 3: Alternative with str.replace()
If you prefer to work directly with the string column first, this does the same as Method 2 but uses str.replace():
df['Density'] = pd.to_numeric(df['Density'].str.replace('M$', '', regex=True), errors='coerce')
After running any of these methods, your DataFrame will match your desired output:
Depth Density 0 0.001 0.3 1 0.002 NaN 2 0.003 0.33
内容的提问来源于stack exchange,提问作者Snaxl

