Pandas:删除_Q后缀含低于阈值数字的单元格并保留合规前缀
Solution for Processing Pandas DataFrame Cells Based on Suffix Digits
Got it, let's work through this problem together. Here's a straightforward way to implement your requirements using Pandas:
Step 1: Define a Helper Function
First, we'll create a function that handles each individual cell. This function will:
- Split the cell value into the prefix (before
_Q) and the digit string (after_Q) - Check if all digits in the suffix are ≥ your threshold (7 in this case)
- Return the prefix if all digits meet the threshold, otherwise return a missing value (effectively "deleting" the cell content)
import pandas as pd def process_cell(cell, threshold=7): # Handle non-string cells or values without the _Q format cell_str = str(cell) if '_Q' not in cell_str: return cell # Split into prefix and the digit suffix prefix, digit_suffix = cell_str.split('_Q', 1) # Check every digit in the suffix all_digits_valid = all(int(d) >= threshold for d in digit_suffix) # Return prefix if valid, else return missing value return prefix if all_digits_valid else pd.NA
Step 2: Apply the Function to the Entire DataFrame
We'll use applymap() to run our helper function on every cell in the DataFrame—this is perfect because we need to process each value independently.
Example Usage
Let's test this with your sample data:
# Create your sample DataFrame df = pd.DataFrame({ 'col1': ['10_Q999999'], 'col2': ['111_Q4987666'], 'col3': ['110_Q277778'], 'col4': ['111_Q999999'] }) # Process all cells processed_df = df.applymap(process_cell) # Print the result print(processed_df)
Expected Output
col1 col2 col3 col4 0 10 <NA> <NA> 111
Optional: Replace Missing Values with Empty Strings
If you prefer empty strings instead of <NA> for the "deleted" cells, just add this line:
processed_df = processed_df.fillna('')
This will give you:
col1 col2 col3 col4 0 10 111
内容的提问来源于stack exchange,提问作者SummerEla
相关产品推荐
相关产品推荐

