如何编写通用脚本检测仅包含单一重复数字值的数据列
Solution for Identifying Columns with Uniform Repeated Digits
Got it, let's solve this problem of identifying columns where every value is made up of the same digit repeated any number of times. Here's a straightforward approach with code examples you can adapt to your workflow:
Core Logic Breakdown
First, we need two key checks:
- For every value in the column, verify it’s composed entirely of a single repeated digit (e.g., "000", 555, "7" all pass; "121", "a99" fail).
- Ensure all non-null values in the column meet this criteria—even one invalid value disqualifies the column.
Example Python Function (Pandas-Friendly)
Since pandas is a common tool for data processing, here’s a reusable function that works with dataframes:
import pandas as pd def is_repeated_digit_col(column): def validate_value(val): # Handle nulls: adjust this if you want to ignore nulls (return True instead) if pd.isna(val): return False val_str = str(val) # Check if all characters are the same, and that character is a digit return len(val_str) > 0 and all(char == val_str[0] for char in val_str) and val_str[0].isdigit() # Check every value in the column and confirm all pass return column.apply(validate_value).all() # How to use it: sample_df = pd.DataFrame({ "valid_col_1": ["00", 111, "2222", 33333], "valid_col_2": [5, "7", 9], # Single digits count as valid (one repetition) "invalid_col": ["123", "111", "456"], "mixed_col": ["999", None, "888"] # Null makes this invalid (as per current logic) }) # Scan all columns and print valid ones for col_name in sample_df.columns: if is_repeated_digit_col(sample_df[col_name]): print(f"✅ Column '{col_name}' only contains repeated single digits.")
Key Customization Options
- Null Handling: The function currently rejects columns with any nulls. If you want to ignore nulls (treat them as non-disqualifying), modify the
validate_valuefunction to returnTruefor nulls, then usecolumn.dropna().apply(validate_value).all()instead of the current return line. - Exclude Single Digits: If you don’t want single-digit values (like "5" or 7) to count as valid, add
len(val_str) > 1to the return condition invalidate_value. - Regex Alternative: For a more concise value check, use regular expressions. This regex matches strings made of one or more identical digits:
import re def validate_value_regex(val): if pd.isna(val): return False return bool(re.fullmatch(r'(\d)\1*', str(val)))
内容的提问来源于stack exchange,提问作者Raghavendra Gupta
相关产品推荐
相关产品推荐

