使用正则表达式清理TSV文件特殊字符并嵌入Python代码实现
Clean Special Characters Before Generating TSV Subsets
Looks like you need to sanitize your TSV data before generating and exporting subsets—let's integrate that regex cleaning step into your existing code properly. I've adjusted the code to handle exactly what you asked for: removing special characters except periods, single spaces, tabs, slashes, and hyphens, plus collapsing consecutive spaces into a single one.
Here's the complete, updated code:
import pandas as pd import csv import re # Required for regex operations from itertools import chain, combinations # Load your source TSV file (updated to match your 'X.tsv' filename) df = pd.read_table('X.tsv') def clean_text(text): """Sanitize text by removing unwanted characters and normalizing spaces""" # Handle missing values to avoid errors if pd.isna(text): return text text_str = str(text) # Collapse consecutive spaces into a single space (preserve tab characters) cleaned = re.sub(r' +', ' ', text_str) # Remove any characters not in our allowed list: letters, numbers, space, tab, ., /, - cleaned = re.sub(r'[^a-zA-Z0-9 \t\./\-]', '', cleaned) # Optional: Trim leading/trailing spaces (remove this line if you want to keep them) return cleaned.strip() # Apply cleaning to all string columns in the DataFrame # For non-string columns (like numbers), they'll remain unchanged df = df.applymap(lambda x: clean_text(x) if isinstance(x, str) else x) def all_subsets(ss): return chain(*map(lambda x: combinations(ss, x), range(0, len(ss) + 1))) # Filter out columns you don't want included in subsets cols = [x for x in df.columns if x not in ['acm_classification', '...']] # Example: Generate and export each subset as a TSV file # Adjust the filename format to fit your needs for subset in all_subsets(cols): if subset: # Skip the empty subset (remove this check if you want it included) subset_columns = list(subset) subset_df = df[subset_columns] # Create a safe filename by replacing spaces/tabs with underscores safe_filename = f'subset_{"_".join(subset_columns)}.tsv' subset_df.to_csv(safe_filename, sep='\t', index=False, quoting=csv.QUOTE_NONE)
Key Notes:
- Regex Breakdown:
re.sub(r' +', ' ', text_str): Targets only consecutive spaces (not tabs) and replaces them with a single space, as requested.re.sub(r'[^a-zA-Z0-9 \t\./\-]', '', cleaned): Removes any character that isn't a letter, number, space, tab, period, slash, or hyphen. The hyphen is placed at the end of the bracket to avoid being interpreted as a range.
- Handling Non-String Data: The
applymapcheck ensures we only clean string columns—numeric or other data types stay untouched. - Exporting Subsets: Added code to export each subset to a TSV file with a descriptive name. The
quoting=csv.QUOTE_NONEensures tabs aren't wrapped in quotes, keeping the TSV format clean.
If you need to clean only specific columns instead of all string columns, replace the df.applymap(...) line with something like:
# Clean specific columns columns_to_clean = ['column1', 'column2'] for col in columns_to_clean: df[col] = df[col].apply(clean_text)
内容的提问来源于stack exchange,提问作者Faqahat Fareed
相关产品推荐
相关产品推荐

