如何从含Dari语的文件中识别并移除英文单词?求代码优化方案
Hey there! Let's break down how to solve your two problems: identifying and stripping English words from Dari text files, plus optimizing the code you've already started working on.
1. Core Approach: Identifying & Removing English Words
Since Dari uses the Arabic script (with modified characters) and English relies on Latin letters, we can leverage this clear script distinction with regular expressions to target English words efficiently. Here's a step-by-step breakdown:
- Key Insight: English words are made up of Latin letters (A-Z, a-z) and sometimes include apostrophes (like "don't" or "Mary's"). We can craft a regex pattern to match these while ignoring Dari's Arabic script entirely.
- Implementation Steps:
- Read the Dari file with explicit UTF-8 encoding (critical for preserving non-Latin characters without garbling).
- Use regex to replace all matching English words with empty strings.
- Clean up extra whitespace left behind after removal (e.g., multiple spaces or line breaks).
- Save the processed content back to a new file.
Example Python Code
import re def remove_english_from_dari(input_file, output_file): # Read the file with UTF-8 encoding (standard for Dari text) with open(input_file, 'r', encoding='utf-8') as f: content = f.read() # Regex pattern to match whole English words (including those with apostrophes) # \b ensures we don't match partial Latin characters embedded in Dari text english_word_pattern = r"\b[A-Za-z']+\b" # Replace all matches with empty string cleaned_content = re.sub(english_word_pattern, '', content) # Clean up redundant whitespace cleaned_content = re.sub(r'\s+', ' ', cleaned_content).strip() # Save the cleaned text with open(output_file, 'w', encoding='utf-8') as f: f.write(cleaned_content) # Usage example remove_english_from_dari("raw_dari_text.txt", "cleaned_dari_text.txt")
2. Optimizing Your Existing Code
If your current code works but feels slow, clunky, or incomplete, here are the most impactful tweaks to refine it:
- Replace Manual Word Checks with Regex: If you were looping through each word and checking if it’s English (e.g.,
all(c.isalpha() for c in word)), regex is far more efficient—especially for large files. It processes the entire text in one pass instead of iterating over every word individually. - Enforce UTF-8 Encoding: Always specify
encoding='utf-8'when reading/writing files. Skipping this can lead to garbled Dari text or errors, since default system encodings vary. - Add Error Handling: Wrap file operations in
try-exceptblocks to catch missing files, permission issues, or encoding errors gracefully. - Handle Edge Cases: Update your logic to catch:
- English words attached to Dari text (e.g., "سلامhello")—adjust the regex to
r"[A-Za-z']+"(remove\b) to target these, but note this may match partial Latin characters if they exist. - English words with trailing punctuation (e.g., "hello!")—modify the regex to
r"\b[A-Za-z']+\b|[A-Za-z']+[.,!?]"to include these cases.
- English words attached to Dari text (e.g., "سلامhello")—adjust the regex to
- Streamline Bulk Processing: If you’re handling multiple files, use a loop to process them in bulk instead of repeating code.
Example Optimized Code (For a Loop-Based Original)
Suppose your original code looked like this:
# Original (less efficient) code with open("input.txt", "r") as f: lines = f.readlines() cleaned_lines = [] for line in lines: words = line.split() cleaned_words = [word for word in words if not all(c.isalpha() for c in word)] cleaned_lines.append(' '.join(cleaned_words)) with open("output.txt", "w") as f: f.write('\n'.join(cleaned_lines))
Here’s the optimized version:
# Optimized code import re def optimize_dari_cleanup(input_file, output_file): try: with open(input_file, 'r', encoding='utf-8') as f: content = f.read() # Target English words with optional apostrophes and trailing punctuation english_pattern = r"\b[A-Za-z']+\b|[A-Za-z']+[.,!?]" cleaned_content = re.sub(english_pattern, '', content) cleaned_content = re.sub(r'\s+', ' ', cleaned_content).strip() with open(output_file, 'w', encoding='utf-8') as f: f.write(cleaned_content) print(f"Successfully cleaned {input_file}") except FileNotFoundError: print(f"Error: File {input_file} not found.") except Exception as e: print(f"An unexpected error occurred: {str(e)}") # Usage example optimize_dari_cleanup("raw_dari_text.txt", "cleaned_dari_text.txt")
This version runs faster, handles edge cases better, and is more robust than the original loop-based approach.
内容的提问来源于stack exchange,提问作者The Afghan

