You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从含Dari语的文件中识别并移除英文单词?求代码优化方案

How to Remove English Words from Dari Files & Optimize Your Code

Hey there! Let's break down how to solve your two problems: identifying and stripping English words from Dari text files, plus optimizing the code you've already started working on.

1. Core Approach: Identifying & Removing English Words

Since Dari uses the Arabic script (with modified characters) and English relies on Latin letters, we can leverage this clear script distinction with regular expressions to target English words efficiently. Here's a step-by-step breakdown:

  • Key Insight: English words are made up of Latin letters (A-Z, a-z) and sometimes include apostrophes (like "don't" or "Mary's"). We can craft a regex pattern to match these while ignoring Dari's Arabic script entirely.
  • Implementation Steps:
    1. Read the Dari file with explicit UTF-8 encoding (critical for preserving non-Latin characters without garbling).
    2. Use regex to replace all matching English words with empty strings.
    3. Clean up extra whitespace left behind after removal (e.g., multiple spaces or line breaks).
    4. Save the processed content back to a new file.

Example Python Code

import re

def remove_english_from_dari(input_file, output_file):
    # Read the file with UTF-8 encoding (standard for Dari text)
    with open(input_file, 'r', encoding='utf-8') as f:
        content = f.read()
    
    # Regex pattern to match whole English words (including those with apostrophes)
    # \b ensures we don't match partial Latin characters embedded in Dari text
    english_word_pattern = r"\b[A-Za-z']+\b"
    # Replace all matches with empty string
    cleaned_content = re.sub(english_word_pattern, '', content)
    
    # Clean up redundant whitespace
    cleaned_content = re.sub(r'\s+', ' ', cleaned_content).strip()
    
    # Save the cleaned text
    with open(output_file, 'w', encoding='utf-8') as f:
        f.write(cleaned_content)

# Usage example
remove_english_from_dari("raw_dari_text.txt", "cleaned_dari_text.txt")

2. Optimizing Your Existing Code

If your current code works but feels slow, clunky, or incomplete, here are the most impactful tweaks to refine it:

  • Replace Manual Word Checks with Regex: If you were looping through each word and checking if it’s English (e.g., all(c.isalpha() for c in word)), regex is far more efficient—especially for large files. It processes the entire text in one pass instead of iterating over every word individually.
  • Enforce UTF-8 Encoding: Always specify encoding='utf-8' when reading/writing files. Skipping this can lead to garbled Dari text or errors, since default system encodings vary.
  • Add Error Handling: Wrap file operations in try-except blocks to catch missing files, permission issues, or encoding errors gracefully.
  • Handle Edge Cases: Update your logic to catch:
    • English words attached to Dari text (e.g., "سلامhello")—adjust the regex to r"[A-Za-z']+" (remove \b) to target these, but note this may match partial Latin characters if they exist.
    • English words with trailing punctuation (e.g., "hello!")—modify the regex to r"\b[A-Za-z']+\b|[A-Za-z']+[.,!?]" to include these cases.
  • Streamline Bulk Processing: If you’re handling multiple files, use a loop to process them in bulk instead of repeating code.

Example Optimized Code (For a Loop-Based Original)

Suppose your original code looked like this:

# Original (less efficient) code
with open("input.txt", "r") as f:
    lines = f.readlines()
cleaned_lines = []
for line in lines:
    words = line.split()
    cleaned_words = [word for word in words if not all(c.isalpha() for c in word)]
    cleaned_lines.append(' '.join(cleaned_words))
with open("output.txt", "w") as f:
    f.write('\n'.join(cleaned_lines))

Here’s the optimized version:

# Optimized code
import re

def optimize_dari_cleanup(input_file, output_file):
    try:
        with open(input_file, 'r', encoding='utf-8') as f:
            content = f.read()
        
        # Target English words with optional apostrophes and trailing punctuation
        english_pattern = r"\b[A-Za-z']+\b|[A-Za-z']+[.,!?]"
        cleaned_content = re.sub(english_pattern, '', content)
        cleaned_content = re.sub(r'\s+', ' ', cleaned_content).strip()
        
        with open(output_file, 'w', encoding='utf-8') as f:
            f.write(cleaned_content)
        print(f"Successfully cleaned {input_file}")
    except FileNotFoundError:
        print(f"Error: File {input_file} not found.")
    except Exception as e:
        print(f"An unexpected error occurred: {str(e)}")

# Usage example
optimize_dari_cleanup("raw_dari_text.txt", "cleaned_dari_text.txt")

This version runs faster, handles edge cases better, and is more robust than the original loop-based approach.

内容的提问来源于stack exchange,提问作者The Afghan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:17:23