如何用Python解析含57000条数据的大型文本文件提取ETEXT NO.?
Problem Context
I've got a large text file (~57,000 book entries) formatted like this example snippet:
《植物生命面面观:特别参考英国植物区系》,作者Robert Lloyd Praeger,ETEXT NO.56900;《莫文斯托教区牧师》,作者Sabine Baring-Gould,ETEXT NO.56899【副标题:Robert Stephen Hawker,M.A.生平】;《Raamatun tutkisteluja IV》,作者Charles T. Russell,ETEXT NO.56898【副标题:Harmagedonin taistelu】【语言:芬兰语】……
I need to extract only the numerical values following "ETEXT NO." using Python. What's an efficient way to do this?
Solution: Use Regular Expressions for Precise Extraction
Regular expressions are perfect here because the "ETEXT NO." format is consistent. Here are two approaches tailored to different file sizes:
1. Efficient Line-by-Line Processing (Best for Large Files)
Since your file has 57k entries, loading the entire thing into memory might not be ideal. Instead, process it line by line to keep memory usage low:
import re # Precompile the regex pattern for better performance etext_pattern = re.compile(r'ETEXT NO\.(\d+)') extracted_numbers = [] # Open and read the file line by line with open('your_book_file.txt', 'r', encoding='utf-8') as input_file: for line in input_file: # Find all matching ETEXT numbers in the current line matches = etext_pattern.findall(line) extracted_numbers.extend(matches) # Save the results to a new file (one number per line) with open('extracted_etext_numbers.txt', 'w', encoding='utf-8') as output_file: for number in extracted_numbers: output_file.write(f"{number}\n")
2. One-Time Read (For Smaller Files)
If your file isn't overly large and fits comfortably in memory, you can read the entire content at once:
import re etext_pattern = re.compile(r'ETEXT NO\.(\d+)') # Read the entire file content with open('your_book_file.txt', 'r', encoding='utf-8') as input_file: full_content = input_file.read() # Extract all matching numbers in one go extracted_numbers = etext_pattern.findall(full_content) # Write results to output with open('extracted_etext_numbers.txt', 'w', encoding='utf-8') as output_file: output_file.write('\n'.join(extracted_numbers))
Key Details & Adjustments
- Regex Explanation:
ETEXT NO\.(\d+)matches the literal string "ETEXT NO." (note the escaped.since it's a special regex character), then captures one or more digits (\d+) as the value we want. - Encoding: If your file uses a non-UTF-8 encoding (like GBK), update the
encodingparameter inopen()to match. - Flexible Matching: If there are variations in the format (e.g., lowercase, extra spaces), modify the regex to be more flexible:
This will match variations likeetext_pattern = re.compile(r'ETEXT\s+NO\.(\d+)', re.IGNORECASE)etext no.123,ETEXT NO.456, etc.
内容的提问来源于stack exchange,提问作者Azazel

