You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python解析含57000条数据的大型文本文件提取ETEXT NO.?

Python: Extract ETEXT NO. Values from Large Book Entry Text File

Problem Context

I've got a large text file (~57,000 book entries) formatted like this example snippet:

《植物生命面面观:特别参考英国植物区系》,作者Robert Lloyd Praeger,ETEXT NO.56900;《莫文斯托教区牧师》,作者Sabine Baring-Gould,ETEXT NO.56899【副标题:Robert Stephen Hawker,M.A.生平】;《Raamatun tutkisteluja IV》,作者Charles T. Russell,ETEXT NO.56898【副标题:Harmagedonin taistelu】【语言:芬兰语】……

I need to extract only the numerical values following "ETEXT NO." using Python. What's an efficient way to do this?


Solution: Use Regular Expressions for Precise Extraction

Regular expressions are perfect here because the "ETEXT NO." format is consistent. Here are two approaches tailored to different file sizes:

1. Efficient Line-by-Line Processing (Best for Large Files)

Since your file has 57k entries, loading the entire thing into memory might not be ideal. Instead, process it line by line to keep memory usage low:

import re

# Precompile the regex pattern for better performance
etext_pattern = re.compile(r'ETEXT NO\.(\d+)')

extracted_numbers = []

# Open and read the file line by line
with open('your_book_file.txt', 'r', encoding='utf-8') as input_file:
    for line in input_file:
        # Find all matching ETEXT numbers in the current line
        matches = etext_pattern.findall(line)
        extracted_numbers.extend(matches)

# Save the results to a new file (one number per line)
with open('extracted_etext_numbers.txt', 'w', encoding='utf-8') as output_file:
    for number in extracted_numbers:
        output_file.write(f"{number}\n")

2. One-Time Read (For Smaller Files)

If your file isn't overly large and fits comfortably in memory, you can read the entire content at once:

import re

etext_pattern = re.compile(r'ETEXT NO\.(\d+)')

# Read the entire file content
with open('your_book_file.txt', 'r', encoding='utf-8') as input_file:
    full_content = input_file.read()

# Extract all matching numbers in one go
extracted_numbers = etext_pattern.findall(full_content)

# Write results to output
with open('extracted_etext_numbers.txt', 'w', encoding='utf-8') as output_file:
    output_file.write('\n'.join(extracted_numbers))

Key Details & Adjustments

  • Regex Explanation: ETEXT NO\.(\d+) matches the literal string "ETEXT NO." (note the escaped . since it's a special regex character), then captures one or more digits (\d+) as the value we want.
  • Encoding: If your file uses a non-UTF-8 encoding (like GBK), update the encoding parameter in open() to match.
  • Flexible Matching: If there are variations in the format (e.g., lowercase, extra spaces), modify the regex to be more flexible:
    etext_pattern = re.compile(r'ETEXT\s+NO\.(\d+)', re.IGNORECASE)
    
    This will match variations like etext no.123, ETEXT NO.456, etc.

内容的提问来源于stack exchange,提问作者Azazel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:05:13