Python提取含1-6位数字(首尾可带字符)的行及ETEXT NO相关信息
Let's break down your problems one by one, starting with why your regex didn't work and then fixing both requirements.
Why Your Regex Failed
Your pattern ^[a-zA-Z]\d{1,6}$ is way too restrictive:
^and$anchor the match to the start and end of the line, so it only matches lines that start with a single letter, followed by 1-6 digits, and nothing else.- This doesn't account for lines with other characters before/after the number, or numbers that aren't immediately after a letter.
1. Extract Lines Containing 1-6 Digit Numbers
If you want lines that contain any 1-6 digit sequence (even if surrounded by other characters), use this regex to match such lines:
.*\d{1,6}.*
.*matches any characters (including none) before the digits.\d{1,6}matches 1 to 6 digits..*matches any characters after the digits.
If you want to ensure you're matching standalone 1-6 digit numbers (not part of a longer number like 1234567), use negative lookarounds to exclude digits before/after:
.*(?<!\d)\d{1,6}(?!\d).*
(?<!\d): Ensures no digit comes right before the sequence.(?!\d): Ensures no digit comes right after the sequence.
Example Implementation (Python)
Here's how to extract these lines using Python:
import re # Choose one pattern below pattern = r".*(?<!\d)\d{1,6}(?!\d).*" # For standalone 1-6 digit numbers # pattern = r".*\d{1,6}.*" # For any line with 1-6 digits (even part of longer numbers) with open("your_file.txt", "r") as f: for line in f: cleaned_line = line.strip() if re.match(pattern, cleaned_line): print(cleaned_line)
2. Look Up Entries by ETEXT NO
Assuming your file has entries structured like this (common for ETEXT-style files):
ETEXT NO: 1234
Title: The Adventures of Tom Sawyer
Author: Mark Twain
...
We can parse the file into a searchable dictionary where keys are ETEXT NOs, and values hold the title and author. Then you can easily look up entries.
Example Python Code
import re def parse_etext_entries(file_path): etext_database = {} current_entry = {} with open(file_path, "r") as f: for line in f: cleaned_line = line.strip() # Match ETEXT NO line etext_match = re.match(r"ETEXT NO:\s*(\d+)", cleaned_line) if etext_match: # Save previous entry if it exists if current_entry: etext_database[current_entry["etext_no"]] = current_entry # Start a new entry current_entry = {"etext_no": etext_match.group(1)} # Match Title line title_match = re.match(r"Title:\s*(.*)", cleaned_line) if title_match: current_entry["title"] = title_match.group(1) # Match Author line author_match = re.match(r"Author:\s*(.*)", cleaned_line) if author_match: current_entry["author"] = author_match.group(1) # Add the final entry to the database if current_entry: etext_database[current_entry["etext_no"]] = current_entry return etext_database # Usage example entries = parse_etext_entries("your_file.txt") # Look up a specific ETEXT NO target_no = "1234" if target_no in entries: entry = entries[target_no] print(f"ETEXT NO: {entry['etext_no']}") print(f"Title: {entry['title']}") print(f"Author: {entry['author']}") else: print(f"No entry found for ETEXT NO {target_no}")
Quick Notes:
- Adjust the regex patterns for
ETEXT NO:,Title:, andAuthor:if your file uses different formatting (e.g., all caps, colons in different places). - If your entries are separated by blank lines, add a check to reset
current_entrywhen a blank line is encountered.
内容的提问来源于stack exchange,提问作者Nazmus Shakib

