Python正则匹配:提取含分类文本及语料行内容的技术问题
1. Fixing the Category + Text Regex Extraction
Your original regex approach didn't work for two key reasons:
- The pattern
\n[A-z]*\tassumes line breaks and tabs exist in your example text (which they don't — your sample is a single line ofA <some text1> B <some text2> C <some text3>). - The greedy
(.*)match eats up too much content before hitting your pattern, leading to incomplete, misaligned results.
Correct Regex Approach
Use a regex that matches each category (A/B/C) plus its associated text, stopping just before the next category or the end of the text. We'll use a positive lookahead to enforce this clean stop condition:
import re test_doc = "A <some text1> B <some text2> C <some text3>" # Regex breakdown: # [A-Z] → Match a single uppercase category letter (adjust to [A-Za-z] if categories use lowercase) # .*? → Non-greedily match any characters until... # (?= [A-Z]|$) → Positive lookahead: either a space + next category letter, or end of string category_pattern = r'([A-Z] .*?)(?= [A-Z]|$)' results = re.findall(category_pattern, test_doc) print(results) # Output: ['A <some text1>', 'B <some text2>', 'C <some text3>']
This gives you exactly the expected list of category-text pairs.
2. Extracting Full Lines from the r8 Test Dataset
For the dataset where each line follows category<tab><sometext>, you don't even need regex for basic extraction — simple line reading is more efficient and straightforward.
Simple Line Reading Solution
Read the file line-by-line, filter out empty lines, and keep each full line intact:
# Open the file and read all non-empty lines with open('r8-test-all-terms.txt', 'r', encoding='utf-8') as file: docs = [line.rstrip('\n') for line in file if line.strip()] # docs will be your desired list: # docs[0] = "category<tab><sometext1>" # docs[1] = "category<tab><sometext2>" # ...
Regex Validation (Optional)
If you want to ensure you only extract lines that strictly follow the category<tab>text format, use a regex with the re.MULTILINE flag to match each line's start/end:
import re with open('r8-test-all-terms.txt', 'r', encoding='utf-8') as file: file_content = file.read() # Match lines with at least one tab, capturing the entire line valid_docs = re.findall(r'^[^\t]+\t.*$', file_content, re.MULTILINE)
The re.MULTILINE flag makes ^ and $ match the start and end of each line (not just the entire file), ensuring you get valid, complete lines.
内容的提问来源于stack exchange,提问作者IISC_Student

