You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则匹配:提取含分类文本及语料行内容的技术问题

Solutions for Your Text Extraction Problems

1. Fixing the Category + Text Regex Extraction

Your original regex approach didn't work for two key reasons:

  • The pattern \n[A-z]*\t assumes line breaks and tabs exist in your example text (which they don't — your sample is a single line of A <some text1> B <some text2> C <some text3>).
  • The greedy (.*) match eats up too much content before hitting your pattern, leading to incomplete, misaligned results.

Correct Regex Approach

Use a regex that matches each category (A/B/C) plus its associated text, stopping just before the next category or the end of the text. We'll use a positive lookahead to enforce this clean stop condition:

import re

test_doc = "A <some text1> B <some text2> C <some text3>"
# Regex breakdown:
# [A-Z] → Match a single uppercase category letter (adjust to [A-Za-z] if categories use lowercase)
# .*? → Non-greedily match any characters until...
# (?= [A-Z]|$) → Positive lookahead: either a space + next category letter, or end of string
category_pattern = r'([A-Z] .*?)(?= [A-Z]|$)'
results = re.findall(category_pattern, test_doc)

print(results)
# Output: ['A <some text1>', 'B <some text2>', 'C <some text3>']

This gives you exactly the expected list of category-text pairs.

2. Extracting Full Lines from the r8 Test Dataset

For the dataset where each line follows category<tab><sometext>, you don't even need regex for basic extraction — simple line reading is more efficient and straightforward.

Simple Line Reading Solution

Read the file line-by-line, filter out empty lines, and keep each full line intact:

# Open the file and read all non-empty lines
with open('r8-test-all-terms.txt', 'r', encoding='utf-8') as file:
    docs = [line.rstrip('\n') for line in file if line.strip()]

# docs will be your desired list:
# docs[0] = "category<tab><sometext1>"
# docs[1] = "category<tab><sometext2>"
# ...

Regex Validation (Optional)

If you want to ensure you only extract lines that strictly follow the category<tab>text format, use a regex with the re.MULTILINE flag to match each line's start/end:

import re

with open('r8-test-all-terms.txt', 'r', encoding='utf-8') as file:
    file_content = file.read()
    # Match lines with at least one tab, capturing the entire line
    valid_docs = re.findall(r'^[^\t]+\t.*$', file_content, re.MULTILINE)

The re.MULTILINE flag makes ^ and $ match the start and end of each line (not just the entire file), ensuring you get valid, complete lines.


内容的提问来源于stack exchange,提问作者IISC_Student

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:38:26