CSV文件正则匹配需求:从'B08-1506'起匹配至下一同类模式并合并多行
Got it, let's work through this problem step by step. Your goal is to split your CSV into blocks starting with identifiers like B08-1506 (following the BXX-XXXX pattern, where X is a digit), capture each block up to the next identifier, and merge any multi-line content in each block into a single line. Here's how to do it with regex:
Regex Pattern (With DOTALL Mode Enabled)
First, we need a regex that can span lines and stop at the next identifier. Use this pattern (make sure your regex engine supports DOTALL mode, which lets . match newlines):
(B\d{2}-\d{4})(.*?)(?=B\d{2}-\d{4}|$)
Breakdown of the Pattern:
(B\d{2}-\d{4}): Captures the starting identifier (e.g.,B08-1506) as Group 1. This matches exactly two digits afterB, a hyphen, then four digits.(.*?): Non-greedy match of all characters (including newlines, thanks to DOTALL) for the rest of the block. The non-greedy?ensures we stop at the next identifier instead of matching all the way to the end of the file.(?=B\d{2}-\d{4}|$): Positive lookahead that tells the regex to stop when it hits either the nextBXX-XXXXidentifier or the end of the file.
Merging Multi-Line Content into Single Lines
Once you capture each block, you'll need to clean up line breaks and extra whitespace. Here's how to implement this in Python as an example:
import re # Your raw CSV content (with line breaks) raw_csv = """B08-1506,324873, st, $0.0, ljkflka,,,,, 1 of 37 jksdfhjfhjk jkdsfh,,,,,,, B08-1606,324873, st, $0.0, ljkflka,,,,, 1 of 37 jksdfhjfhjk jkdsfh,,,,,,, B09-0680,324873, st, $0.0, ljkflka,,,,, 1 of 37 jksdfhjfhjk jkdsfh,,,,,,, B09-1234,...""" # Compile regex with DOTALL mode pattern = re.compile(r'(B\d{2}-\d{4})(.*?)(?=B\d{2}-\d{4}|$)', re.DOTALL) # Process each matched block cleaned_lines = [] for match in pattern.finditer(raw_csv): id_segment = match.group(1) content_segment = match.group(2) # Replace all newlines and multiple spaces with a single space, then trim edges cleaned_content = re.sub(r'\s+', ' ', content_segment).strip() # Combine into a single line and add to results cleaned_lines.append(f"{id_segment}{cleaned_content}") # Print or save the cleaned lines for line in cleaned_lines: print(line)
For Text Editors (Like VS Code)
If you're using a text editor to do this manually:
- Enable regex mode and the "Match newlines with .*" option (DOTALL equivalent)
- Find:
(B\d{2}-\d{4})([\s\S]*?)(?=B\d{2}-\d{4}|$) - Replace:
$1${2//\n/ }(adjust syntax if needed—some editors use\1instead of$1, and you may need to replace newlines with spaces in a separate step if the editor doesn't support inline replacements)
Key Notes:
- Always use non-greedy matching (
.*?) here—without it, the regex will match from the first identifier all the way to the last one, which isn't what you want. - DOTALL mode is essential to capture content across line breaks in each block.
- Cleaning whitespace ensures your final CSV lines don't have messy line breaks or redundant spaces.
内容的提问来源于stack exchange,提问作者veilupearl

