寻求纯文本转对象的最优算法方案(含换行字段处理)
Great question—this is such a common pain point when working with semi-structured text tables, especially when fields like title can include whitespace or line breaks that break naive line-by-line, split-on-whitespace parsing. I’ve dealt with this exact scenario many times (legacy reports, scraped data, messy exports) and there are a few reliable rule-based approaches that work really well, depending on how consistent your data is.
Core Insight: Use Fixed-Format Fields as Anchors
Looking at your example, the critical observation is that the last three fields (quantity, price, sum) have strict, predictable formats:
quantityis an integerpriceandsumare decimal numbers (likely with fixed decimal places, like two digits)codeappears to be a single, whitespace-free identifier (e.g.,item1,item2)
These fixed-format fields act as "anchors" we can use to reverse-engineer each row, regardless of how messy the title field gets (even with line breaks).
Step-by-Step Solution Approach
1. Read the Entire Text Block (Not Line-by-Line)
Since title can span lines, forget about splitting on newlines first. Load the entire text as a single string—this lets us handle multi-line titles seamlessly.
2. Split the Text into Individual Items
Each item ends with a sum value (e.g., 3.00). Use this pattern to split the text into separate item blocks. For example, in regex, you can split on whitespace that follows a decimal number with two digits:
import re items = re.split(r'(?<=\d\.\d{2})\s+', raw_text.strip())
The (?<=...) is a positive lookbehind, so it only splits after a valid sum value, preserving the sum in each item block.
3. Reverse-Parse Each Item Block
For each item, work backwards from the end to extract the fixed fields first, then the remaining content is your title:
- Extract
sum: Grab the last decimal number with two digits. - Extract
price: From the remaining text, grab the next decimal number with two digits. - Extract
quantity: Grab the last integer from the remaining text. - Extract
code: Grab the last whitespace-free string (since yourcodevalues don’t have spaces). - Extract
title: Everything left is thetitle—replace any line breaks or extra whitespace with a single space to clean it up.
Example Code Implementation
Here’s a Python example that handles your sample data (including multi-line titles):
import re def parse_text_table(raw_text): # Split text into item blocks using sum values as anchors item_blocks = re.split(r'(?<=\d\.\d{2})\s+', raw_text.strip()) parsed_items = [] for block in item_blocks: # Extract sum (last decimal with 2 digits) sum_match = re.search(r'(\d+\.\d{2})$', block) if not sum_match: continue sum_val = sum_match.group(1) remaining = block[:sum_match.start()].strip() # Extract price price_match = re.search(r'(\d+\.\d{2})$', remaining) if not price_match: continue price_val = price_match.group(1) remaining = remaining[:price_match.start()].strip() # Extract quantity (integer) quantity_match = re.search(r'(\d+)$', remaining) if not quantity_match: continue quantity_val = quantity_match.group(1) remaining = remaining[:quantity_match.start()].strip() # Extract code (last whitespace-free string) code_match = re.search(r'(\S+)$', remaining) if not code_match: continue code_val = code_match.group(1) title_val = remaining[:code_match.start()].strip() # Clean up title: replace line breaks/extra spaces with single space title_val = re.sub(r'\s+', ' ', title_val) parsed_items.append({ 'title': title_val, 'code': code_val, 'quantity': int(quantity_val), 'price': float(price_val), 'sum': float(sum_val) }) return parsed_items # Test with your sample data (plus a multi-line title) sample_data = """Item1 item1 2 1.50 3.00 Item2 item2 1 3.00 3.00 good Item3 item3 2 1.00 2.00 Item4 item4 3 2.50 7.50 This is a multi-line title for item4""" result = parse_text_table(sample_data) for item in result: print(f"title = {item['title']}, code = {item['code']}, quantity = {item['quantity']}, price = {item['price']}, sum = {item['sum']}")
4. Handling Edge Cases
- If
codecan sometimes include spaces: You’ll need an additional anchor forcode(e.g., if it follows a specific pattern likeitem\d, use regex to match that directly instead of grabbing the last whitespace-free string). - If decimal formats vary: Adjust the regex to handle optional decimal places, different separators (commas instead of periods), etc.
- For extremely unstructured data: If fields have no consistent patterns, you might need to use lightweight NLP (like named entity recognition) to classify each segment, but that’s overkill for most cases like yours.
Final Thoughts
The reverse-parsing approach with fixed-format anchors is the most reliable and efficient solution here. It avoids the ambiguity of splitting on whitespace or line breaks, and leverages the parts of your data that are machine-readable by design.
内容的提问来源于stack exchange,提问作者Ben Sin

