You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

寻求纯文本转对象的最优算法方案(含换行字段处理)

Parsing Text Tables with Variable-Length Fields (Line Breaks in Title/Code)

Great question—this is such a common pain point when working with semi-structured text tables, especially when fields like title can include whitespace or line breaks that break naive line-by-line, split-on-whitespace parsing. I’ve dealt with this exact scenario many times (legacy reports, scraped data, messy exports) and there are a few reliable rule-based approaches that work really well, depending on how consistent your data is.

Core Insight: Use Fixed-Format Fields as Anchors

Looking at your example, the critical observation is that the last three fields (quantity, price, sum) have strict, predictable formats:

  • quantity is an integer
  • price and sum are decimal numbers (likely with fixed decimal places, like two digits)
  • code appears to be a single, whitespace-free identifier (e.g., item1, item2)

These fixed-format fields act as "anchors" we can use to reverse-engineer each row, regardless of how messy the title field gets (even with line breaks).

Step-by-Step Solution Approach

1. Read the Entire Text Block (Not Line-by-Line)

Since title can span lines, forget about splitting on newlines first. Load the entire text as a single string—this lets us handle multi-line titles seamlessly.

2. Split the Text into Individual Items

Each item ends with a sum value (e.g., 3.00). Use this pattern to split the text into separate item blocks. For example, in regex, you can split on whitespace that follows a decimal number with two digits:

import re
items = re.split(r'(?<=\d\.\d{2})\s+', raw_text.strip())

The (?<=...) is a positive lookbehind, so it only splits after a valid sum value, preserving the sum in each item block.

3. Reverse-Parse Each Item Block

For each item, work backwards from the end to extract the fixed fields first, then the remaining content is your title:

  • Extract sum: Grab the last decimal number with two digits.
  • Extract price: From the remaining text, grab the next decimal number with two digits.
  • Extract quantity: Grab the last integer from the remaining text.
  • Extract code: Grab the last whitespace-free string (since your code values don’t have spaces).
  • Extract title: Everything left is the title—replace any line breaks or extra whitespace with a single space to clean it up.

Example Code Implementation

Here’s a Python example that handles your sample data (including multi-line titles):

import re

def parse_text_table(raw_text):
    # Split text into item blocks using sum values as anchors
    item_blocks = re.split(r'(?<=\d\.\d{2})\s+', raw_text.strip())
    parsed_items = []
    
    for block in item_blocks:
        # Extract sum (last decimal with 2 digits)
        sum_match = re.search(r'(\d+\.\d{2})$', block)
        if not sum_match:
            continue
        sum_val = sum_match.group(1)
        remaining = block[:sum_match.start()].strip()
        
        # Extract price
        price_match = re.search(r'(\d+\.\d{2})$', remaining)
        if not price_match:
            continue
        price_val = price_match.group(1)
        remaining = remaining[:price_match.start()].strip()
        
        # Extract quantity (integer)
        quantity_match = re.search(r'(\d+)$', remaining)
        if not quantity_match:
            continue
        quantity_val = quantity_match.group(1)
        remaining = remaining[:quantity_match.start()].strip()
        
        # Extract code (last whitespace-free string)
        code_match = re.search(r'(\S+)$', remaining)
        if not code_match:
            continue
        code_val = code_match.group(1)
        title_val = remaining[:code_match.start()].strip()
        
        # Clean up title: replace line breaks/extra spaces with single space
        title_val = re.sub(r'\s+', ' ', title_val)
        
        parsed_items.append({
            'title': title_val,
            'code': code_val,
            'quantity': int(quantity_val),
            'price': float(price_val),
            'sum': float(sum_val)
        })
    
    return parsed_items

# Test with your sample data (plus a multi-line title)
sample_data = """Item1 item1 2 1.50 3.00
Item2 item2 1 3.00 3.00 good
Item3 item3 2 1.00 2.00
Item4 item4 3 2.50 7.50
This is a multi-line
title for item4"""

result = parse_text_table(sample_data)
for item in result:
    print(f"title = {item['title']}, code = {item['code']}, quantity = {item['quantity']}, price = {item['price']}, sum = {item['sum']}")

4. Handling Edge Cases

  • If code can sometimes include spaces: You’ll need an additional anchor for code (e.g., if it follows a specific pattern like item\d, use regex to match that directly instead of grabbing the last whitespace-free string).
  • If decimal formats vary: Adjust the regex to handle optional decimal places, different separators (commas instead of periods), etc.
  • For extremely unstructured data: If fields have no consistent patterns, you might need to use lightweight NLP (like named entity recognition) to classify each segment, but that’s overkill for most cases like yours.

Final Thoughts

The reverse-parsing approach with fixed-format anchors is the most reliable and efficient solution here. It avoids the ambiguity of splitting on whitespace or line breaks, and leverages the parts of your data that are machine-readable by design.

内容的提问来源于stack exchange,提问作者Ben Sin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:06:59