You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将含startpayment/endpayment块的TXT解析为字典列表

Parse Russian Payment TXT into List of Dictionaries

Got it, let's tackle this parsing problem head-on. You need to turn each payment block (wrapped between startpayment and endpayment) into a dictionary, then gather all those dictionaries into a list. Here's a practical Python solution tailored exactly to your data structure:

Step-by-Step Solution Code

import re

def parse_payments_file(file_path):
    # Read the file with UTF-8 encoding to handle Russian characters properly
    with open(file_path, 'r', encoding='utf-8') as f:
        content = f.read()
    
    # Split content into individual payment blocks (skip the empty first segment before the first startpayment)
    blocks = content.split('startpayment')[1:]
    payment_list = []
    
    for block in blocks:
        # Extract the core content between startpayment and endpayment
        block_content = block.split('endpayment')[0].strip()
        # Remove the trailing "КонецДокумента" marker to clean up the block
        block_content = block_content.replace('КонецДокумента', '').strip()
        
        payment_dict = {}
        remaining_content = block_content
        
        while remaining_content:
            # Find the first '=' to split key and value
            eq_pos = remaining_content.find('=')
            if eq_pos == -1:
                break  # No more valid key-value pairs left
            
            key = remaining_content[:eq_pos].strip()
            # Use regex to detect the next key pattern: whitespace + word + '='
            # This ensures we don't split values that contain spaces (like payer names)
            next_key_match = re.search(r'\s+[A-ZА-Я][A-Za-zА-Яа-я]*=', remaining_content[eq_pos+1:])
            
            if next_key_match:
                # Extract value up to the start of the next key
                value_end = eq_pos + 1 + next_key_match.start()
                value = remaining_content[eq_pos+1:value_end].strip()
                # Update remaining content to process the next pair
                remaining_content = remaining_content[value_end:].strip()
            else:
                # This is the last key-value pair in the block
                value = remaining_content[eq_pos+1:].strip()
                remaining_content = ''
            
            # Handle empty values (e.g., ДатаСписано=) by assigning an empty string
            payment_dict[key] = value if value else ''
        
        payment_list.append(payment_dict)
    
    return payment_list

# Usage example
payments = parse_payments_file('your_payments_file.txt')
# Print the first payment record to verify the output
print(payments[0])

Key Features Explained

  • UTF-8 Encoding: Ensures Russian Cyrillic characters are read and processed without garbling.
  • Block Isolation: Uses startpayment and endpayment as delimiters to split the file into individual payment records.
  • Smart Value Capture: The regex pattern detects when a new key starts, so values with spaces (like full payer names or payment purposes) are captured in full instead of being split incorrectly.
  • Empty Value Handling: Properly assigns empty strings to keys that have no associated value (e.g., ДатаСписано=).
  • Noise Cleanup: Removes the КонецДокумента marker from each block to avoid extra, unwanted entries in your dictionaries.

Testing with Your Sample Data

When you run this code on your provided sample text, it will generate a list of two dictionaries. Each dictionary will include all the key-value pairs from the corresponding payment block—including multi-word values and empty fields—exactly as they appear in the original file.

内容的提问来源于stack exchange,提问作者Eugene Trofimov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:45:36