如何用Python将含startpayment/endpayment块的TXT解析为字典列表
Parse Russian Payment TXT into List of Dictionaries
Got it, let's tackle this parsing problem head-on. You need to turn each payment block (wrapped between startpayment and endpayment) into a dictionary, then gather all those dictionaries into a list. Here's a practical Python solution tailored exactly to your data structure:
Step-by-Step Solution Code
import re def parse_payments_file(file_path): # Read the file with UTF-8 encoding to handle Russian characters properly with open(file_path, 'r', encoding='utf-8') as f: content = f.read() # Split content into individual payment blocks (skip the empty first segment before the first startpayment) blocks = content.split('startpayment')[1:] payment_list = [] for block in blocks: # Extract the core content between startpayment and endpayment block_content = block.split('endpayment')[0].strip() # Remove the trailing "КонецДокумента" marker to clean up the block block_content = block_content.replace('КонецДокумента', '').strip() payment_dict = {} remaining_content = block_content while remaining_content: # Find the first '=' to split key and value eq_pos = remaining_content.find('=') if eq_pos == -1: break # No more valid key-value pairs left key = remaining_content[:eq_pos].strip() # Use regex to detect the next key pattern: whitespace + word + '=' # This ensures we don't split values that contain spaces (like payer names) next_key_match = re.search(r'\s+[A-ZА-Я][A-Za-zА-Яа-я]*=', remaining_content[eq_pos+1:]) if next_key_match: # Extract value up to the start of the next key value_end = eq_pos + 1 + next_key_match.start() value = remaining_content[eq_pos+1:value_end].strip() # Update remaining content to process the next pair remaining_content = remaining_content[value_end:].strip() else: # This is the last key-value pair in the block value = remaining_content[eq_pos+1:].strip() remaining_content = '' # Handle empty values (e.g., ДатаСписано=) by assigning an empty string payment_dict[key] = value if value else '' payment_list.append(payment_dict) return payment_list # Usage example payments = parse_payments_file('your_payments_file.txt') # Print the first payment record to verify the output print(payments[0])
Key Features Explained
- UTF-8 Encoding: Ensures Russian Cyrillic characters are read and processed without garbling.
- Block Isolation: Uses
startpaymentandendpaymentas delimiters to split the file into individual payment records. - Smart Value Capture: The regex pattern detects when a new key starts, so values with spaces (like full payer names or payment purposes) are captured in full instead of being split incorrectly.
- Empty Value Handling: Properly assigns empty strings to keys that have no associated value (e.g.,
ДатаСписано=). - Noise Cleanup: Removes the
КонецДокументаmarker from each block to avoid extra, unwanted entries in your dictionaries.
Testing with Your Sample Data
When you run this code on your provided sample text, it will generate a list of two dictionaries. Each dictionary will include all the key-value pairs from the corresponding payment block—including multi-word values and empty fields—exactly as they appear in the original file.
内容的提问来源于stack exchange,提问作者Eugene Trofimov
相关产品推荐
相关产品推荐

