You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python文本文件清洗:自动分隔无明显间隔的数值

Solution: Automatically Split Numerical Values with Missing Spaces

To solve your problem, we'll use a regular expression to insert spaces before minus signs (-) that aren't preceded by an E (or e, for case insensitivity) — this handles scientific notation correctly while splitting adjacent numerical values. Here's the modified code and breakdown:

Step-by-Step Changes

  1. Import the re module: Required for regex operations.
  2. Add line processing in the data row handler: Use a negative lookbehind regex to insert spaces where needed before splitting the line.

Modified Full Code

import re

class Block:
    def __init__(self):
        self.data = {}
        self.array = []

    def ingest(self, lines):
        for line in lines:
            
            # Skip empty lines
            if line.strip() == '':
                continue
            
            # Process lines starting with letters (e.g., CSYS, NBLOCK)
            elif line[0].isalpha():
                parts = [k.strip() for k in line.split(',')] 
                # Convert numeric strings to floats, leave others as-is
                parts = [float(k) if k.replace('.', '').isdigit() else k for k in parts]
                key = parts[0].lower()
                if key not in self.data:
                    self.data[key] = [parts[1:]]
                else:
                    self.data[key].append( parts[1:] )
            
            # Skip comment/metadata lines
            elif line[0] in '\t/*_!':
                continue

            # Process format specifier lines (e.g., (3i9,6e21.13e3))
            elif line[0] == '(':
                self.data[key].append( [line.strip()] )
                self.data[key].append( [] )
            else:
                # Insert space before '-' unless preceded by E/e
                processed_line = re.sub(r'(?<![Ee])-', r' -', line)
                parts = processed_line.split()
                if parts[0] != '-1':
                    # Convert all parts to floats
                    self.data[key][-1].append( [float(k) for k in parts] )


# Example usage
input_file = "your_input_file.txt"
blk = Block() 
blk.ingest( open(input_file) ) 
my_data = blk.data

# Print the nblock data to verify
print("nblock =", my_data['nblock'])

Key Regex Explanation

The regex r'(?<![Ee])-' uses a negative lookbehind:

  • (?<![Ee]): Ensures the character before - is not E or e (preserving scientific notation like 1.2E-05).
  • We replace matches with ' -' to create a space-separated value that split() can handle correctly.

Verification with Sample Input

For your sample input lines:

  • Original: 1 0 0 4.0000000000000E+001-6.0000000000000E+001
  • Processed: 1 0 0 4.0000000000000E+001 -6.0000000000000E+001
  • Splits into parts that convert to [1.0, 0.0, 0.0, 40.0, -60.0] as expected.

Output

Running the code with your sample input will produce exactly the nblock structure you specified.

内容的提问来源于stack exchange,提问作者livelysteak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 23:24:37