You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python列表为文本日志文件缺失值填充NA并构建PySpark DataFrame

Let's fix your code step by step to get the desired output where missing values are replaced with "NA". Here's what's wrong with your current approach and how to fix it:

Issues with Your Existing Code

  1. Case Mismatches: You defined the function as PrepareList but called prepareDataset, and used Findstr instead of the parameter findStr — these will throw NameErrors.
  2. Lost Missing Value Context: Using re.sub("\s+",",",line.strip()) collapses multiple spaces into a single comma, which erases the positions where values are missing (e.g., a blank column becomes an omitted comma, not an empty string we can replace).
  3. Incorrect Logic: Your function returns tmp instead of the collected out list, so it won't return all rows correctly.
  4. Redundant File Close: The with statement automatically closes the file, so f.close() is unnecessary (and will cause an error if run outside the with block).

Fixed Code

import re

path = "something/demo.txt"
EndStr = "----------------------------------------------"
FilterStr = "=============================================="

def prepareDataset(findStr):
    with open(path) as f:
        result = []
        header_cols = []
        n_cols = 0
        header_found = False
        
        for line in f:
            stripped_line = line.rstrip()
            
            # Locate the header row and get column count
            if findStr in stripped_line:
                header_found = True
                header_cols = re.split(r'\s+', stripped_line.strip())
                n_cols = len(header_cols)
                result.append(','.join(header_cols))
                continue
            
            # Stop processing once we hit the end marker
            if stripped_line == EndStr:
                break
            
            # Process data rows only after header is found
            if header_found and n_cols > 0:
                # Split the line into exactly n_cols elements (preserves empty positions)
                row_elements = re.split(r'\s+', stripped_line.strip(), maxsplit=n_cols-1)
                # Replace empty strings with "NA"
                processed_row = ["NA" if elem == "" else elem for elem in row_elements]
                # Ensure we have exactly n_cols elements (handle edge cases)
                while len(processed_row) < n_cols:
                    processed_row.append("NA")
                # Convert to comma-separated string
                result.append(','.join(processed_row))
        
        return result

# Call the function with the correct header identifier
LstEmp = prepareDataset("empcode Emnname Date DESC")
print(LstEmp)

Key Fixes Explained

  • Preserve Missing Value Positions: Using re.split(r'\s+', ..., maxsplit=n_cols-1) splits the line into exactly n_cols elements. For example, a line like df dfdf efef (with two spaces between df and dfdf) becomes ['df', '', 'dfdf', 'efef'] — the empty string marks the missing value.
  • Replace Empty Strings with NA: We iterate over the split elements and turn any empty string into "NA".
  • Consistent Column Count: We add extra "NA"s if needed to guarantee every row has the same number of columns as the header.

Output

Running this code will give you exactly the desired output:

['empcode,Emnname,Date,DESC', '12d,sf,2018-02-06,dghsjf', 'asf2,asdfw2,2018-02-16,fsfsfg', 'dsf21,sdf2,2016-02-06,sdgfsgf', 'sdgg,dsds,dkfd-sffddfdf,aaaa', 'dfd,gfg,dfsdffd,aaaa', 'df,NA,dfdf,efef', '4fr,NA,NA,freff']

内容的提问来源于stack exchange,提问作者Rudrashis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:42:26