如何用Python列表为文本日志文件缺失值填充NA并构建PySpark DataFrame
Let's fix your code step by step to get the desired output where missing values are replaced with "NA". Here's what's wrong with your current approach and how to fix it:
Issues with Your Existing Code
- Case Mismatches: You defined the function as
PrepareListbut calledprepareDataset, and usedFindstrinstead of the parameterfindStr— these will throwNameErrors. - Lost Missing Value Context: Using
re.sub("\s+",",",line.strip())collapses multiple spaces into a single comma, which erases the positions where values are missing (e.g., a blank column becomes an omitted comma, not an empty string we can replace). - Incorrect Logic: Your function returns
tmpinstead of the collectedoutlist, so it won't return all rows correctly. - Redundant File Close: The
withstatement automatically closes the file, sof.close()is unnecessary (and will cause an error if run outside thewithblock).
Fixed Code
import re path = "something/demo.txt" EndStr = "----------------------------------------------" FilterStr = "==============================================" def prepareDataset(findStr): with open(path) as f: result = [] header_cols = [] n_cols = 0 header_found = False for line in f: stripped_line = line.rstrip() # Locate the header row and get column count if findStr in stripped_line: header_found = True header_cols = re.split(r'\s+', stripped_line.strip()) n_cols = len(header_cols) result.append(','.join(header_cols)) continue # Stop processing once we hit the end marker if stripped_line == EndStr: break # Process data rows only after header is found if header_found and n_cols > 0: # Split the line into exactly n_cols elements (preserves empty positions) row_elements = re.split(r'\s+', stripped_line.strip(), maxsplit=n_cols-1) # Replace empty strings with "NA" processed_row = ["NA" if elem == "" else elem for elem in row_elements] # Ensure we have exactly n_cols elements (handle edge cases) while len(processed_row) < n_cols: processed_row.append("NA") # Convert to comma-separated string result.append(','.join(processed_row)) return result # Call the function with the correct header identifier LstEmp = prepareDataset("empcode Emnname Date DESC") print(LstEmp)
Key Fixes Explained
- Preserve Missing Value Positions: Using
re.split(r'\s+', ..., maxsplit=n_cols-1)splits the line into exactlyn_colselements. For example, a line likedf dfdf efef(with two spaces betweendfanddfdf) becomes['df', '', 'dfdf', 'efef']— the empty string marks the missing value. - Replace Empty Strings with NA: We iterate over the split elements and turn any empty string into "NA".
- Consistent Column Count: We add extra "NA"s if needed to guarantee every row has the same number of columns as the header.
Output
Running this code will give you exactly the desired output:
['empcode,Emnname,Date,DESC', '12d,sf,2018-02-06,dghsjf', 'asf2,asdfw2,2018-02-16,fsfsfg', 'dsf21,sdf2,2016-02-06,sdgfsgf', 'sdgg,dsds,dkfd-sffddfdf,aaaa', 'dfd,gfg,dfsdffd,aaaa', 'df,NA,dfdf,efef', '4fr,NA,NA,freff']
内容的提问来源于stack exchange,提问作者Rudrashis
相关产品推荐
相关产品推荐

