如何在不使用Pandas的情况下通过Python逐行清理文本文件并生成含指定9列的列表的列表
Solution
Here's a Python script that processes your text file line by line, ignores the specified lines, and produces the required list of lists structure without using Pandas:
def process_football_data(file_path): result = [] current_team = None current_year = None with open(file_path, 'r') as file: for line in file: # Strip whitespace and remove leading/trailing | characters cleaned_line = line.strip().strip('|') if not cleaned_line: continue # Skip empty lines # Split into parts, stripping whitespace from each segment parts = [part.strip() for part in cleaned_line.split('|')] # Skip separator lines, header lines, and total summary lines if parts == ['---'] or parts[0] in ('Rnd', 'Total'): continue # Track current team or year from single-element lines if len(parts) == 1: content = parts[0] # Check if it's a 4-digit year if content.isdigit() and len(content) == 4: current_year = content else: current_team = content else: # Combine current team/year with data line to form 9 columns if current_team and current_year: data_row = [current_team, current_year] + parts result.append(data_row) return result # Example usage if __name__ == "__main__": output = process_football_data("football_data.txt") print(output)
How it works:
- Initialization: We set up variables to track the current team and year, plus an empty list to store our final results.
- Safe File Handling: Using a
withstatement ensures the file is properly closed after processing. - Line Cleaning: Each line is stripped of whitespace and surrounding
|characters to get the core content. - Filtering Unwanted Lines: We skip separators (
---), header lines starting withRnd, and total summary lines starting withTotal. - Tracking Context: Single-element lines are identified as either a team name or a 4-digit year, updating our context variables accordingly.
- Building Output Rows: Valid data lines (with 7 elements) are combined with the current team and year to form the required 9-column structure, then added to the result list.
Output:
Running this script with your input file will produce exactly the list structure you requested:
[ ["Arsenal", "2021", "R1", "Tottenham", "1", "0", "0", "Emirates", "March"], ["Arsenal", "2021", "R2", "Man utd", "0", "1", "0", "Old Trafford", "March"], ["Arsenal", "2020", "R1", "Chelsea", "1", "0", "0", "Stamford Bridge", "March"], ["Arsenal", "2020", "R2", "Mancity", "0", "1", "0", "Ethiad", "March"], ["Man Utd", "2021", "R1", "Chelsea", "1", "0", "0", "Emirates", "March"], ["Man Utd", "2021", "R2", "Wolves", "0", "1", "0", "Old Trafford", "March"], ["Man utd", "2020", "R1", "Tottenham", "1", "0", "0", "Stamford Bridge", "March"], ["Man utd", "2020", "R2", "Palace", "0", "1", "0", "Ethiad", "March"] ]
Notes:
- The script handles minor whitespace variations in the input (like extra spaces around
|characters) by stripping whitespace from each segment. - It only adds data lines to the result if both team and year context have been captured, ensuring no incomplete rows are included.
内容的提问来源于stack exchange,提问作者user16603508
相关产品推荐
相关产品推荐

