如何用Pandas从TXT文件创建DataFrame并处理异常行
Here's a practical, step-by-step approach to parse your file.txt into a pandas DataFrame while handling common anomalies like misplaced spaces in fields:
Step 1: Read and Preprocess the File
First, we'll load the file, clean up unnecessary whitespace, and filter out empty lines:
import pandas as pd # Read the file and clean raw lines with open('file.txt', 'r') as f: # Strip whitespace from each line and skip empty lines lines = [line.strip() for line in f if line.strip()]
Step 2: Group Lines into Student Entries
Each student entry is made up of 6 consecutive lines (from NISN to E-mail). We'll split the cleaned lines into chunks of 6:
# Split lines into chunks of 6 (one chunk per student) entry_chunks = [lines[i:i+6] for i in range(0, len(lines), 6)]
Step 3: Process Entries and Fix Anomalies
For each chunk, we'll extract key-value pairs, clean values, and handle common issues like spaces in NISN or school names:
student_entries = [] # Define the exact column order we want in the final DataFrame desired_columns = ['NISN', 'FullName', 'FirstName', 'LastName', 'School', 'E-mail'] for chunk in entry_chunks: # Initialize entry with None for all columns (handles missing lines) entry = {col: None for col in desired_columns} # Skip chunks that don't have exactly 6 lines (incomplete entries) if len(chunk) != 6: print(f"Skipping invalid entry (missing lines): {chunk}") student_entries.append(entry) continue for line in chunk: # Remove the leading "# " prefix from each line line_clean = line[2:] if line.startswith('# ') else line # Split into key and value (split only once to avoid breaking values with colons) try: key, value = line_clean.split(':- ', 1) except ValueError: print(f"Skipping malformed line: {line}") continue # Clean values based on field type if key == 'NISN': # Remove all whitespace from NISN to fix typos like "123456 7" cleaned_value = value.replace(' ', '') # Optional: Warn if NISN isn't numeric after cleaning if not cleaned_value.isdigit(): print(f"Warning: Non-numeric NISN found: {cleaned_value}") else: # For other fields, just strip leading/trailing whitespace (preserve internal spaces) cleaned_value = value.strip() # Update the entry if the key matches our desired columns if key in entry: entry[key] = cleaned_value student_entries.append(entry)
Step 4: Create the Final DataFrame
Convert the list of processed entries into a pandas DataFrame and reorder columns to match your desired structure:
# Convert entries to DataFrame df = pd.DataFrame(student_entries) # Reorder columns to match your preferred layout df = df[desired_columns] # View the result print(df)
Key Anomaly Handlers Breakdown
- Spaces in NISN: We strip all whitespace from NISN to fix typos like
"123456 7"→"1234567". - Spaces in School Names: We keep internal spaces intact (e.g.,
"Kimc il"stays as-is) since school names often include spaces. - Incomplete Entries: Chunks with fewer than 6 lines are flagged, and missing fields are filled with
NaN. - Malformed Lines: Lines that don't follow the
# Key:- Valueformat are skipped, with their corresponding field left empty.
Example Output
For your abnormal input example, the resulting DataFrame will look like this:
| NISN | FullName | FirstName | LastName | School | |
|---|---|---|---|---|---|
| 1234567 | Joe Doe | Joe | Doe | Klima | joe@gmail.com |
| 8901234 | Jenny Low | Jenny | Low | Kimc il | jenny@gmail.com |
内容的提问来源于stack exchange,提问作者Hendra

