You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas从TXT文件创建DataFrame并处理异常行

Solution to Convert Text File to DataFrame with Anomaly Handling

Here's a practical, step-by-step approach to parse your file.txt into a pandas DataFrame while handling common anomalies like misplaced spaces in fields:

Step 1: Read and Preprocess the File

First, we'll load the file, clean up unnecessary whitespace, and filter out empty lines:

import pandas as pd

# Read the file and clean raw lines
with open('file.txt', 'r') as f:
    # Strip whitespace from each line and skip empty lines
    lines = [line.strip() for line in f if line.strip()]

Step 2: Group Lines into Student Entries

Each student entry is made up of 6 consecutive lines (from NISN to E-mail). We'll split the cleaned lines into chunks of 6:

# Split lines into chunks of 6 (one chunk per student)
entry_chunks = [lines[i:i+6] for i in range(0, len(lines), 6)]

Step 3: Process Entries and Fix Anomalies

For each chunk, we'll extract key-value pairs, clean values, and handle common issues like spaces in NISN or school names:

student_entries = []
# Define the exact column order we want in the final DataFrame
desired_columns = ['NISN', 'FullName', 'FirstName', 'LastName', 'School', 'E-mail']

for chunk in entry_chunks:
    # Initialize entry with None for all columns (handles missing lines)
    entry = {col: None for col in desired_columns}
    
    # Skip chunks that don't have exactly 6 lines (incomplete entries)
    if len(chunk) != 6:
        print(f"Skipping invalid entry (missing lines): {chunk}")
        student_entries.append(entry)
        continue
    
    for line in chunk:
        # Remove the leading "# " prefix from each line
        line_clean = line[2:] if line.startswith('# ') else line
        
        # Split into key and value (split only once to avoid breaking values with colons)
        try:
            key, value = line_clean.split(':- ', 1)
        except ValueError:
            print(f"Skipping malformed line: {line}")
            continue
        
        # Clean values based on field type
        if key == 'NISN':
            # Remove all whitespace from NISN to fix typos like "123456 7"
            cleaned_value = value.replace(' ', '')
            # Optional: Warn if NISN isn't numeric after cleaning
            if not cleaned_value.isdigit():
                print(f"Warning: Non-numeric NISN found: {cleaned_value}")
        else:
            # For other fields, just strip leading/trailing whitespace (preserve internal spaces)
            cleaned_value = value.strip()
        
        # Update the entry if the key matches our desired columns
        if key in entry:
            entry[key] = cleaned_value
    
    student_entries.append(entry)

Step 4: Create the Final DataFrame

Convert the list of processed entries into a pandas DataFrame and reorder columns to match your desired structure:

# Convert entries to DataFrame
df = pd.DataFrame(student_entries)

# Reorder columns to match your preferred layout
df = df[desired_columns]

# View the result
print(df)

Key Anomaly Handlers Breakdown

  • Spaces in NISN: We strip all whitespace from NISN to fix typos like "123456 7" → "1234567".
  • Spaces in School Names: We keep internal spaces intact (e.g., "Kimc il" stays as-is) since school names often include spaces.
  • Incomplete Entries: Chunks with fewer than 6 lines are flagged, and missing fields are filled with NaN.
  • Malformed Lines: Lines that don't follow the # Key:- Value format are skipped, with their corresponding field left empty.

Example Output

For your abnormal input example, the resulting DataFrame will look like this:

NISNFullNameFirstNameLastNameSchoolE-mail
1234567Joe DoeJoeDoeKlimajoe@gmail.com
8901234Jenny LowJennyLowKimc iljenny@gmail.com

内容的提问来源于stack exchange,提问作者Hendra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 19:17:40