You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3 文件内容拆分与个人信息分类及性别判定需求

Got it, let's break down how to tackle this task step by step. I’ve worked on similar structured data extraction projects before, so here’s a practical, hands-on approach you can follow:

Step 1: Read and Parse the TXT File

First, you’ll need to load the file content and split each entry into its individual components. Assuming your TXT file uses a consistent delimiter (like commas, tabs, or pipes) to separate fields, Python makes this straightforward.

# Open and read the text file
with open('character_info.txt', 'r', encoding='utf-8') as file:
    # Skip the header line if your file has one (adjust as needed)
    header = file.readline()
    # Store all lines as a list of entries
    entries = [line.strip() for line in file if line.strip()]
Step 2: Classify Information into Ordered Lists

Since you need to preserve the original order, create dedicated lists for each data category. Append each field to its corresponding list as you iterate through the entries.

# Initialize ordered lists for each data type
last_names = []
first_names = []
ssns = []
addresses = []

# Process each entry to populate the lists
for entry in entries:
    # Split the entry using your file's delimiter (replace ',' with your actual delimiter)
    ln, fn, ssn, addr = entry.split(',')
    last_names.append(ln)
    first_names.append(fn)
    ssns.append(ssn)
    addresses.append(addr)
Step 3: Output Classified Information in Original Order

To verify the data is preserved correctly, loop through the lists using their shared index (since all lists will have the same length) and print each character’s full set of info.

# Print classified info in the original order
for idx in range(len(first_names)):
    print(f"Character {idx + 1}:")
    print(f"  Full Name: {first_names[idx]} {last_names[idx]}")
    print(f"  SSN: {ssns[idx]}")
    print(f"  Address: {addresses[idx]}")
    print("-" * 30)
Step 4: Determine Fictional Character Gender

For gender classification, the most reliable method for fictional characters is to use first name patterns. You can either use a pre-built library or maintain a custom dictionary of name-gender mappings (great for unique fictional names).

Option 1: Use a Name-Guessing Library

The gender-guesser library works well for common names. Install it first with pip install gender-guesser, then use this code:

import gender_guesser.detector as gender

# Initialize the gender detector
gd = gender.Detector()

# Guess gender for each character
for idx in range(len(first_names)):
    first_name = first_names[idx]
    gender_result = gd.get_gender(first_name)
    # Format the result for readability
    gender_label = {
        'male': 'Male',
        'female': 'Female',
        'andy': 'Ambiguous',
        'unknown': 'Unknown'
    }.get(gender_result, 'Unknown')
    print(f"{first_names[idx]} {last_names[idx]} | Gender: {gender_label}")

Option 2: Custom Name Dictionary (For Unique Fictional Names)

If your characters have made-up names, create a custom dictionary to map names directly to genders:

# Custom gender mapping for fictional names
fictional_gender_map = {
    "Frodo": "Male",
    "Arwen": "Female",
    "Gimli": "Male"
}

# Assign gender using the custom map
for idx in range(len(first_names)):
    first_name = first_names[idx]
    gender = fictional_gender_map.get(first_name, "Unknown")
    print(f"{first_names[idx]} {last_names[idx]} | Gender: {gender}")
Key Notes to Adjust for Your Use Case
  • If your TXT file has inconsistent formatting (e.g., line breaks between fields), add error handling with try-except blocks to avoid crashes.
  • For large datasets, consider using pandas DataFrames instead of lists—they’re easier to manage for sorting, filtering, and exporting later.

内容的提问来源于stack exchange,提问作者Matt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:24:48