Python3 文件内容拆分与个人信息分类及性别判定需求
Got it, let's break down how to tackle this task step by step. I’ve worked on similar structured data extraction projects before, so here’s a practical, hands-on approach you can follow:
First, you’ll need to load the file content and split each entry into its individual components. Assuming your TXT file uses a consistent delimiter (like commas, tabs, or pipes) to separate fields, Python makes this straightforward.
# Open and read the text file with open('character_info.txt', 'r', encoding='utf-8') as file: # Skip the header line if your file has one (adjust as needed) header = file.readline() # Store all lines as a list of entries entries = [line.strip() for line in file if line.strip()]
Since you need to preserve the original order, create dedicated lists for each data category. Append each field to its corresponding list as you iterate through the entries.
# Initialize ordered lists for each data type last_names = [] first_names = [] ssns = [] addresses = [] # Process each entry to populate the lists for entry in entries: # Split the entry using your file's delimiter (replace ',' with your actual delimiter) ln, fn, ssn, addr = entry.split(',') last_names.append(ln) first_names.append(fn) ssns.append(ssn) addresses.append(addr)
To verify the data is preserved correctly, loop through the lists using their shared index (since all lists will have the same length) and print each character’s full set of info.
# Print classified info in the original order for idx in range(len(first_names)): print(f"Character {idx + 1}:") print(f" Full Name: {first_names[idx]} {last_names[idx]}") print(f" SSN: {ssns[idx]}") print(f" Address: {addresses[idx]}") print("-" * 30)
For gender classification, the most reliable method for fictional characters is to use first name patterns. You can either use a pre-built library or maintain a custom dictionary of name-gender mappings (great for unique fictional names).
Option 1: Use a Name-Guessing Library
The gender-guesser library works well for common names. Install it first with pip install gender-guesser, then use this code:
import gender_guesser.detector as gender # Initialize the gender detector gd = gender.Detector() # Guess gender for each character for idx in range(len(first_names)): first_name = first_names[idx] gender_result = gd.get_gender(first_name) # Format the result for readability gender_label = { 'male': 'Male', 'female': 'Female', 'andy': 'Ambiguous', 'unknown': 'Unknown' }.get(gender_result, 'Unknown') print(f"{first_names[idx]} {last_names[idx]} | Gender: {gender_label}")
Option 2: Custom Name Dictionary (For Unique Fictional Names)
If your characters have made-up names, create a custom dictionary to map names directly to genders:
# Custom gender mapping for fictional names fictional_gender_map = { "Frodo": "Male", "Arwen": "Female", "Gimli": "Male" } # Assign gender using the custom map for idx in range(len(first_names)): first_name = first_names[idx] gender = fictional_gender_map.get(first_name, "Unknown") print(f"{first_names[idx]} {last_names[idx]} | Gender: {gender}")
- If your TXT file has inconsistent formatting (e.g., line breaks between fields), add error handling with
try-exceptblocks to avoid crashes. - For large datasets, consider using
pandasDataFrames instead of lists—they’re easier to manage for sorting, filtering, and exporting later.
内容的提问来源于stack exchange,提问作者Matt

