Python读取美国邮编文件并拆分数据至三个列表的技术问题
Parsing Zip Code Data into Separate Lists in Python
Hey there! Let's wrap up that script to parse your zip code file into three neat lists. I'll walk you through completing and improving your existing code, with clear explanations along the way.
Complete Working Code (Basic Version)
First, here's a straightforward completion of your initial code, with safer file handling and thorough data cleaning:
def main(): zipcode = [] city = [] state = [] # Use a `with` statement to auto-close the file (safer than manual close) with open("zipcodes.txt", "r") as file_object: theList = file_object.readlines() for line in theList: # Remove newline characters and extra spaces from the start/end of the line cleaned_line = line.strip() # Skip empty lines to avoid errors if not cleaned_line: continue # Split the line into three parts using commas as separators parts = cleaned_line.split(',') # Clean each part: remove surrounding quotes and spaces zip_part = parts[0].strip().strip("'") city_part = parts[1].strip().strip("'") state_part = parts[2].strip().strip("'") # Add the cleaned data to the respective lists zipcode.append(zip_part) city.append(city_part) state.append(state_part) # Quick test to verify the data print("First 5 entries:") for z, c, s in zip(zipcode[:5], city[:5], state[:5]): print(f"Zip: {z}, City: {c}, State: {s}") if __name__ == "__main__": main()
Key Improvements & Explanations
withStatement: This replaces your manualopen()call because it automatically closes the file when the code block finishes—no risk of leaving files open accidentally.- Line Cleaning:
line.strip()removes newline characters (\n) and any extra spaces at the start/end of each line. We also skip empty lines to prevent errors from trying to split an empty string. - Data Parsing:
split(',')breaks each line into three chunks. Then we usestrip().strip("'")to remove both surrounding spaces and single quotes, giving us clean, usable strings ready for your lists.
Optimized Version (For Large Files)
If your zip code file is huge, reading the entire thing into memory with readlines() isn't ideal. Here's a more memory-efficient version that reads line-by-line, plus added error handling for malformed lines:
def main(): zipcode = [] city = [] state = [] with open("zipcodes.txt", "r") as file_object: # Iterate directly over the file object (reads one line at a time) for line_num, line in enumerate(file_object, 1): cleaned_line = line.strip() if not cleaned_line: continue parts = cleaned_line.split(',') # Check if the line has exactly 3 parts to avoid index errors if len(parts) != 3: print(f"Warning: Skipping malformed line {line_num}: {line}") continue zip_part = parts[0].strip().strip("'") city_part = parts[1].strip().strip("'") state_part = parts[2].strip().strip("'") zipcode.append(zip_part) city.append(city_part) state.append(state_part) print(f"Successfully processed {len(zipcode)} valid entries!") print("Sample output:") for z, c, s in zip(zipcode[:3], city[:3], state[:3]): print(f"- {z} | {c}, {s}") if __name__ == "__main__": main()
This version:
- Uses direct file iteration to avoid loading the entire file into memory (perfect for large datasets)
- Adds line number tracking to report exactly where malformed lines are, making debugging easier
- Includes a check to skip lines that don't match the expected 3-part format, preventing unexpected crashes
内容的提问来源于stack exchange,提问作者Nick Nasty
相关产品推荐
相关产品推荐

