You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取美国邮编文件并拆分数据至三个列表的技术问题

Parsing Zip Code Data into Separate Lists in Python

Hey there! Let's wrap up that script to parse your zip code file into three neat lists. I'll walk you through completing and improving your existing code, with clear explanations along the way.

Complete Working Code (Basic Version)

First, here's a straightforward completion of your initial code, with safer file handling and thorough data cleaning:

def main():
    zipcode = []
    city = []
    state = []
    
    # Use a `with` statement to auto-close the file (safer than manual close)
    with open("zipcodes.txt", "r") as file_object:
        theList = file_object.readlines()
        
        for line in theList:
            # Remove newline characters and extra spaces from the start/end of the line
            cleaned_line = line.strip()
            # Skip empty lines to avoid errors
            if not cleaned_line:
                continue
            # Split the line into three parts using commas as separators
            parts = cleaned_line.split(',')
            # Clean each part: remove surrounding quotes and spaces
            zip_part = parts[0].strip().strip("'")
            city_part = parts[1].strip().strip("'")
            state_part = parts[2].strip().strip("'")
            # Add the cleaned data to the respective lists
            zipcode.append(zip_part)
            city.append(city_part)
            state.append(state_part)
    
    # Quick test to verify the data
    print("First 5 entries:")
    for z, c, s in zip(zipcode[:5], city[:5], state[:5]):
        print(f"Zip: {z}, City: {c}, State: {s}")

if __name__ == "__main__":
    main()

Key Improvements & Explanations

  • with Statement: This replaces your manual open() call because it automatically closes the file when the code block finishes—no risk of leaving files open accidentally.
  • Line Cleaning: line.strip() removes newline characters (\n) and any extra spaces at the start/end of each line. We also skip empty lines to prevent errors from trying to split an empty string.
  • Data Parsing: split(',') breaks each line into three chunks. Then we use strip().strip("'") to remove both surrounding spaces and single quotes, giving us clean, usable strings ready for your lists.

Optimized Version (For Large Files)

If your zip code file is huge, reading the entire thing into memory with readlines() isn't ideal. Here's a more memory-efficient version that reads line-by-line, plus added error handling for malformed lines:

def main():
    zipcode = []
    city = []
    state = []
    
    with open("zipcodes.txt", "r") as file_object:
        # Iterate directly over the file object (reads one line at a time)
        for line_num, line in enumerate(file_object, 1):
            cleaned_line = line.strip()
            if not cleaned_line:
                continue
            parts = cleaned_line.split(',')
            # Check if the line has exactly 3 parts to avoid index errors
            if len(parts) != 3:
                print(f"Warning: Skipping malformed line {line_num}: {line}")
                continue
            zip_part = parts[0].strip().strip("'")
            city_part = parts[1].strip().strip("'")
            state_part = parts[2].strip().strip("'")
            zipcode.append(zip_part)
            city.append(city_part)
            state.append(state_part)
    
    print(f"Successfully processed {len(zipcode)} valid entries!")
    print("Sample output:")
    for z, c, s in zip(zipcode[:3], city[:3], state[:3]):
        print(f"- {z} | {c}, {s}")

if __name__ == "__main__":
    main()

This version:

  • Uses direct file iteration to avoid loading the entire file into memory (perfect for large datasets)
  • Adds line number tracking to report exactly where malformed lines are, making debugging easier
  • Includes a check to skip lines that don't match the expected 3-part format, preventing unexpected crashes

内容的提问来源于stack exchange,提问作者Nick Nasty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:02:44