You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyShp拆分Shapefile无记录且缺失DBF文件问题求助

Fixing Your PyShp Shapefile Splitting Issue

Hey, I see exactly what's causing your broken Shapefiles—let's break down the problems in your code and fix them step by step:

Key Issues in Your Current Code

  • You're not creating valid Shapefile structures: Right now, you're just writing attribute records as plain text to a file. Shapefiles require three core files (.shp, .shx, .dbf) formatted correctly, which PyShp handles via its Writer class—not manual text writing.
  • You're ignoring geometric data: Your code only extracts attribute records, but Shapefiles rely on matching geometry + attribute pairs. Skipping the geometry means your output .shp will have no valid spatial data.
  • PRJ file naming mismatch: You're copying the PRJ file using the original filename instead of your split part's name, so GIS tools won't associate it with the new Shapefile.
  • Typos and inconsistent paths: There's a typo (file_path vs filepath in your function) that could cause unexpected errors.

Corrected Code

Here's the fixed version of your function, with comments explaining the critical changes:

import os
import math
import zipfile
import shutil
import shapefile

path = '<shapefile_data_path>'
storage_path = '<path_to_extract_zip_file>'
current_dir = '<path_for_divided_shapefiles>'
ALLOWED_SIZE = 10  # MB

def split_large_shapefile(filepath):
    # Fix typo: use filepath consistently
    file_name = filepath.split('/')[-1]
    name = file_name.split('.zip')[0]
    storage_file = os.path.join(storage_path, file_name).replace('\\', '/')
    src = os.path.join(path, filepath)
    
    # Copy and check zip size
    shutil.copy(src, storage_file)
    statinfo = os.stat(storage_file)
    if (statinfo.st_size >> 20) <= ALLOWED_SIZE:
        print(f"{file_name} is under size limit, no split needed.")
        return
    
    # Extract zip contents
    extract_dir = os.path.join(storage_path, name)
    os.makedirs(extract_dir, exist_ok=True)
    with zipfile.ZipFile(storage_file, 'r') as zip_ref:
        zip_ref.extractall(extract_dir)
    
    # Find PRJ and SHP files
    prj_file = None
    shp_file = None
    for f in os.listdir(extract_dir):
        if f.endswith('.prj'):
            prj_file = os.path.join(extract_dir, f)
        if f.endswith('.shp'):
            shp_file = os.path.join(extract_dir, f)
    
    if not shp_file:
        print("No SHP file found in extracted contents!")
        return
    
    # Load original Shapefile
    reader = shapefile.Reader(shp_file)
    num_features = len(reader)
    if num_features == 0:
        print("Original Shapefile has no features!")
        return
    
    # Calculate number of parts
    # Note: You might want to adjust this logic to use uncompressed SHP size instead of zip size
    zip_size_mb = statinfo.st_size >> 20
    parts = math.ceil(zip_size_mb / ALLOWED_SIZE)
    increment = math.ceil(num_features / parts)  # Use ceil to avoid missing features
    
    # Get original metadata for new files
    field_names = [field[0] for field in reader.fields[1:]]  # Skip first dummy field
    field_types = [field[1] for field in reader.fields[1:]]
    field_lengths = [field[2] for field in reader.fields[1:]]
    shape_type = reader.shapeType
    
    # Split and write each part
    start_idx = 0
    part_num = 1
    while start_idx < num_features:
        end_idx = min(start_idx + increment, num_features)
        part_name = f"{name}_part{part_num}"
        part_dir = os.path.join(current_dir, part_name)
        os.makedirs(part_dir, exist_ok=True)
        part_shp_path = os.path.join(part_dir, part_name)
        
        # Initialize PyShp Writer with correct shape type and fields
        writer = shapefile.Writer(part_shp_path, shapeType=shape_type)
        # Add fields matching original Shapefile
        for name, typ, length in zip(field_names, field_types, field_lengths):
            writer.field(name, typ, length=length)
        
        # Write each feature (geometry + attributes)
        for idx in range(start_idx, end_idx):
            shape = reader.shape(idx)
            record = reader.record(idx)
            writer.shape(shape)
            writer.record(*record)
        
        # Save the new Shapefile (creates .shp, .shx, .dbf automatically)
        writer.close()
        
        # Copy PRJ file with matching name
        if prj_file:
            shutil.copy(prj_file, os.path.join(part_dir, f"{part_name}.prj"))
        
        # Zip the part and clean up
        shutil.make_archive(part_dir, 'zip', part_dir)
        shutil.rmtree(part_dir)
        
        start_idx = end_idx
        part_num += 1

# Example usage: replace with your actual zip file path
# split_large_shapefile("your_large_shapefile.zip")

Critical Changes Explained

  1. Using shapefile.Writer: This class handles creating all three required Shapefile components (.shp, .shx, .dbf) in the correct format, so you don't have to manually manage files.
  2. Writing geometry + attributes: For each feature in the split range, we extract both the geometric shape and its associated attribute record, then write both to the new file.
  3. Matching field definitions: We copy the field names, types, and lengths from the original Shapefile to ensure the output DBF has the correct structure.
  4. Fixing PRJ naming: The PRJ file is renamed to match the split part's filename, so GIS software recognizes it as the coordinate system for the new Shapefile.
  5. Handling edge cases: Added checks for missing SHP files, empty original files, and adjusted the increment calculation to avoid missing features.

内容的提问来源于stack exchange,提问作者Venkatesh_CTA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:36:46