PyShp拆分Shapefile无记录且缺失DBF文件问题求助
Fixing Your PyShp Shapefile Splitting Issue
Hey, I see exactly what's causing your broken Shapefiles—let's break down the problems in your code and fix them step by step:
Key Issues in Your Current Code
- You're not creating valid Shapefile structures: Right now, you're just writing attribute records as plain text to a file. Shapefiles require three core files (
.shp,.shx,.dbf) formatted correctly, which PyShp handles via itsWriterclass—not manual text writing. - You're ignoring geometric data: Your code only extracts attribute records, but Shapefiles rely on matching geometry + attribute pairs. Skipping the geometry means your output
.shpwill have no valid spatial data. - PRJ file naming mismatch: You're copying the PRJ file using the original filename instead of your split part's name, so GIS tools won't associate it with the new Shapefile.
- Typos and inconsistent paths: There's a typo (
file_pathvsfilepathin your function) that could cause unexpected errors.
Corrected Code
Here's the fixed version of your function, with comments explaining the critical changes:
import os import math import zipfile import shutil import shapefile path = '<shapefile_data_path>' storage_path = '<path_to_extract_zip_file>' current_dir = '<path_for_divided_shapefiles>' ALLOWED_SIZE = 10 # MB def split_large_shapefile(filepath): # Fix typo: use filepath consistently file_name = filepath.split('/')[-1] name = file_name.split('.zip')[0] storage_file = os.path.join(storage_path, file_name).replace('\\', '/') src = os.path.join(path, filepath) # Copy and check zip size shutil.copy(src, storage_file) statinfo = os.stat(storage_file) if (statinfo.st_size >> 20) <= ALLOWED_SIZE: print(f"{file_name} is under size limit, no split needed.") return # Extract zip contents extract_dir = os.path.join(storage_path, name) os.makedirs(extract_dir, exist_ok=True) with zipfile.ZipFile(storage_file, 'r') as zip_ref: zip_ref.extractall(extract_dir) # Find PRJ and SHP files prj_file = None shp_file = None for f in os.listdir(extract_dir): if f.endswith('.prj'): prj_file = os.path.join(extract_dir, f) if f.endswith('.shp'): shp_file = os.path.join(extract_dir, f) if not shp_file: print("No SHP file found in extracted contents!") return # Load original Shapefile reader = shapefile.Reader(shp_file) num_features = len(reader) if num_features == 0: print("Original Shapefile has no features!") return # Calculate number of parts # Note: You might want to adjust this logic to use uncompressed SHP size instead of zip size zip_size_mb = statinfo.st_size >> 20 parts = math.ceil(zip_size_mb / ALLOWED_SIZE) increment = math.ceil(num_features / parts) # Use ceil to avoid missing features # Get original metadata for new files field_names = [field[0] for field in reader.fields[1:]] # Skip first dummy field field_types = [field[1] for field in reader.fields[1:]] field_lengths = [field[2] for field in reader.fields[1:]] shape_type = reader.shapeType # Split and write each part start_idx = 0 part_num = 1 while start_idx < num_features: end_idx = min(start_idx + increment, num_features) part_name = f"{name}_part{part_num}" part_dir = os.path.join(current_dir, part_name) os.makedirs(part_dir, exist_ok=True) part_shp_path = os.path.join(part_dir, part_name) # Initialize PyShp Writer with correct shape type and fields writer = shapefile.Writer(part_shp_path, shapeType=shape_type) # Add fields matching original Shapefile for name, typ, length in zip(field_names, field_types, field_lengths): writer.field(name, typ, length=length) # Write each feature (geometry + attributes) for idx in range(start_idx, end_idx): shape = reader.shape(idx) record = reader.record(idx) writer.shape(shape) writer.record(*record) # Save the new Shapefile (creates .shp, .shx, .dbf automatically) writer.close() # Copy PRJ file with matching name if prj_file: shutil.copy(prj_file, os.path.join(part_dir, f"{part_name}.prj")) # Zip the part and clean up shutil.make_archive(part_dir, 'zip', part_dir) shutil.rmtree(part_dir) start_idx = end_idx part_num += 1 # Example usage: replace with your actual zip file path # split_large_shapefile("your_large_shapefile.zip")
Critical Changes Explained
- Using
shapefile.Writer: This class handles creating all three required Shapefile components (.shp,.shx,.dbf) in the correct format, so you don't have to manually manage files. - Writing geometry + attributes: For each feature in the split range, we extract both the geometric shape and its associated attribute record, then write both to the new file.
- Matching field definitions: We copy the field names, types, and lengths from the original Shapefile to ensure the output DBF has the correct structure.
- Fixing PRJ naming: The PRJ file is renamed to match the split part's filename, so GIS software recognizes it as the coordinate system for the new Shapefile.
- Handling edge cases: Added checks for missing SHP files, empty original files, and adjusted the increment calculation to avoid missing features.
内容的提问来源于stack exchange,提问作者Venkatesh_CTA
相关产品推荐
相关产品推荐

