Scraping Hub定时脚本调用JSON文件触发IOError:文件不存在求助
I’ve run into similar issues with Scraping Hub’s resource handling before—let’s break down how to resolve this IOError and get your JSON file loading correctly.
1. Verify Your Project Structure First
Scraping Hub relies on proper Python package structure to recognize package_data. Make sure your project follows this structure:
your-project-root/ setup.py project/ __init__.py # This is critical—it marks the folder as a Python package resources/ bmibmi-67d3f1f00f49.json your_script.py # Your timed script/spider file
If the __init__.py is missing in the project folder, Scraping Hub won’t treat it as a valid package, so your package_data configuration won’t take effect.
2. Adjust Setup Configuration for Reliable Resource Deployment
While package_data works in local environments, using include_package_data=True with a MANIFEST.in file is more robust for Scraping Hub’s deployment pipeline:
Step 2.1: Create a MANIFEST.in File
In your project root, add a MANIFEST.in file with this line:
include project/resources/*.json
This explicitly tells the packaging tool to include all JSON files in the resources directory during deployment.
Step 2.2: Update setup.py
Modify your setup.py to use include_package_data=True (you can keep package_data as a fallback, but it’s often redundant with MANIFEST.in):
from setuptools import setup, find_packages setup( name='your-project-name', version='1.0', packages=find_packages(), include_package_data=True, # Optional: Keep for extra redundancy package_data={ 'project': ['resources/*.json'], }, )
3. Load the JSON File Correctly (Don’t Hardcode Paths!)
The root cause of your error is likely hardcoding the file path—Scraping Hub’s runtime working directory doesn’t match your local environment. Use Python’s built-in resource utilities to get the correct path dynamically:
Option A: For Python 3.7+ (Modern Approach)
Use importlib.resources to safely locate the resource:
import json from importlib.resources import files # Get the path to your JSON file json_file = files('project.resources').joinpath('bmibmi-67d3f1f00f49.json') # Load the file content with open(json_file, 'r') as f: json_data = json.load(f)
Option B: Compatible with Older Python Versions
Use pkg_resources (part of setuptools):
import json import pkg_resources # Get the absolute path to the JSON file json_path = pkg_resources.resource_filename('project', 'resources/bmibmi-67d3f1f00f49.json') # Load the file with open(json_path, 'r') as f: json_data = json.load(f)
4. Redeploy and Validate
After making these changes:
- Run
shub deployto push your updated code and resources to Scraping Hub. - In your Scraping Hub project dashboard, go to Files to confirm the
resourcesfolder and JSON file are present in the deployed code. - Test your timed script again—the file not found error should be resolved.
内容的提问来源于stack exchange,提问作者nicolasdavid

