You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scraping Hub定时脚本调用JSON文件触发IOError:文件不存在求助

Fixing "No such file or directory" for JSON Resources in Scraping Hub

I’ve run into similar issues with Scraping Hub’s resource handling before—let’s break down how to resolve this IOError and get your JSON file loading correctly.

1. Verify Your Project Structure First

Scraping Hub relies on proper Python package structure to recognize package_data. Make sure your project follows this structure:

your-project-root/
  setup.py
  project/
    __init__.py  # This is critical—it marks the folder as a Python package
    resources/
      bmibmi-67d3f1f00f49.json
    your_script.py  # Your timed script/spider file

If the __init__.py is missing in the project folder, Scraping Hub won’t treat it as a valid package, so your package_data configuration won’t take effect.

2. Adjust Setup Configuration for Reliable Resource Deployment

While package_data works in local environments, using include_package_data=True with a MANIFEST.in file is more robust for Scraping Hub’s deployment pipeline:

Step 2.1: Create a MANIFEST.in File

In your project root, add a MANIFEST.in file with this line:

include project/resources/*.json

This explicitly tells the packaging tool to include all JSON files in the resources directory during deployment.

Step 2.2: Update setup.py

Modify your setup.py to use include_package_data=True (you can keep package_data as a fallback, but it’s often redundant with MANIFEST.in):

from setuptools import setup, find_packages

setup(
    name='your-project-name',
    version='1.0',
    packages=find_packages(),
    include_package_data=True,
    # Optional: Keep for extra redundancy
    package_data={
        'project': ['resources/*.json'],
    },
)

3. Load the JSON File Correctly (Don’t Hardcode Paths!)

The root cause of your error is likely hardcoding the file path—Scraping Hub’s runtime working directory doesn’t match your local environment. Use Python’s built-in resource utilities to get the correct path dynamically:

Option A: For Python 3.7+ (Modern Approach)

Use importlib.resources to safely locate the resource:

import json
from importlib.resources import files

# Get the path to your JSON file
json_file = files('project.resources').joinpath('bmibmi-67d3f1f00f49.json')

# Load the file content
with open(json_file, 'r') as f:
    json_data = json.load(f)

Option B: Compatible with Older Python Versions

Use pkg_resources (part of setuptools):

import json
import pkg_resources

# Get the absolute path to the JSON file
json_path = pkg_resources.resource_filename('project', 'resources/bmibmi-67d3f1f00f49.json')

# Load the file
with open(json_path, 'r') as f:
    json_data = json.load(f)

4. Redeploy and Validate

After making these changes:

  • Run shub deploy to push your updated code and resources to Scraping Hub.
  • In your Scraping Hub project dashboard, go to Files to confirm the resources folder and JSON file are present in the deployed code.
  • Test your timed script again—the file not found error should be resolved.

内容的提问来源于stack exchange,提问作者nicolasdavid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:13:13