部署Django+Scrapy遇404错误,如何让scrapyd后台持续运行?
Hey there, let's work through your Scrapyd issues step by step—first resolving that 404 error with scrapyd-deploy, then setting up Scrapyd to run persistently even after you close the console.
Part 1: Troubleshooting the 404 Error with scrapyd-deploy
First, let's make sure Scrapyd is actually running correctly and accessible:
Verify Scrapyd is listening on port 6800
Run this command to check if Scrapyd is active on the expected port:netstat -tulpn | grep 6800If you don't see a Scrapyd process listed here, your Scrapyd startup is failing. Check the error messages from your console launch—Scrapyd logs are also usually stored in
~/.scrapyd/logs/or/var/log/scrapyd/(depending on your setup) for more details. A common issue on Ubuntu 16.04 is incompatible Python versions: Ubuntu 16.04 ships with Python 3.5, which isn't supported by the latest Scrapyd releases. Try installing a compatible version like:pip install scrapyd==1.2.1Double-check your
scrapy.cfgconfiguration
Ensure your[deploy:default]section looks exactly like this (don't forget the trailing slash in the URL—it can cause 404s):[deploy:default] url = http://your_ip:6800/ project = your_project_nameAlso, confirm your server's firewall (and cloud security group if using a VPS) allows traffic on port 6800. For Ubuntu's UFW firewall, run:
sudo ufw allow 6800Test the Scrapyd API directly
Try accessinghttp://your_ip:6800/in a browser or withcurl:curl http://your_ip:6800/If you see Scrapyd's welcome page, the service is running—your
scrapyd-deployissue might be a typo in your project name or configuration. If you still get a 404, go back to debugging Scrapyd's startup errors.
Part 2: Running Scrapyd in the Background (No Console Required)
Here are three reliable methods, ordered from simplest to most robust:
1. Quick Temporary Fix: nohup
This is the easiest way to run Scrapyd in the background for short-term use:
# Run Scrapyd in background, log output to a file nohup scrapyd > /var/log/scrapyd/scrapyd.log 2>&1 &
- First, create the log directory if it doesn't exist:
sudo mkdir -p /var/log/scrapyd && sudo chown $USER /var/log/scrapyd - To stop Scrapyd, find its process ID and kill it:
ps aux | grep scrapyd kill <your_scrapyd_pid>
2. For Interactive Access: screen
Use this if you need to reconnect to the Scrapyd console later:
- Install screen first:
sudo apt-get install screen - Create a dedicated screen session:
screen -S scrapyd_session - Start Scrapyd inside the session:
scrapyd - Detach from the session (keep Scrapyd running) with
Ctrl+AthenD - Reconnect to the session anytime:
screen -r scrapyd_session - To stop Scrapyd, go back to the session, press
Ctrl+C, then typeexitto close the screen session.
3. Robust Long-Term Solution: Systemd Service
This is the best option for production—Scrapyd will start on boot and automatically restart if it crashes:
- Create a systemd service file:
sudo nano /etc/systemd/system/scrapyd.service - Paste the following configuration (adjust paths to match your setup, especially if using a virtual environment):
[Unit] Description=Scrapyd Crawler Service After=network.target [Service] User=your_username Group=your_username WorkingDirectory=/path/to/your/scrapy/project ExecStart=/usr/local/bin/scrapyd # Use your virtual env path like ~/venv/bin/scrapyd if needed Restart=always RestartSec=5 StandardOutput=journal+console StandardError=journal+console [Install] WantedBy=multi-user.target - Save and exit with
Ctrl+O, press Enter, thenCtrl+X - Reload systemd and start the service:
sudo systemctl daemon-reload sudo systemctl start scrapyd - Enable auto-start on boot:
sudo systemctl enable scrapyd - Check service status or view logs:
sudo systemctl status scrapyd journalctl -u scrapyd -f # Real-time log viewing
内容的提问来源于stack exchange,提问作者Chiefir

