求助:如何合并Web Scraping响应文本脚本与服务器测试脚本
Hey there! Let's break this down step by step since you're already halfway there—great job getting both scripts working independently! Below are practical, actionable ways to merge your server test script and crawler based on common use cases:
First, let's align on the core goal: we want to connect your crawler's HTML output to your DigitalOcean server, either to store the content, serve it to users, or process it further. Here are two straightforward approaches:
Option 1: Direct Function Call (Same Language)
If both scripts use the same language (e.g., Python), this is the simplest method.
Step 1: Refactor your crawler into a reusable function
Extract the core scraping logic from your crawler script so it can be imported and called elsewhere:# crawler.py import requests def fetch_target_html(): target_url = "https://your-target-site.com" # Replace with your URL try: response = requests.get(target_url) response.raise_for_status() # Catch HTTP errors return response.text except Exception as e: print(f"Crawl failed: {str(e)}") return NoneStep 2: Integrate the crawler into your server script
Import the crawler function into your server code, then use it to handle requests or run scheduled tasks. For example, a Flask server that serves crawled content:# server.py from flask import Flask from crawler import fetch_target_html app = Flask(__name__) @app.route('/crawled-content') def serve_crawled_content(): html_content = fetch_target_html() if html_content: # You can also save this to a file/database here instead of returning it return html_content else: return "Failed to fetch content", 500 if __name__ == '__main__': app.run(host='0.0.0.0', port=8080) # Bind to all interfaces for server accessStep 3: Run on DigitalOcean
Install dependencies (e.g.,pip install flask requests), then start the server withpython server.py. Test it by visitinghttp://your-server-ip:8080/crawled-content.
Option 2: File/Database Middleman (Different Languages)
If your scripts use different languages (e.g., Python crawler + Node.js server), use a shared storage layer to pass data:
Step 1: Configure the crawler to save output
Have your crawler write the HTML to a file or lightweight database on your DigitalOcean server. For example:# crawler.py import requests def crawl_and_save(): html = fetch_target_html() if html: with open('/var/crawl-data/latest.html', 'w') as f: f.write(html) print("Content saved successfully")Step 2: Have the server read from the storage
Update your server script to pull content from the shared file/database. Example with Node.js:// server.js const express = require('express'); const fs = require('fs'); const app = express(); app.get('/latest-crawl', (req, res) => { fs.readFile('/var/crawl-data/latest.html', 'utf8', (err, data) => { if (err) return res.status(500).send("No crawled content available"); res.send(data); }); }); app.listen(3000, '0.0.0.0', () => { console.log("Server running on port 3000"); });
Key Tips for DigitalOcean Servers
- Permissions: Make sure your crawler has write access to the target directory (use
sudo chownorchmodto adjust permissions). - Error Handling: Add retries for crawler network failures and server file/database access to avoid crashes.
- Automation: Use
cron(Linux) to schedule your crawler to run at intervals, so your server always has fresh content.
If you can share more details about exactly what you want the merged script to do (e.g., store content in a database, serve an API, trigger crawls via user request), I can refine this further!
内容的提问来源于stack exchange,提问作者Frederik 84

