基于.NET Core的双Web应用:爬虫应用触发更新数据库方案咨询
Hey there, let's walk through the best approach to build your .NET Core solution based on your requirements. Since you need two separate apps—one for web scraping and database writes, another for data processing with a trigger to kick off the scraper—here's a structured, production-ready plan:
First, let's clarify what each app should be to meet your needs:
1.1 Web Scraper App: ASP.NET Core Web API
This app needs to be triggerable via a button click, so a Web API is the most straightforward choice (it natively supports HTTP requests from your second app). Alternatively, you could use a Worker Service with a message queue, but the Web API is simpler for direct triggering.
Core features to implement:
- Expose a
POST /api/scraper/startendpoint to receive trigger requests. - Use a background task system (like
IHostedService/BackgroundService) to run scraping jobs asynchronously—this prevents the API request from timing out during long-running scrapes. - Integrate with a web scraping library (HtmlAgilityPack or AngleSharp are both great for .NET Core) to parse HTML from target sites.
- Use Entity Framework Core (EF Core) to write scraped data to your database—keep your data models clean and reusable.
1.2 Data Processing App: ASP.NET Core Web App (MVC/Razor Pages)
For a user-facing app with a trigger button, go with either MVC or Razor Pages (Razor Pages is simpler for small to medium UIs).
Core features to implement:
- A UI with a button that, when clicked, sends a request to the scraper API's trigger endpoint.
- Logic to read raw data from the database, process it (cleaning, aggregation, etc.), and display results to users.
- Error handling to notify users if the scraper fails to start.
The button in your processing app needs to kick off the scraper—here are your best options:
2.1 Direct HTTP Call (Simplest)
Use HttpClient in your processing app to send a POST request to the scraper's /api/scraper/start endpoint. Key considerations:
- Store the scraper API's base URL in
appsettings.jsonfor easy configuration. - Add retry logic (using
Pollylibrary, for example) to handle transient failures like network blips. - Secure the endpoint: Add API key authentication or JWT tokens to ensure only your processing app can trigger the scraper. For example, add an
[Authorize]attribute to the scraper endpoint and validate an API key in the request header.
2.2 Message Queue (More Reliable for Long Jobs)
If your scrapes take minutes to complete, a message queue (like RabbitMQ or Azure Service Bus) decouples the two apps:
- The processing app sends a "start scrape" message to the queue when the button is clicked.
- The scraper app (configured as a Worker Service or Web API with a background listener) picks up the message and runs the scrape.
- This avoids HTTP timeouts and lets you track job status more easily.
Since both apps need to access the same data:
- Use a shared class library for your EF Core data models and DbContext. Both apps reference this library to ensure consistent data schemas—no duplicate model definitions!
- Configure connection strings in each app's
appsettings.json(pointing to the same database instance). - Handle concurrency: If the scraper is writing data while the processing app is reading/updating it, use EF Core's optimistic concurrency (add a
Versionproperty to your models) to avoid conflicts.
Let's look at some key code pieces to make this concrete:
Scraper API: Trigger Endpoint with Background Tasks
[ApiController] [Route("api/scraper")] [Authorize(AuthenticationSchemes = "ApiKey")] // Secure the endpoint public class ScraperController : ControllerBase { private readonly IBackgroundTaskQueue _taskQueue; private readonly ILogger<ScraperController> _logger; public ScraperController(IBackgroundTaskQueue taskQueue, ILogger<ScraperController> logger) { _taskQueue = taskQueue; _logger = logger; } [HttpPost("start")] public IActionResult StartScrape() { _taskQueue.QueueBackgroundWorkItem(async token => { _logger.LogInformation("Starting web scrape job"); // Replace with your actual scraping logic var scraper = new MyWebScraper(); var scrapedData = await scraper.ScrapeSitesAsync(token); // Write to database using EF Core using var dbContext = new MyDbContext(); await dbContext.ScrapedData.AddRangeAsync(scrapedData); await dbContext.SaveChangesAsync(token); _logger.LogInformation("Scrape job completed successfully"); }); return Accepted("Scrape job has been queued and will run in the background"); } } // Background task queue implementation public interface IBackgroundTaskQueue { void QueueBackgroundWorkItem(Func<CancellationToken, Task> workItem); Task<Func<CancellationToken, Task>> DequeueAsync(CancellationToken cancellationToken); } public class BackgroundTaskQueue : IBackgroundTaskQueue { private readonly ConcurrentQueue<Func<CancellationToken, Task>> _workItems = new(); private readonly SemaphoreSlim _signal = new(0); public void QueueBackgroundWorkItem(Func<CancellationToken, Task> workItem) { if (workItem == null) throw new ArgumentNullException(nameof(workItem)); _workItems.Enqueue(workItem); _signal.Release(); } public async Task<Func<CancellationToken, Task>> DequeueAsync(CancellationToken cancellationToken) { await _signal.WaitAsync(cancellationToken); _workItems.TryDequeue(out var workItem); return workItem; } } // Background service to process queued tasks public class ScraperBackgroundService : BackgroundService { private readonly IBackgroundTaskQueue _taskQueue; private readonly ILogger<ScraperBackgroundService> _logger; public ScraperBackgroundService(IBackgroundTaskQueue taskQueue, ILogger<ScraperBackgroundService> logger) { _taskQueue = taskQueue; _logger = logger; } protected override async Task ExecuteAsync(CancellationToken stoppingToken) { _logger.LogInformation("Scraper background service started"); while (!stoppingToken.IsCancellationRequested) { var workItem = await _taskQueue.DequeueAsync(stoppingToken); try { await workItem(stoppingToken); } catch (Exception ex) { _logger.LogError(ex, "Error executing scrape task"); } } } }
Processing App: Button Trigger with HttpClient
(Razor Pages example):
public class IndexModel : PageModel { private readonly HttpClient _httpClient; private readonly IConfiguration _config; public IndexModel(HttpClient httpClient, IConfiguration config) { _httpClient = httpClient; _config = config; } public string? StatusMessage { get; set; } public async Task<IActionResult> OnPostStartScraperAsync() { try { var scraperApiUrl = $"{_config["ScraperApi:BaseUrl"]}/api/scraper/start"; // Add API key to request header for authentication _httpClient.DefaultRequestHeaders.Add("X-Api-Key", _config["ScraperApi:ApiKey"]); var response = await _httpClient.PostAsync(scraperApiUrl, null); response.EnsureSuccessStatusCode(); StatusMessage = "Scraper started successfully! Data will be updated shortly."; } catch (HttpRequestException ex) { StatusMessage = $"Failed to start scraper: {ex.Message}"; } return Page(); } }
And the Razor Page markup:
@page @model IndexModel <h1>Data Processing Dashboard</h1> @if (!string.IsNullOrEmpty(Model.StatusMessage)) { <div class="alert @(Model.StatusMessage.Contains("Failed") ? "alert-danger" : "alert-success")"> @Model.StatusMessage </div> } <form method="post" asp-page-handler="StartScraper"> <button type="submit" class="btn btn-primary">Update Data via Web Scraper</button> </form> <!-- Add your data processing/display logic below -->
- Deployment: Both apps can be deployed independently to IIS, Docker containers, or cloud services like Azure App Service. Use Docker Compose if you want to deploy them together locally or in a containerized environment.
- Logging: Add structured logging (with Serilog or NLog) to both apps to track scrape jobs, errors, and data processing steps.
- Monitoring: Use tools like Application Insights to monitor app performance, track scrape job duration, and receive alerts for failures.
内容的提问来源于stack exchange,提问作者Dov95

