You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Parse Cloud Code抓取网页表格数据为JSON格式求助

Troubleshooting Your Parse Cloud Code Web Scraping Issue

Hey there! I see you're new to JS and trying to scrape table data into JSON using Parse Cloud Code—let's figure out why you're only getting mismatched HTML instead of the target data, and fix this step by step.

First, Let's Break Down the Core Issues

The problems you're hitting are tied to how the target website loads content and how your request is structured:

  • Anchor tags (#) are client-side only: The #k=thinkwater in your URL is a browser-side anchor. When you send a request to this URL, the server ignores everything after the # and sends back the initial page HTML. The table data you want is almost certainly loaded dynamically by JavaScript after the initial page loads—so your raw HTTP request never sees it.
  • Incorrect request setup: You're using a POST request with form-specific parameters, but the initial page load is a GET request. Those parameters aren't necessary here and might be throwing off the server's response.
  • No JavaScript execution: Parse Cloud Code's httpRequest only fetches raw server HTML—it doesn't run the browser JavaScript that renders the dynamic table content you see when visiting the site.

Step-by-Step Fixes

1. Find the Actual Data Source

First, we need to locate where the table data is being pulled from. Here's how:

  • Open the target URL in your browser, right-click > Inspect to open DevTools.
  • Go to the Network tab, then refresh the page.
  • Filter requests by "XHR" (look for the XHR/fetch tab) to find requests that load dynamic data related to "thinkwater". You'll likely see a request that returns the table data (either as JSON or HTML fragments).
  • Copy that request's URL, method (usually GET), and any required headers/parameters.

2. Update Your Parse Cloud Code Request

Once you have the actual data endpoint, modify your code to target it directly. For example, if the data loads via a GET request to a specific API endpoint, your code might look like this:

Parse.Cloud.define('scrapeThinkWaterData', function(req, res) {
  // Replace this URL with the actual data endpoint you found in DevTools
  const dataUrl = 'http://www.capetown.gov.za/ActualDataEndpointHere';
  
  Parse.Cloud.httpRequest({
    method: 'GET',
    url: dataUrl,
    headers: {
      // Add a valid User-Agent to avoid being blocked (matches a typical browser)
      'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
  }).then(function(httpResponse) {
    // If the response is JSON, process it directly
    try {
      const rawData = JSON.parse(httpResponse.text);
      // Map the raw data to your desired JSON format
      const tableJson = rawData.items.map(item => ({
        title: item.Title,
        publishDate: item.PublishDate,
        summary: item.Summary
        // Add other fields you need from the table
      }));
      res.success(tableJson);
    } catch (e) {
      // If the response is HTML, we'll need to parse the table from it
      res.success('Need to parse HTML response—see next section');
    }
  }, function(httpResponse) {
    res.error(`Request failed: ${httpResponse.status} - ${httpResponse.text}`);
  });
});

3. Parsing HTML Responses (If Needed)

If the data endpoint returns HTML instead of JSON, you can use a lightweight parsing library like cheerio (you'll need to add it to your Parse project dependencies):

  • Install cheerio via npm: npm install cheerio
  • Then use it to extract table rows:
const cheerio = require('cheerio');
// ... inside the success callback
const $ = cheerio.load(httpResponse.text);
const tableRows = [];

// Target the specific table (adjust the selector to match the site's table)
$('table.thinkwater-table tr').each((i, rowEl) => {
  const rowData = {};
  $(rowEl).find('td').each((j, cellEl) => {
    // Map cells to meaningful keys (adjust based on table column names)
    const columnNames = ['Title', 'Date', 'Details'];
    rowData[columnNames[j]] = $(cellEl).text().trim();
  });
  tableRows.push(rowData);
});

res.success(tableRows);

Quick Best Practices

  • Check robots.txt: Always verify if the website allows scraping before proceeding.
  • Rate limit requests: Don't flood the server with requests—you might get blocked.
  • Test incrementally: Start by logging the raw response to confirm you're getting the right data before parsing.

内容的提问来源于stack exchange,提问作者BLE

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:13:50