使用Parse Cloud Code抓取网页表格数据为JSON格式求助
Hey there! I see you're new to JS and trying to scrape table data into JSON using Parse Cloud Code—let's figure out why you're only getting mismatched HTML instead of the target data, and fix this step by step.
First, Let's Break Down the Core Issues
The problems you're hitting are tied to how the target website loads content and how your request is structured:
- Anchor tags (
#) are client-side only: The#k=thinkwaterin your URL is a browser-side anchor. When you send a request to this URL, the server ignores everything after the#and sends back the initial page HTML. The table data you want is almost certainly loaded dynamically by JavaScript after the initial page loads—so your raw HTTP request never sees it. - Incorrect request setup: You're using a
POSTrequest with form-specific parameters, but the initial page load is aGETrequest. Those parameters aren't necessary here and might be throwing off the server's response. - No JavaScript execution: Parse Cloud Code's
httpRequestonly fetches raw server HTML—it doesn't run the browser JavaScript that renders the dynamic table content you see when visiting the site.
Step-by-Step Fixes
1. Find the Actual Data Source
First, we need to locate where the table data is being pulled from. Here's how:
- Open the target URL in your browser, right-click > Inspect to open DevTools.
- Go to the Network tab, then refresh the page.
- Filter requests by "XHR" (look for the XHR/fetch tab) to find requests that load dynamic data related to "thinkwater". You'll likely see a request that returns the table data (either as JSON or HTML fragments).
- Copy that request's URL, method (usually
GET), and any required headers/parameters.
2. Update Your Parse Cloud Code Request
Once you have the actual data endpoint, modify your code to target it directly. For example, if the data loads via a GET request to a specific API endpoint, your code might look like this:
Parse.Cloud.define('scrapeThinkWaterData', function(req, res) { // Replace this URL with the actual data endpoint you found in DevTools const dataUrl = 'http://www.capetown.gov.za/ActualDataEndpointHere'; Parse.Cloud.httpRequest({ method: 'GET', url: dataUrl, headers: { // Add a valid User-Agent to avoid being blocked (matches a typical browser) 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } }).then(function(httpResponse) { // If the response is JSON, process it directly try { const rawData = JSON.parse(httpResponse.text); // Map the raw data to your desired JSON format const tableJson = rawData.items.map(item => ({ title: item.Title, publishDate: item.PublishDate, summary: item.Summary // Add other fields you need from the table })); res.success(tableJson); } catch (e) { // If the response is HTML, we'll need to parse the table from it res.success('Need to parse HTML response—see next section'); } }, function(httpResponse) { res.error(`Request failed: ${httpResponse.status} - ${httpResponse.text}`); }); });
3. Parsing HTML Responses (If Needed)
If the data endpoint returns HTML instead of JSON, you can use a lightweight parsing library like cheerio (you'll need to add it to your Parse project dependencies):
- Install
cheeriovia npm:npm install cheerio - Then use it to extract table rows:
const cheerio = require('cheerio'); // ... inside the success callback const $ = cheerio.load(httpResponse.text); const tableRows = []; // Target the specific table (adjust the selector to match the site's table) $('table.thinkwater-table tr').each((i, rowEl) => { const rowData = {}; $(rowEl).find('td').each((j, cellEl) => { // Map cells to meaningful keys (adjust based on table column names) const columnNames = ['Title', 'Date', 'Details']; rowData[columnNames[j]] = $(cellEl).text().trim(); }); tableRows.push(rowData); }); res.success(tableRows);
Quick Best Practices
- Check robots.txt: Always verify if the website allows scraping before proceeding.
- Rate limit requests: Don't flood the server with requests—you might get blocked.
- Test incrementally: Start by logging the raw response to confirm you're getting the right data before parsing.
内容的提问来源于stack exchange,提问作者BLE

