You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何构造Elasticsearch分页Payload以通过Curl批量提取Bubble.io插件页面数据?

Can I use a Bash + Curl script to scrape bubble.io/plugins via their Elasticsearch API?

Absolutely! This is totally feasible—you just need to replicate the exact API requests that the bubble.io plugins page makes to its Elasticsearch backend, including the pagination payload parameters. Here's a step-by-step breakdown to pull this off:

Step 1: Capture full request details with browser dev tools

First, you need to get all the specifics of the API call the page makes:

  • Open the bubble.io/plugins page in Chrome/Firefox.
  • Launch DevTools (F12), switch to the Network tab.
  • Refresh the page, then look for requests to an Elasticsearch endpoint (usually something like /search or with "elasticsearch" in the URL).
  • Click on that request to view:
    • The Request URL (the actual Elasticsearch endpoint bubble is hitting).
    • Headers: Pay close attention to Authorization, Content-Type, Referer, and any custom headers—you’ll need these in your curl command to avoid being blocked.
    • Payload: The JSON body the page sends. Your x, y, z parameters are likely Elasticsearch pagination tools like from (starting index), size (results per page), or search_after (cursor-based pagination to avoid duplicates).

Step 2: Test a single curl request

Once you have all the details, test a single curl command to confirm it works. For example, if the payload uses from and size:

curl -X POST "https://your-captured-elasticsearch-endpoint.com/search" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer your-token-here" \
  -H "Referer: https://bubble.io/plugins" \
  -d '{
    "query": { /* copy bubble''s exact query here */ },
    "from": 0,
    "size": 20,
    "sort": [{"plugin_name": "asc"}] /* match the sort from the original payload */
  }'

Replace placeholders with values you captured. If this returns valid JSON results, you’re ready to scale.

Step 3: Build a Bash script for pagination

Wrap the curl command in a loop to fetch all pages. The approach depends on the pagination type:

Case 1: Offset-based pagination (using from and size)

If the payload uses from to increment the starting index, loop until the response returns fewer results than size:

#!/bin/bash

# Configuration
ENDPOINT="https://your-captured-endpoint.com/search"
HEADERS=(
  "Content-Type: application/json"
  "Authorization: Bearer your-token"
  "Referer: https://bubble.io/plugins"
)
PAGE_SIZE=20
FROM=0
OUTPUT_FILE="bubble_plugins.json"

# Initialize output file
echo "[" > "$OUTPUT_FILE"

while true; do
  echo "Fetching page starting at index $FROM..."
  
  # Send curl request and capture response
  RESPONSE=$(curl -s -X POST "$ENDPOINT" \
    "${HEADERS[@]/#/-H }" \
    -d '{
      "query": { /* paste bubble''s query here */ },
      "from": '"$FROM"',
      "size": '"$PAGE_SIZE"'
    }')
  
  # Get number of results returned
  RESULTS_COUNT=$(echo "$RESPONSE" | jq '.hits.total.value')
  
  if [ "$RESULTS_COUNT" -eq 0 ]; then
    echo "No more results. Exiting loop."
    break
  fi
  
  # Append results to output file (clean up formatting)
  echo "$RESPONSE" | jq '.hits.hits' | sed '1d;$d' >> "$OUTPUT_FILE"
  echo "," >> "$OUTPUT_FILE"
  
  # Increment starting index
  FROM=$((FROM + PAGE_SIZE))
  
  # Add delay to avoid rate limiting
  sleep 2
done

# Clean up output file (remove trailing comma and close array)
sed -i '$ s/,$//' "$OUTPUT_FILE"
echo "]" >> "$OUTPUT_FILE"

echo "All data saved to $OUTPUT_FILE"

Case 2: Cursor-based pagination (using search_after)

If the payload uses search_after (common for large datasets), extract the last document’s sort values from each response to use in the next request:

#!/bin/bash

ENDPOINT="https://your-captured-endpoint.com/search"
HEADERS=(...)
PAGE_SIZE=20
SEARCH_AFTER=""
OUTPUT_FILE="bubble_plugins.json"

echo "[" > "$OUTPUT_FILE"

while true; do
  PAYLOAD='{
    "query": { /* paste bubble''s query here */ },
    "size": '"$PAGE_SIZE"',
    "sort": [{"plugin_name": "asc"}]
  }'
  
  # Add search_after to payload if it exists
  if [ -n "$SEARCH_AFTER" ]; then
    PAYLOAD=$(echo "$PAYLOAD" | jq --argjson sa "$SEARCH_AFTER" '. + {"search_after": $sa}')
  fi
  
  RESPONSE=$(curl -s -X POST "$ENDPOINT" "${HEADERS[@]/#/-H }" -d "$PAYLOAD")
  
  RESULTS=$(echo "$RESPONSE" | jq '.hits.hits')
  RESULTS_COUNT=$(echo "$RESULTS" | jq 'length')
  
  if [ "$RESULTS_COUNT" -eq 0 ]; then
    break
  fi
  
  # Append results to output
  echo "$RESULTS" | sed '1d;$d' >> "$OUTPUT_FILE"
  echo "," >> "$OUTPUT_FILE"
  
  # Get last document's sort values for next request
  SEARCH_AFTER=$(echo "$RESPONSE" | jq '.hits.hits[-1].sort')
  
  sleep 2
done

# Clean up output file
sed -i '$ s/,$//' "$OUTPUT_FILE"
echo "]" >> "$OUTPUT_FILE"

Important Notes

  • Rate Limiting: Bubble might block your requests if you hit their API too quickly. Add a sleep between requests (2-5 seconds) to avoid this.
  • API Changes: Bubble could modify their payload or endpoint at any time. If your script breaks, re-inspect the request in dev tools to update parameters.
  • Authentication: If the API uses an auth token, ensure it’s included in headers. Some tokens expire, so you might need to refresh them periodically.
  • jq Dependency: The scripts use jq to parse JSON. Install it with sudo apt install jq (Debian/Ubuntu) or brew install jq (macOS).

内容的提问来源于stack exchange,提问作者Santosh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 01:12:48