You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python 3抓取URL不变的多页面数据?

抓取香港中原地产交易数据的分页问题

目标页面:香港中原地产英文站的交易列表页,需获取全部417页数据,但点击下一页时URL无变化,无法通过URL参数直接分页。当前需先实现获取第一、第二页数据。

已尝试两种方法均未成功:

  • 用Beautiful Soup + Requests抓取页面元素:仅能获取第一页内容,无法触发分页加载
  • 调用Network面板中找到的API接口:无法正确获取分页数据,推测是API参数传递方式错误

已尝试的代码

方法1:Beautiful Soup + Requests

import requests 
import pandas as pd
import re
from bs4 import BeautifulSoup
import json

web_url = 'https://hk.centanet.com/findproperty/en/list/transaction?gclid=Cj0KCQjwnbmaBhD-ARIsAGTPcfVae1prjf_9aKh0dbnaBbzYvi3VhKn4qEXDAQJMS6ZvOiet8GLqzaAaAqH_EALw_wcB&q=3qoOuFNgwUeioKQCtZ9KFA'
data = {"name": "Value"}

r = requests.get(web_url)

print(f'status: {r.status_code}')

if r.status_code == 200:    
    raw_data = BeautifulSoup(r.content, 'html.parser')
    
table_content = raw_data.find_all('div', {'class':'cv-structured-list-item cv-structured-list-item--standard bx--structured-list-row'})
print(len(table_content))

方法2:调用API接口

import requests 
import pandas as pd
import re
from bs4 import BeautifulSoup
import json

web_url = 'https://hk.centanet.com/findproperty/api/Transaction/Search'

_header = {'lang':'en',"postType":"Both","day":"Day1095","sort":"InsOrRegDate","order":"Descending","size":"24","offset":"24","pageSource":"search","gclid":"Cj0KCQjwnbmaBhD-ARIsAGTPcfVae1prjf_9aKh0dbnaBbzYvi3VhKn4qEXDAQJMS6ZvOiet8GLqzaAaAqH_EALw_wcB","q":"3qoOuFNgwUeioKQCtZ9KFA"}
_data = {"name":"Value"}

req = requests.post(web_url, headers=_header, json=_data)

print(f'Status Code: {req.status_code}')

if req.status_code==200:    
    data = json.loads(req.content)['data']

print(len(data))
print(data[0].get('displayText').get('addr'))

解决方案

问题分析

你调用的API接口分页逻辑是通过offset和size参数控制的,但请求存在两个错误:

  1. 分页参数被错误放在了请求头里,实际需放在请求Body中
  2. 请求Body内容不完整,缺少正确的筛选参数

修正后的API调用代码(获取前两页)

import requests
import json

api_url = 'https://hk.centanet.com/findproperty/api/Transaction/Search'

# 模拟浏览器请求头,避免被拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36',
    'Referer': 'https://hk.centanet.com/findproperty/en/list/transaction?gclid=Cj0KCQjwnbmaBhD-ARIsAGTPcfVae1prjf_9aKh0dbnaBbzYvi3VhKn4qEXDAQJMS6ZvOiet8GLqzaAaAqH_EALw_wcB&q=3qoOuFNgwUeioKQCtZ9KFA',
    'lang': 'en'
}

# 第一页参数:offset=0,取前24条
page1_payload = {
    "postType": "Both",
    "day": "Day1095",
    "sort": "InsOrRegDate",
    "order": "Descending",
    "size": 24,
    "offset": 0,
    "pageSource": "search",
    "gclid": "Cj0KCQjwnbmaBhD-ARIsAGTPcfVae1prjf_9aKh0dbnaBbzYvi3VhKn4qEXDAQJMS6ZvOiet8GLqzaAaAqH_EALw_wcB",
    "q": "3qoOuFNgwUeioKQCtZ9KFA"
}

response1 = requests.post(api_url, headers=headers, json=page1_payload)
if response1.status_code == 200:
    page1_data = response1.json()['data']
    print(f"第一页数据条数: {len(page1_data)}")
    print(f"第一页第一条地址: {page1_data[0]['displayText']['addr']}")

# 第二页参数:offset=24,取第25-48条
page2_payload = page1_payload.copy()
page2_payload['offset'] = 24

response2 = requests.post(api_url, headers=headers, json=page2_payload)
if response2.status_code == 200:
    page2_data = response2.json()['data']
    print(f"\n第二页数据条数: {len(page2_data)}")
    print(f"第二页第一条地址: {page2_data[0]['displayText']['addr']}")

批量获取所有页面的思路

  1. 先调用一次API,从返回结果的total字段获取总数据条数
  2. 计算总页数:total_pages = (总条数 + 每页条数 - 1) // 每页条数
  3. 循环遍历每一页,修改offset参数(offset = 页码索引 × 每页条数),发送请求并保存数据

注意事项

  • 添加time.sleep(1)控制请求频率,避免被网站封禁
  • 若请求失败,可先访问一次页面获取Cookie后再调用API

内容的提问来源于stack exchange,提问作者humphrey chan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 01:31:25