如何用Python 3抓取URL不变的多页面数据?
抓取香港中原地产交易数据的分页问题
目标页面:香港中原地产英文站的交易列表页,需获取全部417页数据,但点击下一页时URL无变化,无法通过URL参数直接分页。当前需先实现获取第一、第二页数据。
已尝试两种方法均未成功:
- 用Beautiful Soup + Requests抓取页面元素:仅能获取第一页内容,无法触发分页加载
- 调用Network面板中找到的API接口:无法正确获取分页数据,推测是API参数传递方式错误
已尝试的代码
方法1:Beautiful Soup + Requests
import requests import pandas as pd import re from bs4 import BeautifulSoup import json web_url = 'https://hk.centanet.com/findproperty/en/list/transaction?gclid=Cj0KCQjwnbmaBhD-ARIsAGTPcfVae1prjf_9aKh0dbnaBbzYvi3VhKn4qEXDAQJMS6ZvOiet8GLqzaAaAqH_EALw_wcB&q=3qoOuFNgwUeioKQCtZ9KFA' data = {"name": "Value"} r = requests.get(web_url) print(f'status: {r.status_code}') if r.status_code == 200: raw_data = BeautifulSoup(r.content, 'html.parser') table_content = raw_data.find_all('div', {'class':'cv-structured-list-item cv-structured-list-item--standard bx--structured-list-row'}) print(len(table_content))
方法2:调用API接口
import requests import pandas as pd import re from bs4 import BeautifulSoup import json web_url = 'https://hk.centanet.com/findproperty/api/Transaction/Search' _header = {'lang':'en',"postType":"Both","day":"Day1095","sort":"InsOrRegDate","order":"Descending","size":"24","offset":"24","pageSource":"search","gclid":"Cj0KCQjwnbmaBhD-ARIsAGTPcfVae1prjf_9aKh0dbnaBbzYvi3VhKn4qEXDAQJMS6ZvOiet8GLqzaAaAqH_EALw_wcB","q":"3qoOuFNgwUeioKQCtZ9KFA"} _data = {"name":"Value"} req = requests.post(web_url, headers=_header, json=_data) print(f'Status Code: {req.status_code}') if req.status_code==200: data = json.loads(req.content)['data'] print(len(data)) print(data[0].get('displayText').get('addr'))
解决方案
问题分析
你调用的API接口分页逻辑是通过offset和size参数控制的,但请求存在两个错误:
- 分页参数被错误放在了请求头里,实际需放在请求Body中
- 请求Body内容不完整,缺少正确的筛选参数
修正后的API调用代码(获取前两页)
import requests import json api_url = 'https://hk.centanet.com/findproperty/api/Transaction/Search' # 模拟浏览器请求头,避免被拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36', 'Referer': 'https://hk.centanet.com/findproperty/en/list/transaction?gclid=Cj0KCQjwnbmaBhD-ARIsAGTPcfVae1prjf_9aKh0dbnaBbzYvi3VhKn4qEXDAQJMS6ZvOiet8GLqzaAaAqH_EALw_wcB&q=3qoOuFNgwUeioKQCtZ9KFA', 'lang': 'en' } # 第一页参数:offset=0,取前24条 page1_payload = { "postType": "Both", "day": "Day1095", "sort": "InsOrRegDate", "order": "Descending", "size": 24, "offset": 0, "pageSource": "search", "gclid": "Cj0KCQjwnbmaBhD-ARIsAGTPcfVae1prjf_9aKh0dbnaBbzYvi3VhKn4qEXDAQJMS6ZvOiet8GLqzaAaAqH_EALw_wcB", "q": "3qoOuFNgwUeioKQCtZ9KFA" } response1 = requests.post(api_url, headers=headers, json=page1_payload) if response1.status_code == 200: page1_data = response1.json()['data'] print(f"第一页数据条数: {len(page1_data)}") print(f"第一页第一条地址: {page1_data[0]['displayText']['addr']}") # 第二页参数:offset=24,取第25-48条 page2_payload = page1_payload.copy() page2_payload['offset'] = 24 response2 = requests.post(api_url, headers=headers, json=page2_payload) if response2.status_code == 200: page2_data = response2.json()['data'] print(f"\n第二页数据条数: {len(page2_data)}") print(f"第二页第一条地址: {page2_data[0]['displayText']['addr']}")
批量获取所有页面的思路
- 先调用一次API,从返回结果的
total字段获取总数据条数 - 计算总页数:
total_pages = (总条数 + 每页条数 - 1) // 每页条数 - 循环遍历每一页,修改
offset参数(offset = 页码索引 × 每页条数),发送请求并保存数据
注意事项
- 添加
time.sleep(1)控制请求频率,避免被网站封禁 - 若请求失败,可先访问一次页面获取Cookie后再调用API
内容的提问来源于stack exchange,提问作者humphrey chan
相关产品推荐
相关产品推荐

