You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法使用Python抓取指定网站图表数据的技术求助

网页动态数据抓取问题解决方案

问题说明

尝试抓取https://www.poder360.com.br/agregador-de-pesquisas/网站内图表对应的表格数据,先后使用requests、requests-html、BeautifulSoup工具均无法获取完整页面内容,右键查看源码也找不到id为method-table的目标表格。

以下是两次尝试的代码:

尝试1(requests-html)

from requests_html import HTMLSession

def get_data(url_path):
    from requests_html import HTMLSession
    session = HTMLSession()

    r = session.get(url_path)
    r.html.render(wait = 8, sleep = 8)

    return r.html

url_path = 'https://www.poder360.com.br/agregador-de-pesquisas'
content = get_data(url_path)
print(content.html)

尝试2(requests + BeautifulSoup)

import requests
import json
from bs4 import BeautifulSoup

url = 'https://www.poder360.com.br/agregador-de-pesquisas'

r = requests.get(url)

soup = BeautifulSoup(r.content, 'html.parser')

print(soup)

问题根源

目标网站的表格数据是通过AJAX动态加载的,不会直接包含在初始返回的HTML源码中。requests仅能获取静态页面内容;requests-html的渲染功能可能因网站反爬机制(如无头浏览器检测)或渲染时机不匹配,未能成功加载出目标数据。

可行解决方法

方法1:抓包获取数据接口

  1. 打开浏览器开发者工具(F12),切换至「网络」标签页后刷新页面。
  2. 过滤XHR/Fetch类型请求,定位返回表格数据的API接口(通常带有json后缀或包含data/survey等关键词)。
  3. 直接请求该API接口,获取结构化JSON数据,无需解析HTML。

示例代码:

import requests

# 替换为实际抓包得到的API地址
api_url = "https://www.poder360.com.br/api/xxx"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
    # 从抓包结果中复制必要请求头,如Referer、Cookie等
}

response = requests.get(api_url, headers=headers)
data = response.json()

# 按需处理数据
print(data)

方法2:使用Playwright模拟真实浏览器

requests-html的无头浏览器兼容性有限,改用Playwright可更好模拟真实浏览器行为,绕过反爬限制:

  1. 安装依赖:
pip install playwright
playwright install chromium
  1. 示例代码:
from playwright.sync_api import sync_playwright

def get_dynamic_table(url):
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        page = browser.new_page(user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36")
        page.goto(url, wait_until="networkidle")
        # 等待目标表格加载完成
        page.wait_for_selector("#method-table")
        # 获取表格HTML内容
        table_html = page.inner_html("#method-table")
        browser.close()
        return table_html

url = "https://www.poder360.com.br/agregador-de-pesquisas"
table_content = get_dynamic_table(url)
print(table_content)

注意事项

  • 抓包时需完整复制请求头,尤其是Cookie和User-Agent,避免被网站拦截。
  • 若API需认证,需模拟登录或提取有效会话信息。
  • 遵守网站robots.txt规则,控制请求频率,防止IP被封禁。

内容的提问来源于stack exchange,提问作者encrypted

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 00:50:36