无法使用Python抓取指定网站图表数据的技术求助
网页动态数据抓取问题解决方案
问题说明
尝试抓取https://www.poder360.com.br/agregador-de-pesquisas/网站内图表对应的表格数据,先后使用requests、requests-html、BeautifulSoup工具均无法获取完整页面内容,右键查看源码也找不到id为method-table的目标表格。
以下是两次尝试的代码:
尝试1(requests-html)
from requests_html import HTMLSession def get_data(url_path): from requests_html import HTMLSession session = HTMLSession() r = session.get(url_path) r.html.render(wait = 8, sleep = 8) return r.html url_path = 'https://www.poder360.com.br/agregador-de-pesquisas' content = get_data(url_path) print(content.html)
尝试2(requests + BeautifulSoup)
import requests import json from bs4 import BeautifulSoup url = 'https://www.poder360.com.br/agregador-de-pesquisas' r = requests.get(url) soup = BeautifulSoup(r.content, 'html.parser') print(soup)
问题根源
目标网站的表格数据是通过AJAX动态加载的,不会直接包含在初始返回的HTML源码中。requests仅能获取静态页面内容;requests-html的渲染功能可能因网站反爬机制(如无头浏览器检测)或渲染时机不匹配,未能成功加载出目标数据。
可行解决方法
方法1:抓包获取数据接口
- 打开浏览器开发者工具(F12),切换至「网络」标签页后刷新页面。
- 过滤XHR/Fetch类型请求,定位返回表格数据的API接口(通常带有
json后缀或包含data/survey等关键词)。 - 直接请求该API接口,获取结构化JSON数据,无需解析HTML。
示例代码:
import requests # 替换为实际抓包得到的API地址 api_url = "https://www.poder360.com.br/api/xxx" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36", # 从抓包结果中复制必要请求头,如Referer、Cookie等 } response = requests.get(api_url, headers=headers) data = response.json() # 按需处理数据 print(data)
方法2:使用Playwright模拟真实浏览器
requests-html的无头浏览器兼容性有限,改用Playwright可更好模拟真实浏览器行为,绕过反爬限制:
- 安装依赖:
pip install playwright playwright install chromium
- 示例代码:
from playwright.sync_api import sync_playwright def get_dynamic_table(url): with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page(user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36") page.goto(url, wait_until="networkidle") # 等待目标表格加载完成 page.wait_for_selector("#method-table") # 获取表格HTML内容 table_html = page.inner_html("#method-table") browser.close() return table_html url = "https://www.poder360.com.br/agregador-de-pesquisas" table_content = get_dynamic_table(url) print(table_content)
注意事项
- 抓包时需完整复制请求头,尤其是Cookie和User-Agent,避免被网站拦截。
- 若API需认证,需模拟登录或提取有效会话信息。
- 遵守网站
robots.txt规则,控制请求频率,防止IP被封禁。
内容的提问来源于stack exchange,提问作者encrypted
相关产品推荐
相关产品推荐

