如何抓取<script>标签内未在soup=page.content中输出的表格及登录爬取问题
解决动态表格爬取与登录后页面抓取问题
一、先完善你的登录流程代码
看你提供的代码片段,execution部分还没写完,先把完整的登录会话流程补全,确保能成功建立已登录状态:
import requests from lxml import html from bs4 import BeautifulSoup import json # 后续处理JS数据会用到 username = "your_username" password = "your_password" login_url = "https://mywebsite.com/login" url = "https://mywebsite.com/profile" def main(): with requests.Session() as s: # 先获取登录页面的csrf token与execution参数(这俩是Spring Security登录的常见必填项) result = s.get(login_url) tree = html.fromstring(result.text) authenticity_token = list(set(tree.xpath("//input[@name='lt']/@value")))[0] execution = list(set(tree.xpath("//input[@name='execution']/@value")))[0] # 构造登录表单数据,注意要和浏览器实际提交的参数完全匹配 login_data = { "username": username, "password": password, "lt": authenticity_token, "execution": execution, "_eventId": "submit" # 多数登录页需要这个参数,可通过浏览器F12抓包确认 } # 发送登录请求 login_response = s.post(login_url, data=login_data) # 简单验证登录是否成功,比如检查页面是否包含欢迎语或用户名 if "Welcome" in login_response.text: print("登录成功!") else: print("登录失败,请检查账号密码或表单参数") return # 访问目标页面 profile_response = s.get(url) soup = BeautifulSoup(profile_response.text, "html.parser")
二、抓取
相关产品推荐
相关产品推荐

