如何使用BeautifulSoup抓取多页URL?
网页抓取瓶颈解决思路
- 确认页面加载类型:先判断目标页面是静态HTML还是JS动态渲染。直接在浏览器右键查看「页面源代码」,搜索你要抓取的核心数据(比如合同名称、金额),如果找不到,说明页面依赖JavaScript加载内容,此时
requests+BeautifulSoup无法获取完整数据,需要改用Selenium、Playwright这类工具模拟浏览器运行。 - 精准定位目标元素:如果是静态页面,用浏览器开发者工具(F12)定位数据所在的HTML标签、类名或ID。比如要提取合同列表表格:
# 示例:定位class为特定值的表格 contract_table = soup.find("table", class_="listAWC") if contract_table: # 遍历表格行(跳过表头) for row in contract_table.find_all("tr")[1:]: columns = row.find_all("td") # 提取每行的具体字段 contract_no = columns[0].text.strip() procuring_entity = columns[1].text.strip() # 按需提取其他字段 - 实现分页抓取:你的URL中
d-3998960-p=1是页码参数,可通过循环修改该参数值实现多页数据抓取:base_url = "https://www.taneps.go.tz/epps/viewAllAwardedContracts.do?d-3998960-p={}&selectedItem=viewAllAwardedContracts.do&T01_ps=100" # 假设抓取前5页,可根据实际总页数调整 for page_num in range(1, 6): target_url = base_url.format(page_num) response = session.get(target_url) soup = bs(response.content, "html.parser") # 处理当前页数据逻辑 - 规避反爬限制:除了设置User-Agent,还可添加以下措施:
- 从浏览器复制真实Cookie添加到Session中
- 每次请求后设置1-3秒的间隔(
import time; time.sleep(2)) - 若频繁被封,使用代理IP池轮换请求IP
- 完善错误处理:添加异常捕获避免程序中断:
import time try: response = session.get(target_url, timeout=15) # 检查HTTP状态码 response.raise_for_status() soup = bs(response.content, "html.parser") except requests.exceptions.RequestException as e: print(f"第{page_num}页请求失败: {str(e)}") time.sleep(5) # 失败后延迟重试 continue - 数据持久化存储:将抓取的数据存入CSV或JSON文件,方便后续分析:
import csv # 初始化CSV文件 with open("awarded_contracts.csv", "w", newline="", encoding="utf-8") as f: writer = csv.writer(f) # 写入表头 writer.writerow(["合同编号", "采购实体", "中标方"]) # 循环写入数据行 writer.writerow([contract_no, procuring_entity, vendor_name])
内容的提问来源于stack exchange,提问作者tony michael
相关产品推荐
相关产品推荐

