You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python BeautifulSoup遍历URL列表并爬取指定td标签数据

问题根因

你当前的td标签提取代码写在for循环外部,循环执行过程中每次请求新的URL都会覆盖soup变量的存储内容,循环结束后soup仅保留最后一个URL的页面解析结果,因此只能拿到最后一页的td数据。

调整后可用代码

import numpy as np
import pandas as pd
from datetime import datetime
import pytz
import requests
import json
from bs4 import BeautifulSoup

# 新增请求头避免被站点反爬规则拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

url_list = ['https://www.coingecko.com/en/coins/ethereum/historical_data/usd?start_date=2021-08-06&end_date=2021-09-05#panel',
            'https://www.coingecko.com/en/coins/cardano/historical_data/usd?start_date=2021-08-06&end_date=2021-09-05#panel',
            'https://www.coingecko.com/en/coins/chainlink/historical_data/usd?start_date=2021-08-06&end_date=2021-09-05#panel']

# 全局容器存储所有页面的提取结果
all_td_data = []

for url in url_list:
    response = requests.get(url, headers=headers)
    # 新增状态码校验,过滤请求失败的无效数据
    if response.status_code == 200:
        soup = BeautifulSoup(response.text, 'html.parser')
        # 提取当前页面的目标td标签,如果需要纯文本可以改为 [td.get_text(strip=True) for td in soup.find_all("td", class_ = "text-center")]
        current_td = soup.find_all("td", class_ = "text-center")
        all_td_data.append({
            "source_url": url,
            "td_list": current_td
        })
    else:
        print(f"请求异常,URL:{url},状态码:{response.status_code}")

# 查看所有结果示例
for item in all_td_data:
    print(f"来源URL:{item['source_url']}")
    print(f"提取到td标签数量:{len(item['td_list'])}")
    # 打印前2个标签做示例,避免输出内容过多
    print(item['td_list'][:2])

核心调整点

  • 循环外新增全局存储容器,避免每次循环的数据被覆盖
  • 将td提取逻辑移入循环内部,每个页面解析完成后立即提取对应内容
  • 优化循环写法直接遍历URL列表,代码更简洁易读
  • 新增请求头和状态码校验,降低反爬拦截概率,过滤无效请求

内容的提问来源于stack exchange,提问作者Roxana Slj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 06:45:03