You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于动态URL更新Pandas列值?爬虫项目遇难题

问题分析

你遇到的核心问题是每次循环都会重新定义df变量,导致前一个州的数据被完全覆盖,最终只保留最后一个州的结果。给df['state']赋值的操作只是给当前循环的临时DataFrame添加列,无法保留之前州的数据。

解决方案

下面提供两种高效的解决方式,避免逐个合并DataFrame:

方案1:先收集所有行数据,最后一次性生成DataFrame

这种方式效率更高,因为列表的追加操作比反复修改DataFrame更快,尤其适合数据量较大的场景。

import requests
from bs4 import BeautifulSoup as bs
import pandas as pd

# 初始化空列表存储所有行数据
all_rows = []

# 先获取表头(假设所有州页面的表头一致)
first_state = list_of_states[0]
first_url = f'url_here{first_state}'
response = requests.get(first_url)
soup = bs(response.content, 'html.parser')
table_titles = soup.find_all('th')
table_headers = [title.text.strip() for title in table_titles]
# 新增state列的表头
table_headers.append('state')

for i in list_of_states:
    url = f'url_here{i}'
    response = requests.get(url)
    soup = bs(response.content, 'html.parser')
    table = soup.find('table', class_='table_class_here')
    
    column_data = table.find_all('tr')
    # 提取当前州名称
    state = url[x:y]
    
    # 遍历每一行数据
    for row in column_data[1:-1]:
        row_data = row.find_all('td')
        individual_row_data = [data.text.strip() for data in row_data]
        # 给当前行添加州信息
        individual_row_data.append(state)
        # 将行数据加入总列表
        all_rows.append(individual_row_data)

# 最后一次性生成完整的DataFrame
df = pd.DataFrame(all_rows, columns=table_headers)

方案2:初始化主DataFrame,循环追加每个州的数据

如果你更习惯直接操作DataFrame,可以初始化一个空的主DataFrame,每次将单个州的DataFrame追加进去(使用pd.concat替代已废弃的df.append)。

import requests
from bs4 import BeautifulSoup as bs
import pandas as pd

# 初始化空的主DataFrame
df = pd.DataFrame()

for i in list_of_states:
    url = f'url_here{i}'
    response = requests.get(url)
    soup = bs(response.content, 'html.parser')
    table = soup.find('table', class_='table_class_here')
    
    table_titles = soup.find_all('th')
    table_headers = [title.text.strip() for title in table_titles]
    # 生成当前州的临时DataFrame
    state_df = pd.DataFrame(columns=table_headers)
    
    column_data = table.find_all('tr')

    for row in column_data[1:-1]:
        row_data = row.find_all('td')
        individual_row_data = [data.text.strip() for data in row_data]
        length = len(state_df)
        state_df.loc[length] = individual_row_data

    # 添加州信息列
    state = url[x:y]
    state_df['state'] = state
    
    # 将当前州的数据追加到主DataFrame
    df = pd.concat([df, state_df], ignore_index=True)
注意事项
  • 如果不同州页面的表头不一致,方案1需要调整为动态合并表头(可以先收集所有可能的表头,再统一生成DataFrame);方案2的pd.concat会自动处理列名不一致的情况,缺失的列会填充NaN。
  • 建议给requests.get添加超时设置和异常处理,避免爬虫因网络问题中断。

内容的提问来源于stack exchange,提问作者Carl Lennartson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 19:08:30