如何基于动态URL更新Pandas列值?爬虫项目遇难题
问题分析
你遇到的核心问题是每次循环都会重新定义df变量,导致前一个州的数据被完全覆盖,最终只保留最后一个州的结果。给df['state']赋值的操作只是给当前循环的临时DataFrame添加列,无法保留之前州的数据。
解决方案
下面提供两种高效的解决方式,避免逐个合并DataFrame:
方案1:先收集所有行数据,最后一次性生成DataFrame
这种方式效率更高,因为列表的追加操作比反复修改DataFrame更快,尤其适合数据量较大的场景。
import requests from bs4 import BeautifulSoup as bs import pandas as pd # 初始化空列表存储所有行数据 all_rows = [] # 先获取表头(假设所有州页面的表头一致) first_state = list_of_states[0] first_url = f'url_here{first_state}' response = requests.get(first_url) soup = bs(response.content, 'html.parser') table_titles = soup.find_all('th') table_headers = [title.text.strip() for title in table_titles] # 新增state列的表头 table_headers.append('state') for i in list_of_states: url = f'url_here{i}' response = requests.get(url) soup = bs(response.content, 'html.parser') table = soup.find('table', class_='table_class_here') column_data = table.find_all('tr') # 提取当前州名称 state = url[x:y] # 遍历每一行数据 for row in column_data[1:-1]: row_data = row.find_all('td') individual_row_data = [data.text.strip() for data in row_data] # 给当前行添加州信息 individual_row_data.append(state) # 将行数据加入总列表 all_rows.append(individual_row_data) # 最后一次性生成完整的DataFrame df = pd.DataFrame(all_rows, columns=table_headers)
方案2:初始化主DataFrame,循环追加每个州的数据
如果你更习惯直接操作DataFrame,可以初始化一个空的主DataFrame,每次将单个州的DataFrame追加进去(使用pd.concat替代已废弃的df.append)。
import requests from bs4 import BeautifulSoup as bs import pandas as pd # 初始化空的主DataFrame df = pd.DataFrame() for i in list_of_states: url = f'url_here{i}' response = requests.get(url) soup = bs(response.content, 'html.parser') table = soup.find('table', class_='table_class_here') table_titles = soup.find_all('th') table_headers = [title.text.strip() for title in table_titles] # 生成当前州的临时DataFrame state_df = pd.DataFrame(columns=table_headers) column_data = table.find_all('tr') for row in column_data[1:-1]: row_data = row.find_all('td') individual_row_data = [data.text.strip() for data in row_data] length = len(state_df) state_df.loc[length] = individual_row_data # 添加州信息列 state = url[x:y] state_df['state'] = state # 将当前州的数据追加到主DataFrame df = pd.concat([df, state_df], ignore_index=True)
注意事项
- 如果不同州页面的表头不一致,方案1需要调整为动态合并表头(可以先收集所有可能的表头,再统一生成DataFrame);方案2的
pd.concat会自动处理列名不一致的情况,缺失的列会填充NaN。 - 建议给
requests.get添加超时设置和异常处理,避免爬虫因网络问题中断。
内容的提问来源于stack exchange,提问作者Carl Lennartson
相关产品推荐
相关产品推荐

