You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新手求助:用Beautiful Soup爬取加州参议员网页并生成Pandas DataFrame

解决加州参议员网页数据爬取问题

问题描述

作为Beautiful Soup和HTML新手,尝试爬取加州参议员网页,目标提取参议员姓名、党派、选区及议会办公室电话号码并整理为Pandas DataFrame。查看网页源码发现h3标签关联姓名与党派信息,地址和电话在p标签中,但查找所有h3标签得到201个结果,远超实际参议员数量,无法精准筛选提取所需信息。已完成请求和HTML解析,但提取目标数据时遇到困难,附上尝试的代码。

原尝试代码:

import requests
from bs4 import BeautifulSoup
import pandas as pd

# Send a GET request to the website
url = "https://www.senate.ca.gov/senators"
response = requests.get(url)

# Use Beautiful Soup to parse the HTML   
soup = BeautifulSoup(response.content, "html.parser")

# Find the table that contains the senator    information
table = soup.find("table", {"class": "views-  table cols-4"})

# Create lists to store the data
names = []
districts = []
parties = []
phones = []

# Extract the senator information from each row in the table
for row in table.find_all("tr"):
     cells = row.find_all("td")
     if len(cells) == 4:
         name = cells[0].get_text().strip()
         district =   cells[1].get_text().strip()
         party = cells[2].get_text().strip()
         phone = cells[3].get_text().strip()
    
    # Append the data to the lists
    names.append(name)
    districts.append(district)
    parties.append(party)
    phones.append(phone)

 # Create a Pandas dataframe from the lists
 df = pd.DataFrame({"Senator Name": names,  "District": districts, "Party": parties, "Phone Number": phones})

# Print the dataframe
 print(df)

问题分析

原代码错误地假设数据存放在views-table cols-4类的表格中,但实际网页的参议员数据是在独立的.views-row卡片容器里,并非表格结构,因此无法定位到正确的数据区域,导致提取失败。

修正后的代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

# 请求网页并解析
url = "https://www.senate.ca.gov/senators"
response = requests.get(url)
soup = BeautifulSoup(response.content, "html.parser")

# 存储数据的列表
senator_names = []
districts = []
parties = []
office_phones = []

# 定位每个参议员的卡片容器
senator_cards = soup.find_all("div", class_="views-row")

for card in senator_cards:
    # 提取姓名与党派
    h3_text = card.find("h3").get_text(strip=True)
    # 拆分姓名和党派(格式为"姓名 (党派)")
    name_part = h3_text.split(" (")[0]
    party = h3_text.split("(")[1].replace(")", "").strip()
    
    # 提取选区
    district = card.find("div", class_="senator-district").get_text(strip=True).replace("District ", "")
    
    # 提取办公室电话
    phone_tag = card.find("p", string=lambda text: text and "Phone:" in text)
    phone = phone_tag.get_text(strip=True).replace("Phone:", "").strip() if phone_tag else "N/A"
    
    # 添加到列表
    senator_names.append(name_part)
    districts.append(district)
    parties.append(party)
    office_phones.append(phone)

# 生成DataFrame并输出
df = pd.DataFrame({
    "Senator Name": senator_names,
    "District": districts,
    "Party": parties,
    "Office Phone": office_phones
})

print(df)

代码说明

  • 定位数据容器:使用.views-row类定位每个参议员的独立卡片,确保只获取有效数据项,避免无关的h3标签干扰。
  • 拆分姓名与党派:从h3标签文本中拆分出姓名和党派信息,处理格式为"姓名 (党派)"的文本内容。
  • 提取选区:直接从.senator-district类元素中提取选区编号,清理多余文本。
  • 提取电话:通过筛选包含"Phone:"的p标签定位电话信息,处理无电话的情况为"N/A"。

内容的提问来源于stack exchange,提问作者Kris L

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 23:42:25