新手求助:用Beautiful Soup爬取加州参议员网页并生成Pandas DataFrame
解决加州参议员网页数据爬取问题
问题描述
作为Beautiful Soup和HTML新手,尝试爬取加州参议员网页,目标提取参议员姓名、党派、选区及议会办公室电话号码并整理为Pandas DataFrame。查看网页源码发现h3标签关联姓名与党派信息,地址和电话在p标签中,但查找所有h3标签得到201个结果,远超实际参议员数量,无法精准筛选提取所需信息。已完成请求和HTML解析,但提取目标数据时遇到困难,附上尝试的代码。
原尝试代码:
import requests from bs4 import BeautifulSoup import pandas as pd # Send a GET request to the website url = "https://www.senate.ca.gov/senators" response = requests.get(url) # Use Beautiful Soup to parse the HTML soup = BeautifulSoup(response.content, "html.parser") # Find the table that contains the senator information table = soup.find("table", {"class": "views- table cols-4"}) # Create lists to store the data names = [] districts = [] parties = [] phones = [] # Extract the senator information from each row in the table for row in table.find_all("tr"): cells = row.find_all("td") if len(cells) == 4: name = cells[0].get_text().strip() district = cells[1].get_text().strip() party = cells[2].get_text().strip() phone = cells[3].get_text().strip() # Append the data to the lists names.append(name) districts.append(district) parties.append(party) phones.append(phone) # Create a Pandas dataframe from the lists df = pd.DataFrame({"Senator Name": names, "District": districts, "Party": parties, "Phone Number": phones}) # Print the dataframe print(df)
问题分析
原代码错误地假设数据存放在views-table cols-4类的表格中,但实际网页的参议员数据是在独立的.views-row卡片容器里,并非表格结构,因此无法定位到正确的数据区域,导致提取失败。
修正后的代码
import requests from bs4 import BeautifulSoup import pandas as pd # 请求网页并解析 url = "https://www.senate.ca.gov/senators" response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") # 存储数据的列表 senator_names = [] districts = [] parties = [] office_phones = [] # 定位每个参议员的卡片容器 senator_cards = soup.find_all("div", class_="views-row") for card in senator_cards: # 提取姓名与党派 h3_text = card.find("h3").get_text(strip=True) # 拆分姓名和党派(格式为"姓名 (党派)") name_part = h3_text.split(" (")[0] party = h3_text.split("(")[1].replace(")", "").strip() # 提取选区 district = card.find("div", class_="senator-district").get_text(strip=True).replace("District ", "") # 提取办公室电话 phone_tag = card.find("p", string=lambda text: text and "Phone:" in text) phone = phone_tag.get_text(strip=True).replace("Phone:", "").strip() if phone_tag else "N/A" # 添加到列表 senator_names.append(name_part) districts.append(district) parties.append(party) office_phones.append(phone) # 生成DataFrame并输出 df = pd.DataFrame({ "Senator Name": senator_names, "District": districts, "Party": parties, "Office Phone": office_phones }) print(df)
代码说明
- 定位数据容器:使用
.views-row类定位每个参议员的独立卡片,确保只获取有效数据项,避免无关的h3标签干扰。 - 拆分姓名与党派:从h3标签文本中拆分出姓名和党派信息,处理格式为"姓名 (党派)"的文本内容。
- 提取选区:直接从
.senator-district类元素中提取选区编号,清理多余文本。 - 提取电话:通过筛选包含"Phone:"的p标签定位电话信息,处理无电话的情况为"N/A"。
内容的提问来源于stack exchange,提问作者Kris L
相关产品推荐
相关产品推荐

