You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:无法抓取<svg>标签内不可查看数据及政党得票率、得票数

解决印度选举委员会政党结果页面的数据抓取问题

看起来你遇到的核心问题有三个:一是没定位到包含**得票率(vote%)和得票数(vote count)**的正确数据区域,二是不知道如何提取<svg>标签内的数值,三是原代码没处理页面的分页内容。我来一步步帮你解决:

问题分析

  1. 原代码只聚焦了id="piecharts26"的div,但得票率和得票数其实分散在页面的政党详情表格、SVG标签中;
  2. <svg>里的投票占比数值并非隐藏数据,而是直接嵌在<text>标签内,BeautifulSoup完全可以解析;
  3. 页面分页有规律(第一页是PartyWiseResultS26.htm,第二页是PartyWiseResultS26_2.htm),需要主动遍历分页URL。

解决方案代码

我重写了代码,用requests替代urllib(更简洁易用),同时处理分页、提取SVG数据,并输出你需要的CSV格式:

import requests
from bs4 import BeautifulSoup
import csv
import os

def get_soup(url):
    # 模拟浏览器请求头,避免被网站拦截
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"}
    response = requests.get(url, headers=headers)
    response.encoding = 'utf-8'
    return BeautifulSoup(response.text, "html.parser")

def extract_party_data(soup, state_code):
    party_data = []
    # 提取政党详情表格(包含得票数)
    for table in soup.find_all("table", {"class": "table table-striped table-bordered"}):
        rows = table.find_all("tr")
        for row in rows[1:]:  # 跳过表头行
            cols = row.find_all("td")
            if len(cols) >= 5:
                party_name = cols[1].text.strip()
                vote_count = cols[4].text.strip().replace(",", "")  # 去除逗号格式化数字
                
                # 从SVG中提取对应政党的得票率
                svg = soup.find("svg", {"id": f"piechart{state_code}"})
                vote_percent = None
                if svg:
                    # SVG里的得票率在<text>标签中,筛选带%的文本
                    text_elements = svg.find_all("text")
                    for text in text_elements:
                        text_content = text.text.strip()
                        if "%" in text_content:
                            vote_percent = text_content
                            break  # 匹配到第一个百分比后跳出(可根据页面结构调整匹配逻辑)
                
                party_data.append([party_name, vote_percent, vote_count, state_code])
    return party_data

# 配置要抓取的邦代码和分页URL
state_codes = ["S26"]
base_url = "http://eciresults.nic.in/PartyWiseResult{}.htm"
all_data = []
header = ["Party", "Vote%", "Vote Count", "State Code"]

for code in state_codes:
    # 抓取第一页
    url_page1 = base_url.format(code)
    soup_page1 = get_soup(url_page1)
    data_page1 = extract_party_data(soup_page1, code)
    all_data.extend(data_page1)
    
    # 抓取第二页
    url_page2 = base_url.format(f"{code}_2")
    soup_page2 = get_soup(url_page2)
    data_page2 = extract_party_data(soup_page2, code)
    all_data.extend(data_page2)

# 写入CSV文件
output_file = os.path.expanduser("party_election_results.csv")
with open(output_file, "w", newline="", encoding="utf-8") as f:
    writer = csv.writer(f)
    writer.writerow(header)
    writer.writerows(all_data)

print(f"数据已成功写入 {output_file}")

关键说明

  • 请求头设置:添加User-Agent模拟浏览器访问,避免被网站识别为爬虫拦截;
  • SVG数据提取:直接解析SVG标签下的<text>元素,筛选包含%的文本即可获取得票率;如果页面中SVG文本和表格行有严格对应关系,可进一步调整索引精准匹配;
  • 分页处理:根据URL规律拼接第二页地址,确保抓取全部政党数据;
  • CSV格式化:使用Python内置csv模块,避免手动拼接字符串出现格式错误,同时支持UTF-8编码处理特殊字符。

如果SVG的得票率匹配不够精准,你可以查看页面SVG的结构(比如每个政党的百分比文本是否有特定class或位置属性),再调整提取逻辑即可。

内容的提问来源于stack exchange,提问作者Vignesh Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:30:05