You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫问题:无法从表格td标签中正确提取Href属性

问题:爬取维基表格时无法提取Club Name列中的链接(返回None)

我是Python新手,查过Stack Overflow相关问题但没解决。现在爬取维基百科美国冰壶俱乐部列表时,DataFrame其他列都正常,但最后一列URL总是返回None,明明第一列的俱乐部名称都带有链接。之前爬非表格页面的href没问题,但表格里提取就遇到了困难,求帮忙。

我的代码:

import requests
import pandas as pd
from bs4 import BeautifulSoup
from urllib.parse import urlparse

url = "https://en.wikipedia.org/wiki/List_of_curling_clubs_in_the_United_States"
data = requests.get(url).text

soup = BeautifulSoup(data, 'lxml')
table = soup.find('table', class_='wikitable sortable')

df = pd.DataFrame(columns=['Club Name', 'City/Town', 'State', 'Type', 'Sheets', 'Memberships', 'Year Founded', 'Notes', 'URL'])

for row in table.tbody.find_all('tr'):    
    # Find all data for each column
    columns = row.find_all('td')
    
    if(columns != []):
        club_name = columns[0].text.strip()
        city = columns[1].text.strip()
        state = columns[2].text.strip()
        type_arena = columns[3].text.strip()
        sheets = columns[4].text.strip()
        memberships = columns[5].text.strip()
        year_founded = columns[6].text.strip()
        notes = columns[7].text.strip()
        club_url = columns[0].find('a').get('href')
        
        df = df.append({'Club Name': club_name,  'City/Town': city, 'State': state, 'Type': type_arena, 'Sheets': sheets, 'Memberships': memberships, 'Year Founded': year_founded, 'Notes': notes, 'URL': club_url}, ignore_index=True)

解决方案

问题核心是:部分俱乐部名称的<a>标签嵌套在其他元素(比如<span>)内部,直接用find('a')无法定位;另外df.append()已被pandas弃用,建议改用更高效的列表收集数据方式。

修改后的代码:

import requests
import pandas as pd
from bs4 import BeautifulSoup

url = "https://en.wikipedia.org/wiki/List_of_curling_clubs_in_the_United_States"
data = requests.get(url).text

soup = BeautifulSoup(data, 'lxml')
table = soup.find('table', class_='wikitable sortable')

# 用列表存储行数据,比循环append更高效
data_rows = []

for row in table.tbody.find_all('tr'):    
    columns = row.find_all('td')
    
    if columns:
        club_name = columns[0].text.strip()
        city = columns[1].text.strip()
        state = columns[2].text.strip()
        type_arena = columns[3].text.strip()
        sheets = columns[4].text.strip()
        memberships = columns[5].text.strip()
        year_founded = columns[6].text.strip()
        notes = columns[7].text.strip()
        
        # 用select_one查找a标签,支持深层嵌套的元素定位
        a_tag = columns[0].select_one('a')
        # 增加判断,避免找不到a标签时抛出报错
        club_url = a_tag.get('href') if a_tag else None
        
        data_rows.append({
            'Club Name': club_name,
            'City/Town': city,
            'State': state,
            'Type': type_arena,
            'Sheets': sheets,
            'Memberships': memberships,
            'Year Founded': year_founded,
            'Notes': notes,
            'URL': club_url
        })

# 一次性生成DataFrame
df = pd.DataFrame(data_rows)

关键修改点:

  • 用columns[0].select_one('a')替代find('a'):select_one支持CSS选择器逻辑,能定位到嵌套在其他标签里的<a>元素
  • 添加空值判断:当单元格无链接时返回None,避免抛出AttributeError
  • 改用列表收集数据后生成DataFrame:彻底替代已弃用的df.append(),提升代码效率和兼容性

内容的提问来源于stack exchange,提问作者Jason Christian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 01:32:26