Python爬虫问题:无法从表格td标签中正确提取Href属性
问题:爬取维基表格时无法提取Club Name列中的链接(返回None)
我是Python新手,查过Stack Overflow相关问题但没解决。现在爬取维基百科美国冰壶俱乐部列表时,DataFrame其他列都正常,但最后一列URL总是返回None,明明第一列的俱乐部名称都带有链接。之前爬非表格页面的href没问题,但表格里提取就遇到了困难,求帮忙。
我的代码:
import requests import pandas as pd from bs4 import BeautifulSoup from urllib.parse import urlparse url = "https://en.wikipedia.org/wiki/List_of_curling_clubs_in_the_United_States" data = requests.get(url).text soup = BeautifulSoup(data, 'lxml') table = soup.find('table', class_='wikitable sortable') df = pd.DataFrame(columns=['Club Name', 'City/Town', 'State', 'Type', 'Sheets', 'Memberships', 'Year Founded', 'Notes', 'URL']) for row in table.tbody.find_all('tr'): # Find all data for each column columns = row.find_all('td') if(columns != []): club_name = columns[0].text.strip() city = columns[1].text.strip() state = columns[2].text.strip() type_arena = columns[3].text.strip() sheets = columns[4].text.strip() memberships = columns[5].text.strip() year_founded = columns[6].text.strip() notes = columns[7].text.strip() club_url = columns[0].find('a').get('href') df = df.append({'Club Name': club_name, 'City/Town': city, 'State': state, 'Type': type_arena, 'Sheets': sheets, 'Memberships': memberships, 'Year Founded': year_founded, 'Notes': notes, 'URL': club_url}, ignore_index=True)
解决方案
问题核心是:部分俱乐部名称的<a>标签嵌套在其他元素(比如<span>)内部,直接用find('a')无法定位;另外df.append()已被pandas弃用,建议改用更高效的列表收集数据方式。
修改后的代码:
import requests import pandas as pd from bs4 import BeautifulSoup url = "https://en.wikipedia.org/wiki/List_of_curling_clubs_in_the_United_States" data = requests.get(url).text soup = BeautifulSoup(data, 'lxml') table = soup.find('table', class_='wikitable sortable') # 用列表存储行数据,比循环append更高效 data_rows = [] for row in table.tbody.find_all('tr'): columns = row.find_all('td') if columns: club_name = columns[0].text.strip() city = columns[1].text.strip() state = columns[2].text.strip() type_arena = columns[3].text.strip() sheets = columns[4].text.strip() memberships = columns[5].text.strip() year_founded = columns[6].text.strip() notes = columns[7].text.strip() # 用select_one查找a标签,支持深层嵌套的元素定位 a_tag = columns[0].select_one('a') # 增加判断,避免找不到a标签时抛出报错 club_url = a_tag.get('href') if a_tag else None data_rows.append({ 'Club Name': club_name, 'City/Town': city, 'State': state, 'Type': type_arena, 'Sheets': sheets, 'Memberships': memberships, 'Year Founded': year_founded, 'Notes': notes, 'URL': club_url }) # 一次性生成DataFrame df = pd.DataFrame(data_rows)
关键修改点:
- 用
columns[0].select_one('a')替代find('a'):select_one支持CSS选择器逻辑,能定位到嵌套在其他标签里的<a>元素 - 添加空值判断:当单元格无链接时返回None,避免抛出
AttributeError - 改用列表收集数据后生成DataFrame:彻底替代已弃用的
df.append(),提升代码效率和兼容性
内容的提问来源于stack exchange,提问作者Jason Christian
相关产品推荐
相关产品推荐

