使用BeautifulSoup爬取网页时提取<td>标签内国籍文本的问题咨询
你报错的原因是artist_nationality_list_items是所有标签组成的列表,列表本身没有contents属性,只有单个BeautifulSoup标签元素才能调用该属性。
核心实现思路
页面内的每条艺术家数据都是固定结构:第一个存带链接的姓名,相邻的下一个存国籍信息,直接通过节点的相邻关系定位国籍信息最稳妥,避免索引错位问题。
修改后可运行的完整代码
import requests import csv from bs4 import BeautifulSoup def findName(): page = requests.get('https://web.archive.org/web/20121007172955/https://www.nga.gov/collection/anB1.htm') soup = BeautifulSoup(page.text, 'html.parser') last_links = soup.find(class_='AlphaNav') last_links.decompose() # 写入csv时指定编码避免乱码 f = csv.writer(open('h-artist_lastname.csv', 'w', encoding='utf-8')) f.writerow(['Last Name, First Name', 'Nationality', 'Link']) artist_name_list = soup.find(class_='BodyText') artist_name_list_items = artist_name_list.find_all('a') for artist_name in artist_name_list_items: names = artist_name.contents[0] links = 'https://web.archive.org' + artist_name.get('href') # 定位国籍信息:先拿到a标签所属的td,再找下一个相邻的td nationality_td = artist_name.parent.find_next_sibling('td') nationalities = nationality_td.text.strip() if nationality_td else '无国籍信息' f.writerow([names, nationalities, links]) findName()
关键修改说明
- 去掉了原来无意义的全量抓取逻辑,通过节点关系直接定位对应国籍标签,不会出现数据错位
- 增加了空值判断,避免个别行没有国籍信息时触发报错
- 写入csv时增加了utf-8编码指定,避免特殊字符乱码
内容的提问来源于stack exchange,提问作者Azure21
相关产品推荐
相关产品推荐

