Python爬取URL嵌套标签数据:获取图片名称及创作者生卒日期
解决方案
首先修正你基础代码里的变量错误:你把请求结果存在了page变量里,但创建BeautifulSoup时用了未定义的webpage,需要统一变量名。
接下来是完整的爬取代码,针对目标数据(图片对应的艺术家名称、出生日期、死亡日期)进行提取:
import pandas as pd import requests from bs4 import BeautifulSoup # 发起请求并解析页面 url = 'https://www.christies.com/en/auction/modern-british-and-irish-art-evening-sale-29987/' response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') # 定位main标签下的所有目标li元素 target_items = soup.select('main li.lots-list-item') # 初始化存储数据的列表 artists_data = [] for item in target_items: # 提取艺术家名称和生卒年(通常在作品标题的子元素里) artist_info = item.find('h3', class_='lot-info__title') if artist_info: # 拆分名称和生卒年,格式一般为 "Artist Name (YYYY-YYYY)" full_text = artist_info.get_text(strip=True) # 处理括号里的生卒年 if '(' in full_text and ')' in full_text: name_part = full_text.split('(')[0].strip() dates_part = full_text.split('(')[1].replace(')', '').strip() # 拆分出生和死亡日期 if '-' in dates_part: birth_date, death_date = dates_part.split('-', 1) birth_date = birth_date.strip() death_date = death_date.strip() else: # 部分可能只有出生或死亡日期,做兼容处理 birth_date = dates_part if dates_part else None death_date = None else: name_part = full_text birth_date = None death_date = None # 提取图片链接(如果需要) img_element = item.find('img', class_='lot-image__img') img_url = img_element['src'] if img_element else None # 将数据添加到列表 artists_data.append({ '艺术家名称': name_part, '出生日期': birth_date, '死亡日期': death_date, '图片链接': img_url }) # 转换为DataFrame并输出 df = pd.DataFrame(artists_data) print(df) # 可选:保存为CSV # df.to_csv('christies_artists.csv', index=False, encoding='utf-8-sig')
代码说明
- 变量修正:将原代码中的
page改为response,确保解析的是正确的请求响应文本。 - 元素定位:使用
select('main li.lots-list-item')精准定位main标签下的目标li元素(若页面结构更新,需对应调整class名称)。 - 信息提取:从每个li的标题元素中拆分艺术家名称和生卒年,兼容只有单日期的特殊情况。
- 图片链接提取:同步获取对应图片的URL,若不需要可删除相关代码块。
内容的提问来源于stack exchange,提问作者Ashish
相关产品推荐
相关产品推荐

