Python爬虫数据错位问题:店铺名称与网站对应不符修复求助
漫画店铺爬取数据错位修复
问题背景
使用BeautifulSoup和requests爬取https://www.comicshoplocator.com的漫画店铺信息时,因部分店铺不存在LocationShopProfile或StoreWeb元素,导致DataFrame中Name列与Website列对应错位,出现名称和网站不匹配的情况。
原代码
from bs4 import BeautifulSoup import requests data0 = [] data1 = [] response = requests.get( "https://www.comicshoplocator.com/StoreLocatorPremier?query=75077&showCsls=true" ) soup = BeautifulSoup(response.text, "html.parser") for tag in soup.find_all('div', class_="LocationName"): title = tag.text data0.append({ 'title': title }) for button in soup.find_all('div', class_="LocationDetails"): for childdiv in button.find_all('div', class_="LocationShopProfile"): for zb in childdiv.find_all('a'): if zb.get_text() == 'Shop Profile': website = zb.get('href') forsite = requests.get('https://www.comicshoplocator.com/' + website) soup = BeautifulSoup(forsite.text, "html.parser") for tag in soup.find_all('div', class_="StoreWeb"): site = tag.text.replace('Web: http://', '') data7.append({ 'site': site }) df = pd.DataFrame(columns=['Name', 'Website']) df[df.columns[0]] = pd.DataFrame(data0) df[df.columns[1]] = pd.DataFrame(data1)
当前输出结果
Name Website 0 TWENTY ELEVEN COMICS WWW.TWENTYELEVENCOMICS.COM 1 READ COMICS www.boomerangcomics.com 2 BOOMERANG COMICS www.facebook.com/morefuncomics 3 MORE FUN COMICS AND GAMES www.madnesscomicsandgames.com 4 MADNESS COMICS & GAMES NaN 5 SANCTUARY BOOKS AND GAMES NaN
预期正确结果
Name Website 0 TWENTY ELEVEN COMICS WWW.TWENTYELEVENCOMICS.COM 1 READ COMICS NaN 2 BOOMERANG COMICS www.boomerangcomics.com 3 MORE FUN COMICS AND GAMES www.facebook.com/morefuncomics 4 MADNESS COMICS & GAMES www.madnesscomicsandgames.com 5 SANCTUARY BOOKS AND GAMES NaN
修复方案
核心问题是原代码分开收集名称和网站信息,没有将每个店铺的名称与对应网站绑定,导致无网站的店铺无法生成对应空值,最终索引错位。修复思路是:遍历每个店铺的完整容器,在同一个循环内获取名称和网站信息,确保每条店铺数据的完整性。
修改后的代码
from bs4 import BeautifulSoup import requests import pandas as pd # 存储每条店铺的完整信息 shop_data = [] response = requests.get( "https://www.comicshoplocator.com/StoreLocatorPremier?query=75077&showCsls=true" ) soup = BeautifulSoup(response.text, "html.parser") # 遍历每个店铺的名称标签,同时关联对应的详情容器 for name_tag in soup.find_all('div', class_="LocationName"): shop_name = name_tag.text.strip() shop_website = None # 默认设为None,转DataFrame后自动转为NaN # 获取当前名称对应的详情容器 details_tag = name_tag.find_next_sibling('div', class_="LocationDetails") if details_tag: # 查找店铺profile入口 profile_div = details_tag.find('div', class_="LocationShopProfile") if profile_div: shop_profile_link = profile_div.find('a', text='Shop Profile') if shop_profile_link: # 请求详情页获取网站信息 profile_url = f"https://www.comicshoplocator.com/{shop_profile_link.get('href')}" profile_response = requests.get(profile_url) profile_soup = BeautifulSoup(profile_response.text, "html.parser") web_tag = profile_soup.find('div', class_="StoreWeb") if web_tag: shop_website = web_tag.text.replace('Web: http://', '').strip() # 将单条店铺信息加入列表 shop_data.append({ 'Name': shop_name, 'Website': shop_website }) # 转换为DataFrame df = pd.DataFrame(shop_data) print(df)
关键修复点
- 绑定单条数据:在同一个循环内处理单个店铺的名称和网站,确保每条数据一一对应,避免错位。
- 空值兜底:默认将网站设为
None,即使找不到对应元素也保留空值占位,保证列表长度与店铺数量一致。 - 节点关联:通过
find_next_sibling将名称和对应详情容器绑定,确保操作的是同一店铺的信息。 - 修正变量错误:修复原代码中
data7未定义、data1未使用的问题。
内容的提问来源于stack exchange,提问作者Ave
相关产品推荐
相关产品推荐

