You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup进行网页数据抓取时出现报错的问题求解

代码问题排查与修复方案

存在的问题

  • 循环请求URL时参数错误:遍历urls列表时,requests.get()传入了整个列表urls,应该传入单次循环的单个URLi
  • find_all()返回值用法错误:soup.find_all('h1')返回的是h1标签列表,不能直接调用get_text(),需改为用find('h1')获取单个h1标签,或遍历列表提取内容
  • 未做非空判断:查找类名为organism-3-col的div时,若页面不存在该元素,直接调用find_all('img')会抛出空对象属性报错,需增加非空判断
  • 属性提取语法错误:提取src、alt属性赋值时误用了圆括号image('src'),正确语法为方括号image['src']
  • 缩进逻辑混乱:h1提取、图片查找、数据追加等逻辑都应该放在URL遍历的循环内部,原代码缩进错误会导致仅处理最后一个URL,或抛出变量未定义报错
  • 列名不统一:初始DataFrame的列名和新增DataFrame的列名不一致,会导致数据追加错位
  • 冗余代码:声明的csv写入对象writer全程未使用,可直接删除
  • 弃用方法问题:新版pandas已移除append方法,需改用pd.concat实现数据追加

修复后的完整代码

from bs4 import BeautifulSoup
import requests 
import pandas as pd

# 定义所有待抓取的URL列表
urls = ['https://www.rpsgroup.com/services/','https://www.rpsgroup.com']

# 创建空白数据框
df = pd.DataFrame(columns=['Header1','Image Source', 'Alt Tag'])

# HTML解析逻辑
for i in urls:               
    Web_page = requests.get(i)
    soup = BeautifulSoup(Web_page.content, 'html.parser')

    # 提取页面h1内容
    h1 = soup.find('h1')
    header1 = h1.get_text().strip() if h1 else ''
    print(header1)

    # 查找指定类名下的所有图片元素
    img_container = soup.find('div',class_='organism-3-col')
    if not img_container:
        continue
    image_link = img_container.find_all('img')
    
    # 提取图片src属性与alt标签内容
    for image in image_link:
        image_source = image['src'] if image.has_attr('src') else 'Source not found'
        print(image_source)
        image_alt = image['alt'] if image.has_attr('alt') else 'Alt not found'
        print(image_alt)
        # 逐行追加数据
        new_row = pd.DataFrame([[header1, image_source, image_alt]], columns=df.columns)
        df = pd.concat([df, new_row], ignore_index=True)

# 导出结果到csv文件
df.to_csv('scrape.csv', index=False)

内容的提问来源于stack exchange,提问作者Mandeep Singh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 16:45:03