You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python能否在函数内调用其他函数?爬虫场景合并数据采集函数问题

合并后实现代码

你原来的代码存在两次重复请求同一页面的冗余问题,合并后仅需发起一次请求,同时完成两类信息的采集,还能避免分别采集后拼接出现的对应关系错位问题:

import requests
from bs4 import BeautifulSoup
import pandas as pd

headers = {'User-Agent': 'Mozilla/5.0 (X11; Linux i586; rv:31.0) Gecko/20100101 Firefox/31.0'}
link = "https://www.allmusic.com/mood/tender-xa0000001119/songs"

def get_song_performer(url):
    # 单次请求页面
    source_code = requests.get(url, headers=headers)
    soup = BeautifulSoup(source_code.text, 'html.parser')
    songs = []
    performers = []
    # 逐行解析表格内容,保证歌曲和演唱者对应关系正确
    for tr in soup.select('table tr'):
        # 提取当前行的歌曲名和演唱者
        title_td = tr.find('td', class_='title')
        performer_td = tr.find('td', class_='performer')
        if title_td and performer_td:
            # 取歌曲名
            song_tag = title_td.find('a')
            song = song_tag.string.strip() if song_tag else ''
            # 取演唱者
            performer_tag = performer_td.find('a')
            performer = performer_tag.string.strip() if performer_tag else ''
            songs.append(song)
            performers.append(performer)
    # 直接生成DataFrame返回
    return pd.DataFrame({'song': songs, 'performer': performers})

# 仅需调用一次函数即可得到结果DataFrame
df = get_song_performer(link)
优化说明
  • 去掉了两次重复的页面请求,降低被网站反爬拦截的概率,同时提升采集效率
  • 改为逐行解析表格行数据,从根源上避免了分别采集两类字段后,因数据条数不一致导致的拼接错位问题
  • 所有变量改为函数内部变量,不会污染全局命名空间
  • 增加了空值判断逻辑,遇到字段缺失的行也不会报错中断采集

内容的提问来源于stack exchange,提问作者user12076260

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 14:06:04