You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用requests爬取DataFrame链接报错:InvalidSchema问题求助

问题解决:InvalidSchema 错误处理

问题原因

看错误提示里的URL:'https://en.wikipedia.org/wiki/Wagner_Group',能发现URL前后多了一对单引号。这是因为你的数据集里,链接是带着单引号被存入DataFrame的。单独测试的URL没有多余引号所以正常运行,但apply时传入的是带单引号的字符串,requests无法识别这种格式的URL,因此抛出InvalidSchema错误。

解决方案

方案1:先清理DataFrame的链接列

在调用apply前,先去掉链接前后的单引号:

# 清理links列的单引号
df['links'] = df['links'].str.strip("'")
# 再调用爬虫函数
df['links'].apply(get_data)

方案2:在爬虫函数内处理URL

如果不想修改原DataFrame,可以在get_data函数里先清理URL:

def get_data(url): 
    # 移除URL前后的单引号
    cleaned_url = url.strip("'")
    page = requests.get(cleaned_url)
    soup = BeautifulSoup(page.content,'html.parser')
    text = ""
    for paragraph in soup.find_all('p'):
        text += paragraph.text
    return(text)

df['links'].apply(get_data)

额外优化建议

可以给函数加异常处理,避免单个URL出错导致整个apply流程中断:

def get_data(url): 
    try:
        cleaned_url = url.strip("'")
        page = requests.get(cleaned_url)
        page.raise_for_status()  # 检查请求是否成功(比如404、500错误)
        soup = BeautifulSoup(page.content,'html.parser')
        # 用join简化文本拼接
        text = "".join([p.text for p in soup.find_all('p')])
        return text
    except Exception as e:
        print(f"处理URL {url} 失败: {str(e)}")
        return None

# 将爬取结果存入新列
df['text_content'] = df['links'].apply(get_data)

内容的提问来源于stack exchange,提问作者cps

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 17:57:26