You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

遍历URL列表提取页面标题时遇requests.exceptions.InvalidSchema错误求助

问题描述

尝试遍历一个URL列表,使用requests和BeautifulSoup提取每个URL对应的页面标题,但持续遇到以下错误:

requests.exceptions.InvalidSchema: No connection adapters were found for "['https://reddit.com/?feed=home', 'https://reddit.com/chunkCSS/CollectionCommentsPageCommentsPageCountryPageFrontpageGovernanceReleaseNotesModalModListingMod~e3d63e32.74eb929a3827c754ba25_.css', 'https://reddit.com/chunkCSS/CountryPageFrontpageModListingMultiredditProfileCommentsProfileOverviewProfilePosts~Subreddit.e72fce90a7f3165091b9_.css', 'https://reddit.com/chunkCSS/Frontpage.85a25b7700617eafa94b_.css', 'https://reddit.com/?feed=home', 'https://reddit.com/r/popular/',]

相关代码如下:

pages = []
for admin_login_pages in domains:
    with open("urls.txt", "w") as f:
        f.write(admin_login_pages)
    if "admin" in admin_login_pages:
        if "login" in admin_login_pages:
            pages.append(admin_login_pages)
    with open("urls.txt", "r") as fread:
        url_list = [x.strip() for x in fread.readlines()]
        r = requests.get(str(url_list))
        soup = BeautifulSoup(r.content, 'html.parser')
        for title in soup.find_all('title'):
            print(f"{admin_login_pages} - {title.get_text()}")
if not pages:
    print(f"{Fore.RED} No admin or login pages Found")
else:
    for page_list in pages:
        print(f"{Fore.GREEN} {page_list}")
错误原因

核心问题是你把整个URL列表转成字符串传给了requests.get(),而不是遍历列表里的每个单独URL。str(url_list)会生成类似"['url1', 'url2']"的格式,这不是合法的URL结构,requests无法识别对应的连接协议,因此抛出InvalidSchema错误。

解决方案

1. 修正核心请求逻辑

遍历url_list中的每个URL,单独发起请求,而不是把整个列表作为参数传入:

pages = []
for admin_login_pages in domains:
    with open("urls.txt", "w") as f:
        f.write(admin_login_pages)
    # 合并条件判断,简化代码
    if "admin" in admin_login_pages and "login" in admin_login_pages:
        pages.append(admin_login_pages)
    
    with open("urls.txt", "r") as fread:
        url_list = [x.strip() for x in fread.readlines()]
        # 遍历每个URL单独请求
        for url in url_list:
            # 跳过空字符串,避免无效请求
            if not url:
                continue
            try:
                r = requests.get(url)
                r.raise_for_status()  # 主动抛出HTTP状态码错误(如404、500)
                soup = BeautifulSoup(r.content, 'html.parser')
                # 简化标题获取逻辑
                title_text = soup.title.get_text() if soup.title else "无标题"
                print(f"{url} - {title_text}")
            except requests.exceptions.RequestException as e:
                print(f"{url} 请求失败: {str(e)}")

# 翻译提示文本
if not pages:
    print(f"{Fore.RED} 未找到包含admin或login的页面")
else:
    for page_list in pages:
        print(f"{Fore.GREEN} {page_list}")

2. 额外优化建议

  • 去掉冗余文件操作:如果domains本身就是URL列表,完全不需要写入再读取urls.txt,直接遍历domains即可,减少IO开销:
pages = []
for url in domains:
    if "admin" in url and "login" in url:
        pages.append(url)
    
    try:
        r = requests.get(url, timeout=10)  # 添加超时限制,避免无限等待
        r.raise_for_status()
        soup = BeautifulSoup(r.content, 'html.parser')
        title_text = soup.title.string if soup.title else "无标题"
        print(f"{url} - {title_text}")
    except requests.exceptions.RequestException as e:
        print(f"{url} 请求出错: {str(e)}")

if not pages:
    print(f"{Fore.RED} 未找到包含admin或login的页面")
else:
    for page in pages:
        print(f"{Fore.GREEN} {page}")

内容的提问来源于stack exchange,提问作者c0d3ninja

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 12:52:36