You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多线程批量检测URL meta内容报AttributeError问题求助

问题根因与修复方案

报错直接原因

你调用executor.map(networkCall, sites, headers)时,map方法会按位置依次从后面的可迭代对象里取参数传给目标函数:第一个参数从sites取URL,第二个参数从headers取。而headers是字典,直接迭代字典只能拿到键(也就是字符串'User-Agent'),等于你传给networkCall的第二个参数是字符串,当requests.get接收字符串类型的headers参数时,内部要调用headers.items()遍历请求头,自然就报'str' object has no attribute 'items'的错误。

其他需同步修复的隐藏问题

  • networkCall返回的是原始字节流response.content,getMeta里直接调用find_all方法会报错,需要先把字节流解析为BeautifulSoup对象,同时要导入bs4库的BeautifulSoup
  • main函数里你手动把sites赋值为测试URL列表,会覆盖之前从CSV读取的URL数据,实际运行时需要删除这段测试代码
  • getMeta里的判断逻辑写反了:你现有逻辑是检测到非空的meta内容就标记Absent,和你「检测是否存在指定meta内容」的需求相反
  • 全局变量out有线程安全风险,直接放在getMeta内部生成更稳妥
  • 新增了超时控制和异常捕获,避免个别请求超时、报错卡住整个流程

修复后完整代码

import concurrent.futures
import requests
import threading
import time
import pandas as pd
import requests_cache
from bs4 import BeautifulSoup
import itertools

thread_local = threading.local()

# 读取CSV中的URL
df = pd.read_csv("test.csv")
sites = df['URLS'].tolist()

user_agent = 'Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.9.0.7) Gecko/2009021910 Firefox/3.0.7'
headers = {'User-Agent': user_agent}

requests_cache.install_cache('network_call', backend='sqlite', expire_after=2592000)

def getSess():
    if not hasattr(thread_local, "session"):
        thread_local.session = requests.Session()
    return thread_local.session

def networkCall(url, headers):
    session = getSess()
    try:
        with session.get(url, headers=headers, timeout=10) as response:
            print(f"Read {len(response.content)} from {url}")
            # 直接返回解析后的BeautifulSoup对象,避免后续重复解析
            return BeautifulSoup(response.content, 'html.parser')
    except Exception as e:
        print(f"请求{url}失败:{str(e)}")
        return None

def getMeta(meta_res):
    out = []
    for each in meta_res:
        has_valid_meta = False
        if each is None:
            out.append("请求失败")
            continue
        meta = each.find_all('meta')
        for tag in meta:
            if 'name' in tag.attrs and tag.attrs['name'].strip().lower() in ['description', 'keywords']:
                content = tag.attrs.get('content', '').strip()
                if content:
                    has_valid_meta = True
                    break
        # 逻辑修正:存在有效meta内容标记为Present,否则为Absent
        out.append("Present" if has_valid_meta else "Absent")
    return out

def allSites(sites, headers):
    with concurrent.futures.ThreadPoolExecutor(max_workers=50) as executor:
        # 用itertools.repeat把headers重复为和sites长度一致的可迭代对象,保证每个请求都拿到完整的headers字典
        out = executor.map(networkCall, sites, itertools.repeat(headers, len(sites)))
        return list(out)

if __name__ == "__main__":
    # 测试时可以保留下面这段,实际跑CSV数据时删除即可
    # sites = [
    # "https://www.jython.org",
    # "http://olympus.realpython.org/dice",
    # ] * 10
    start_time = time.time()
    list_meta = allSites(sites, headers)
    duration = time.time() - start_time
    print(f"处理完{len(sites)}个URL,耗时{duration}秒")
    output = getMeta(list_meta)
    df["is it there"] = pd.Series(output)
    df.to_csv('new.csv', index=False, header=True)

内容的提问来源于stack exchange,提问作者chixcy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 13:45:07