Python多线程批量检测URL meta内容报AttributeError问题求助
问题根因与修复方案
报错直接原因
你调用executor.map(networkCall, sites, headers)时,map方法会按位置依次从后面的可迭代对象里取参数传给目标函数:第一个参数从sites取URL,第二个参数从headers取。而headers是字典,直接迭代字典只能拿到键(也就是字符串'User-Agent'),等于你传给networkCall的第二个参数是字符串,当requests.get接收字符串类型的headers参数时,内部要调用headers.items()遍历请求头,自然就报'str' object has no attribute 'items'的错误。
其他需同步修复的隐藏问题
networkCall返回的是原始字节流response.content,getMeta里直接调用find_all方法会报错,需要先把字节流解析为BeautifulSoup对象,同时要导入bs4库的BeautifulSoupmain函数里你手动把sites赋值为测试URL列表,会覆盖之前从CSV读取的URL数据,实际运行时需要删除这段测试代码getMeta里的判断逻辑写反了:你现有逻辑是检测到非空的meta内容就标记Absent,和你「检测是否存在指定meta内容」的需求相反- 全局变量
out有线程安全风险,直接放在getMeta内部生成更稳妥 - 新增了超时控制和异常捕获,避免个别请求超时、报错卡住整个流程
修复后完整代码
import concurrent.futures import requests import threading import time import pandas as pd import requests_cache from bs4 import BeautifulSoup import itertools thread_local = threading.local() # 读取CSV中的URL df = pd.read_csv("test.csv") sites = df['URLS'].tolist() user_agent = 'Mozilla/5.0 (Windows; U; Windows NT 5.1; en-US; rv:1.9.0.7) Gecko/2009021910 Firefox/3.0.7' headers = {'User-Agent': user_agent} requests_cache.install_cache('network_call', backend='sqlite', expire_after=2592000) def getSess(): if not hasattr(thread_local, "session"): thread_local.session = requests.Session() return thread_local.session def networkCall(url, headers): session = getSess() try: with session.get(url, headers=headers, timeout=10) as response: print(f"Read {len(response.content)} from {url}") # 直接返回解析后的BeautifulSoup对象,避免后续重复解析 return BeautifulSoup(response.content, 'html.parser') except Exception as e: print(f"请求{url}失败:{str(e)}") return None def getMeta(meta_res): out = [] for each in meta_res: has_valid_meta = False if each is None: out.append("请求失败") continue meta = each.find_all('meta') for tag in meta: if 'name' in tag.attrs and tag.attrs['name'].strip().lower() in ['description', 'keywords']: content = tag.attrs.get('content', '').strip() if content: has_valid_meta = True break # 逻辑修正:存在有效meta内容标记为Present,否则为Absent out.append("Present" if has_valid_meta else "Absent") return out def allSites(sites, headers): with concurrent.futures.ThreadPoolExecutor(max_workers=50) as executor: # 用itertools.repeat把headers重复为和sites长度一致的可迭代对象,保证每个请求都拿到完整的headers字典 out = executor.map(networkCall, sites, itertools.repeat(headers, len(sites))) return list(out) if __name__ == "__main__": # 测试时可以保留下面这段,实际跑CSV数据时删除即可 # sites = [ # "https://www.jython.org", # "http://olympus.realpython.org/dice", # ] * 10 start_time = time.time() list_meta = allSites(sites, headers) duration = time.time() - start_time print(f"处理完{len(sites)}个URL,耗时{duration}秒") output = getMeta(list_meta) df["is it there"] = pd.Series(output) df.to_csv('new.csv', index=False, header=True)
内容的提问来源于stack exchange,提问作者chixcy
相关产品推荐
相关产品推荐

