You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup遇TypeError,是否由多线程引发?

问题分析与修复方案

哎,我看了你的问题和代码,马上就发现问题所在了——你遇到的TypeError: 'NoneType' object is not subscriptable,核心原因是有些页面根本没有<meta name="keywords">标签,这时候soup.find()会返回None,你直接去访问['content']肯定会触发错误!

你看到能打印出关键词,是那些存在该标签的URL正常执行到了print(keywords1)语句,但一旦遇到无标签的URL,报错会直接打断当前线程的worker函数,导致后续写入data的逻辑根本跑不完,甚至线程直接挂掉。

下面是具体的修复步骤和完整代码:

关键问题拆解

  1. 无meta标签时的空值访问:当页面没有目标meta标签,soup.find()返回None,直接链式访问['content']会触发类型错误。
  2. 缺乏网络请求异常处理:网络请求可能超时、返回HTTP错误(比如404),这些情况会直接导致线程崩溃。
  3. 锁的使用不够安全:直接用mutex.acquire()和mutex.release(),如果中间出现异常,锁可能无法释放,导致其他线程阻塞。

修复后的完整代码

from threading import Thread, Lock
from bs4 import BeautifulSoup
import requests

mutex = Lock()
threads = []
data = []
websitelist = [
    'http://www.wilmsmetall.de/',
    'https://www.prokilo.de/maerkte/koeln-muelheim/',
    'http://www.droesser.de/'
]

class ReqThread(Thread):
    def __init__(self, threadID, u_list):
        Thread.__init__(self)
        self.threadID = threadID
        self.urllist = u_list

    def run(self):
        print(f"Starting {self.threadID}")
        for url in self.urllist:
            worker(url)
        print(f"Exiting {self.threadID}")

def worker(url):
    global data
    try:
        # 捕获网络请求的所有异常,避免线程崩溃
        html = requests.get(url, timeout=3)
        html.raise_for_status()  # 触发HTTP错误(如404、500)的异常
        html_page = html.content
        soup = BeautifulSoup(html_page, 'html.parser')
        
        # 先判断meta标签是否存在,再提取内容
        meta_keywords = soup.find("meta", {"name":"keywords"})
        if meta_keywords:
            # 用get方法更安全,避免content属性不存在的情况
            keywords1 = meta_keywords.get('content', '')
            if keywords1:
                print(f"Extracted keywords from {url}: {keywords1}")
                # 拆分关键词并写入data,用with自动管理锁
                for keyword in keywords1.split(','):
                    cleaned_keyword = keyword.strip()
                    if cleaned_keyword:  # 过滤空的关键词
                        with mutex:
                            data.append(cleaned_keyword)
        else:
            print(f"No meta keywords found for {url}")
    except Exception as e:
        print(f"Error processing {url}: {str(e)}")
        return

# 启动所有线程
for index, website in enumerate(websitelist):
    t = ReqThread(index, [website])
    t.start()
    threads.append(t)

# 等待所有线程执行完成
for thread in threads:
    thread.join()

print("\nFinal collected keywords:")
print(data)

核心修改说明

  • 空值安全处理:先判断meta_keywords是否存在,再用get('content', '')提取内容,避免直接下标访问的错误。
  • 网络异常捕获:用try-except包裹请求逻辑,即使某个URL请求失败,线程也不会崩溃,能继续处理其他任务。
  • 锁的安全管理:用with mutex:替代手动的acquire()和release(),确保锁一定会被释放,不会出现死锁。
  • 关键词过滤:增加了空关键词的过滤,避免把空字符串加入data列表。

内容的提问来源于stack exchange,提问作者LGR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 09:52:38