使用BeautifulSoup抓取h1时报'str' object is not callable如何解决
问题修复方案
你遇到的错误和BS4本身无关,核心是两处基础逻辑错误:
- 变量重复覆盖错误:
get_archive_h1函数中,你先将bsh赋值为BeautifulSoup解析对象,随后又将其重新赋值为bsh.h1.text.strip()的字符串结果,返回时再次尝试对字符串调用.h1属性,直接触发TypeError: 'str' object is not callable报错。 - concurrent.futures API误用:
executor.map返回的是任务执行结果的迭代器,而非Future对象集合,不能传入as_completed遍历,API使用错误导致你观察到执行次数异常、结果返回异常等问题。 - 额外兼容建议:部分页面可能不存在H1标签,
bsh.h1会返回None,直接调用.text会触发NoneType错误,需要增加判空逻辑。
修正后完整代码
import concurrent.futures from bs4 import BeautifulSoup from urllib.request import urlopen CONNECTIONS = 1 archive_url_list = [ "https://web.archive.org/web/20171220015929/http://www.manueldrivingschool.co.uk:80/prices.php", "https://web.archive.org/web/20160313085709/http://www.manueldrivingschool.co.uk/lessons_prices.php", "https://web.archive.org/web/20171220002420/http://www.manueldrivingschool.co.uk:80/prices", "https://web.archive.org/web/20201202094502/https://www.manueldrivingschool.co.uk/success", ] archive_h1_list = [] def get_archive_h1(h1_url): try: html = urlopen(h1_url) bsh = BeautifulSoup(html.read(), 'lxml') h1_tag = bsh.h1 return h1_tag.text.strip() if h1_tag else "无H1标签" except Exception: return "请求/解析失败" def concurrent_calls(): # 若要使用as_completed,需先提交任务获取Future列表 futures = [] with concurrent.futures.ThreadPoolExecutor(max_workers=CONNECTIONS) as executor: for url in archive_url_list: futures.append(executor.submit(get_archive_h1, url)) for future in concurrent.futures.as_completed(futures): try: data = future.result() archive_h1_list.append(data) except Exception: archive_h1_list.append("No Data Received!") if __name__ == '__main__': concurrent_calls() print(archive_h1_list)
如果你不需要异步获取结果顺序,直接用executor.map更简洁,替换concurrent_calls函数即可:
def concurrent_calls(): with concurrent.futures.ThreadPoolExecutor(max_workers=CONNECTIONS) as executor: # map按提交顺序返回结果,直接遍历即可 for result in executor.map(get_archive_h1, archive_url_list): archive_h1_list.append(result)
内容的提问来源于stack exchange,提问作者Lee Roy
相关产品推荐
相关产品推荐

