Python BeautifulSoup报错TypeError: find()不接受关键字参数求解决
问题原因
你代码里的read_news函数接收的参数是字符串类型的网址,但函数里直接用link.find(...)调用了字符串的find方法——字符串的find只能用来查找子串,不支持class_这类关键字参数,这就是报错的根源。
修正后的代码
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin # 用来安全拼接URL def read_news(news_url): news_detail = {} # 先请求新闻详情页,解析成BeautifulSoup对象 res = requests.get(news_url) soup = BeautifulSoup(res.content, "html.parser") # 提取标题并获取文本内容 title_tag = soup.find("h1", class_='post_title') news_detail['title'] = title_tag.get_text(strip=True) if title_tag else "无标题" # 提取正文内容并获取文本 content_tag = soup.find("span", class_='news_text') news_detail['content'] = content_tag.get_text(strip=True) if content_tag else "无内容" print(news_detail) base_url = 'https://www.baki-xeber.com' r = requests.get(base_url) bs4 = BeautifulSoup(r.content, "html.parser") elements = bs4.find_all("div", class_='post_content_without_pic') for element in elements: a_tag = element.find("a") if a_tag and a_tag.get("href"): # 用urljoin拼接绝对URL,避免路径错误 news_link = urljoin(base_url, a_tag.get("href")) read_news(news_link)
关键修改点
- 在
read_news内部新增了请求详情页并解析的逻辑,确保调用find的是BeautifulSoup的解析对象,而非字符串。 - 增加了
if xxx else的判断,避免找不到元素时抛出异常。 - 使用
urljoin拼接URL,比直接字符串拼接更可靠(比如处理href已经是绝对路径的情况)。 - 用
.get_text(strip=True)获取元素的纯文本内容,而不是返回Tag对象。
内容的提问来源于stack exchange,提问作者TheFlee
相关产品推荐
相关产品推荐

