使用BeautifulSoup解析HTML:无法提取完整href字符串问题
问题解决:获取完整的href链接而非单个字符
你的代码之所以会把href拆成单个字符,核心问题是错误地遍历了href字符串的每个字符,同时对BeautifulSoup的返回值处理逻辑有误:
soup.find('a', class_="kib-product-title")返回的是单个<a>标签对象,而非列表。循环for items in test其实是在遍历这个a标签内部的子节点(比如文本节点)。- 当你调用
items.get("href")时,子节点本身没有href属性,若意外拿到href字符串,for link in items.get("href")会把字符串当成可迭代对象,逐个取出每个字符。
修正方案
场景1:只需要单个目标链接
直接从找到的a标签中提取href即可,不需要嵌套循环:
data = requests.get("https://www.chewy.com/b/food_c332_p2", auth=('user', 'pass'), headers={'User-Agent': user_agent}) with open("dogfoodpage/dg2.html","w+") as f: f.write(data.text) with open("dogfoodpage/dg2.html") as f: page = f.read() soup = BeautifulSoup(page,"html.parser") test = soup.find('a', class_="kib-product-title") productlink = [] if test: # 先判断是否找到标签,避免报错 full_link = test.get("href") productlink.append(full_link)
场景2:需要页面中所有符合class的a标签链接
改用find_all获取所有匹配的a标签,再逐个提取href:
data = requests.get("https://www.chewy.com/b/food_c332_p2", auth=('user', 'pass'), headers={'User-Agent': user_agent}) with open("dogfoodpage/dg2.html","w+") as f: f.write(data.text) with open("dogfoodpage/dg2.html") as f: page = f.read() soup = BeautifulSoup(page,"html.parser") test_links = soup.find_all('a', class_="kib-product-title") productlink = [] for link_tag in test_links: full_link = link_tag.get("href") if full_link: # 过滤没有href的标签 productlink.append(full_link)
额外优化建议
不需要把页面内容写入本地文件再读取,直接用data.text传给BeautifulSoup即可,减少IO操作:
data = requests.get("https://www.chewy.com/b/food_c332_p2", auth=('user', 'pass'), headers={'User-Agent': user_agent}) soup = BeautifulSoup(data.text,"html.parser")
内容的提问来源于stack exchange,提问作者jenie
相关产品推荐
相关产品推荐

