如何用Python的BeautifulSoup爬取网页时获取完整商品描述文本?
问题:爬取商品描述仅获取到部分内容
当前输出
about 0 [\n, [Descrizione], \n, [], [Allenamento per p...
我的代码
def perform_search(kodovi): i = 0 c = 0 for x in kodovi: k = requests.get(searchlink).text soup=BeautifulSoup(k,'html.parser') productlist = soup.find_all("ul",{"class":"products-grid"}) for product in productlist: link = product.find("a",{"class":"product-image"}).get('href') productlinks.append(link) print("Done") def scraping_data(productlinks,r): for link in productlinks: f = requests.get(link,headers=headers).text hun=BeautifulSoup(f,'html.parser') #Here I get description of product try: about=hun.find("div",{"class":"std"}) except: about=None print("nothing found") try: name=hun.find("h1",{"class":"product-main__name"}).text.replace('\n',"") except: name=None whisky = {"about":about} data.append(whisky) r=r+1 print("completed",r) df = pd.DataFrame(data) print(df)
商品描述的HTML结构
<div class="std"> <h2>Descrizione</h2> <p></p><p>A LOT OF TEXT WITH PRODUCT DESCRIPTION......</p> </div>
解决方案
你现在的问题是直接把BeautifulSoup返回的Tag对象存进字典了,输出时自然只会显示Tag内部的结构片段,而非完整文本。以下是两种修复方式:
方法1:提取标签内所有文本(含子标签)
修改获取about的代码:
try: # 获取div内所有文本,自动清理空白并拼接 about = hun.find("div", {"class": "std"}).get_text(strip=True, separator=" ") except AttributeError: about = None print("nothing found")
get_text()会提取当前标签及其所有子标签的文本内容strip=True去除文本首尾的空白字符separator=" "用空格分隔不同子标签的文本,避免内容挤在一起
方法2:仅提取<p>标签内的正文(不含标题)
如果不需要<h2>的标题,只想要正文内容:
try: p_tags = hun.find("div", {"class": "std"}).find_all("p") # 过滤空的p标签,拼接有效文本 about = " ".join([p.get_text(strip=True) for p in p_tags if p.get_text(strip=True)]) except AttributeError: about = None print("nothing found")
额外提示
原代码里的except:捕获范围太广,改成except AttributeError:更精准——只有当find()返回None时,调用get_text()才会触发这个异常,避免误抓其他无关错误。
内容的提问来源于stack exchange,提问作者Want Help
相关产品推荐
相关产品推荐

