You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的BeautifulSoup爬取网页时获取完整商品描述文本?

问题:爬取商品描述仅获取到部分内容

当前输出

about
0  [\n, [Descrizione], \n, [], [Allenamento per p...

我的代码

def perform_search(kodovi):
    i = 0
    c = 0
    for x in kodovi:
        k = requests.get(searchlink).text
        soup=BeautifulSoup(k,'html.parser')
    
        productlist = soup.find_all("ul",{"class":"products-grid"})

        for product in productlist:
            link = product.find("a",{"class":"product-image"}).get('href')
            productlinks.append(link)
            print("Done")


def scraping_data(productlinks,r):
    for link in productlinks:
        f = requests.get(link,headers=headers).text
        hun=BeautifulSoup(f,'html.parser')

#Here I get description of product
        try:
            about=hun.find("div",{"class":"std"})
        except:
            about=None
            print("nothing found")

        try:
            name=hun.find("h1",{"class":"product-main__name"}).text.replace('\n',"")
        except:
            name=None

        whisky = {"about":about}

        data.append(whisky)
        r=r+1
        print("completed",r)

    df = pd.DataFrame(data)

    print(df)

商品描述的HTML结构

<div class="std">
<h2>Descrizione</h2>
<p></p><p>A LOT OF TEXT WITH PRODUCT DESCRIPTION......</p>
</div>

解决方案

你现在的问题是直接把BeautifulSoup返回的Tag对象存进字典了,输出时自然只会显示Tag内部的结构片段,而非完整文本。以下是两种修复方式:

方法1:提取标签内所有文本(含子标签)

修改获取about的代码:

try:
    # 获取div内所有文本,自动清理空白并拼接
    about = hun.find("div", {"class": "std"}).get_text(strip=True, separator=" ")
except AttributeError:
    about = None
    print("nothing found")
  • get_text()会提取当前标签及其所有子标签的文本内容
  • strip=True去除文本首尾的空白字符
  • separator=" "用空格分隔不同子标签的文本,避免内容挤在一起

方法2:仅提取<p>标签内的正文(不含标题)

如果不需要<h2>的标题,只想要正文内容:

try:
    p_tags = hun.find("div", {"class": "std"}).find_all("p")
    # 过滤空的p标签,拼接有效文本
    about = " ".join([p.get_text(strip=True) for p in p_tags if p.get_text(strip=True)])
except AttributeError:
    about = None
    print("nothing found")

额外提示

原代码里的except:捕获范围太广,改成except AttributeError:更精准——只有当find()返回None时,调用get_text()才会触发这个异常,避免误抓其他无关错误。

内容的提问来源于stack exchange,提问作者Want Help

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 07:10:32