You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python(bs4)将HTML标签转为字典结构化内容?爬取电影页遇阻

爬取Empire电影榜单时的描述提取问题

问题背景

爬取Wayback Machine存档的Empire杂志最佳电影榜单页面时,提取电影描述环节遇到阻碍:目标描述元素无专属类名,仅通过带类名的父节点和兄弟节点关联,期望将描述中的HTML标签内容转换为字典内的结构化块内容,曾尝试Flask CKEditor未解决该问题。

现有爬取代码

# Imports

from bs4 import BeautifulSoup
from requests import *

# Flask

from flask import Flask, render_template

# Scraper

URL = "https://web.archive.org/web/20200518073855/https://www.empireonline.com/movies/features/best-movies-2/"
res = get(URL)
html_data = res.text

soup = BeautifulSoup(html_data, 'html.parser')
film_data = soup.find_all(name='div', class_='article-title-description__text')

total_films = []

for film in film_data:
    name = str(film.find('h3', class_='title')).split(')')[-1].split(':')[-1].split('</h3>')[0]
    year = str(film.find('strong')).split('<strong>')[-1].split('</strong>')[0]
    image= film.find('img', class_='landscape')
    # print(image)
    desc = film.findChildren('p')
    for strong_tag in desc[1].find_all('a'):
        desc__txt = strong_tag.text, strong_tag.next_sibling
    new_obj = {
        "name": name,
        "year": year,
        'image': image,
        "desc": desc
    }
    total_films.append(new_obj)

# Flask

app = Flask(__name__)

@app.route('/')
def show_data():
    return render_template('index.html', total_films=total_films)

if __name__ == '__main__':
    app.run(debug=True, port=5100)

当前输出结果

<h1>Movie Scraper</h1>
    <h1>Stand By Me</h1>
    <h2>1986</h2>
    [<p class="description"><p><strong>1986</strong><br/><a href="https://web.archive.org/web/20200518073855/https://www.empireonline.com/people/rob-reiner/">Rob Reiner</a>'s adaptation of <a href="https://web.archive.org/web/20200518073855/https://www.empireonline.com/people/stephen-king/">Stephen King</a>'s novella The Body is a stirring, touching adventure film which knows the real world is exciting and scary enough just as it is. It's also a coming-of-age movie which celebrates the intensity of childhood friendship, while gently mourning the transience of such bonds. Which is why, unlike its central character, it'll never get old.<br/><br/><a href="https://web.archive.org/web/20200518073855/https://www.empireonline.com/movies/reviews/stand-review/">Read Empire's review of Stand By Me</a><br/><a class="amazon-link" href="https://web.archive.org/web/20200518073855/https://www.amazon.co.uk/gp/product/B003BNY6YE/ref=as_li_tl?ie=UTF8&amp;tag=baucitnet-21&amp;camp=1634&amp;creative=6738&amp;linkCode=as2&amp;creativeASIN=B003BNY6YE&amp;linkId=5dd32860f2c5126a8a53f2668f114778" rel="nofollow">Buy the film now</a><br/></p> </p>, <p><strong>1986</strong><br/><a href="https://web.archive.org/web/20200518073855/https://www.empireonline.com/people/rob-reiner/">Rob Reiner</a>'s adaptation of <a href="https://web.archive.org/web/20200518073855/https://www.empireonline.com/people/stephen-king/">Stephen King</a>'s novella The Body is a stirring, touching adventure film which knows the real world is exciting and scary enough just as it is. It's also a coming-of-age movie which celebrates the intensity of childhood friendship, while gently mourning the transience of such bonds. Which is why, unlike its central character, it'll never get old.<br/><br/><a href="https://web.archive.org/web/20200518073855/https://www.empireonline.com/movies/reviews/stand-review/">Read Empire's review of Stand By Me</a><br/><a class="amazon-link" href="https://web.archive.org/web/20200518073855/https://www.amazon.co.uk/gp/product/B003BNY6YE/ref=as_li_tl?ie=UTF8&amp;tag=baucitnet-21&amp;camp=1634&amp;creative=6738&amp;linkCode=as2&amp;creativeASIN=B003BNY6YE&amp;linkId=5dd32860f2c5126a8a53f2668f114778" rel="nofollow">Buy the film now</a><br/></p>]

解决方案

核心问题在于原代码直接存储了BeautifulSoup节点对象,且未准确定位描述父节点,导致输出重复且非结构化。以下是优化后的代码,重点处理描述提取和结构化:

优化后的完整代码

from bs4 import BeautifulSoup
from requests import get
from flask import Flask, render_template

URL = "https://web.archive.org/web/20200518073855/https://www.empireonline.com/movies/features/best-movies-2/"
res = get(URL)
html_data = res.text

soup = BeautifulSoup(html_data, 'html.parser')
film_data = soup.find_all('div', class_='article-title-description__text')

total_films = []

for film in film_data:
    # 更可靠地提取电影名称
    title_tag = film.find('h3', class_='title')
    title_text = title_tag.get_text(strip=True)
    name = title_text.split(')')[-1].strip()
    
    # 提取年份
    year_tag = film.find('strong')
    year = year_tag.get_text(strip=True) if year_tag else "Unknown"
    
    # 提取图片链接
    image_tag = film.find('img', class_='landscape')
    image_url = image_tag.get('src') if image_tag else None
    
    # 提取并结构化描述内容
    desc_container = film.find('p', class_='description')
    structured_desc = {
        "main_text": "",
        "related_links": []
    }
    
    if desc_container:
        # 遍历描述内的所有子节点,分离文本和链接
        for child in desc_container.children:
            if isinstance(child, str):
                # 处理文本节点,清理多余空格
                structured_desc["main_text"] += child.strip() + " "
            elif child.name == 'a':
                # 处理链接节点,保存文本和地址
                structured_desc["related_links"].append({
                    "text": child.get_text(strip=True),
                    "url": child.get('href')
                })
        # 清理主文本的尾部空格
        structured_desc["main_text"] = structured_desc["main_text"].strip()
    else:
        structured_desc["main_text"] = "No description available"

    total_films.append({
        "name": name,
        "year": year,
        "image": image_url,
        "desc": structured_desc
    })

# Flask 部分保持不变
app = Flask(__name__)

@app.route('/')
def show_data():
    return render_template('index.html', total_films=total_films)

if __name__ == '__main__':
    app.run(debug=True, port=5100)

关键改动说明

  1. 准确定位描述父节点:直接通过film.find('p', class_='description')定位描述容器,避免获取嵌套的重复p标签。
  2. 结构化描述内容:将描述拆分为main_text(纯文本内容)和related_links(关联链接列表,含文本和地址),方便模板渲染时灵活调用。
  3. 优化基础信息提取:
    • 电影名称改用get_text()提取后拆分,避免直接处理HTML字符串的脆弱性。
    • 图片提取src属性值而非整个节点对象,便于在Flask模板中直接作为图片链接使用。
  4. 异常处理:对可能不存在的标签(如年份、图片、描述)添加判断,避免运行报错。

模板渲染示例(index.html)

若要在模板中展示结构化内容,可参考以下写法:

<h1>Movie Scraper</h1>
{% for movie in total_films %}
    <h2>{{ movie.name }}</h2>
    <h3>{{ movie.year }}</h3>
    {% if movie.image %}
        <img src="{{ movie.image }}" alt="{{ movie.name }}">
    {% endif %}
    <p>{{ movie.desc.main_text }}</p>
    {% if movie.desc.related_links %}
        <h4>Related Links:</h4>
        <ul>
            {% for link in movie.desc.related_links %}
                <li><a href="{{ link.url }}">{{ link.text }}</a></li>
            {% endfor %}
        </ul>
    {% endif %}
{% endfor %}

内容的提问来源于stack exchange,提问作者mentossN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 04:27:25