You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python IMDB网页抓取:如何分离导演与演员的纯文本信息

Extracting Director and Actor Text from Scraped HTML

Hey there! Since you're new to Python web scraping, let's walk through how to pull out the pure text for directors and actors from that <p> tag you've retrieved with first.find('p',{'class':''}).

First, let's assume your scraped HTML content looks like one of these common structures (feel free to adjust based on what you actually see):

<p class="">导演: 张艺谋 演员: 邓超, 孙俪, 沈腾</p>

Or with line breaks:
<p class="">导演: 李安<br>演员: 汤姆·汉克斯, 海伦·亨特</p>

Here are a few simple, reliable methods to extract the info you need:

Method 1: Split Text Directly (Best for Consistent Labels)

First, grab the full clean text from the tag using .get_text(), then split it using the consistent labels ("导演:" and "演员:") as markers:

# Get the full stripped text from the p element
full_text = first.find('p', {'class': ''}).get_text(strip=True)

# Split to isolate director and actor sections
director = full_text.split('导演:')[1].split('演员:')[0].strip()
actors = full_text.split('演员:')[1].strip()

# Output the results
print(f"导演: {director}")
print(f"演员: {actors}")

The strip() method removes any extra spaces, tabs, or newline characters that might clutter the output.

Method 2: Use Regular Expressions (More Flexible)

If the text has minor variations (like extra spaces or inconsistent spacing around labels), regex is a great tool to target the content precisely:

import re

# Get the full text content
full_text = first.find('p', {'class': ''}).get_text(strip=True)

# Match director content: everything after "导演:" up to "演员:" or the end
director_match = re.search(r'导演:\s*(.*?)\s*(演员:|$)', full_text)
director = director_match.group(1) if director_match else "未找到导演信息"

# Match actor content: everything after "演员:" to the end of the text
actor_match = re.search(r'演员:\s*(.*)', full_text)
actors = actor_match.group(1) if actor_match else "未找到演员信息"

print(f"导演: {director}")
print(f"演员: {actors}")

This method handles cases where there might be extra spaces or unexpected formatting around the labels.

Method 3: Handle Line Breaks (For <br> Separated Content)

If the director and actor info are on separate lines using <br> tags, split the text by newlines to process each line individually:

# Get text with line breaks preserved (using '\n' as the separator)
full_text = first.find('p', {'class': ''}).get_text('\n', strip=True)
lines = full_text.split('\n')

director = ""
actors = ""

# Loop through each line to find the relevant sections
for line in lines:
    if '导演:' in line:
        director = line.split('导演:')[1].strip()
    elif '演员:' in line:
        actors = line.split('演员:')[1].strip()

print(f"导演: {director}")
print(f"演员: {actors}")

This works perfectly when the content is split across lines in the HTML.

Quick Pro Tip

Always print out full_text first to see the exact structure of the content you're working with—this will help you pick the method that fits best. If your labels are in English (like "Director:" or "Actors:"), just swap those strings in the code!

内容的提问来源于stack exchange,提问作者Dikshit Bhanushali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 16:42:53