Python IMDB网页抓取:如何分离导演与演员的纯文本信息
Hey there! Since you're new to Python web scraping, let's walk through how to pull out the pure text for directors and actors from that <p> tag you've retrieved with first.find('p',{'class':''}).
First, let's assume your scraped HTML content looks like one of these common structures (feel free to adjust based on what you actually see):
<p class="">导演: 张艺谋 演员: 邓超, 孙俪, 沈腾</p>Or with line breaks:
<p class="">导演: 李安<br>演员: 汤姆·汉克斯, 海伦·亨特</p>
Here are a few simple, reliable methods to extract the info you need:
Method 1: Split Text Directly (Best for Consistent Labels)
First, grab the full clean text from the tag using .get_text(), then split it using the consistent labels ("导演:" and "演员:") as markers:
# Get the full stripped text from the p element full_text = first.find('p', {'class': ''}).get_text(strip=True) # Split to isolate director and actor sections director = full_text.split('导演:')[1].split('演员:')[0].strip() actors = full_text.split('演员:')[1].strip() # Output the results print(f"导演: {director}") print(f"演员: {actors}")
The strip() method removes any extra spaces, tabs, or newline characters that might clutter the output.
Method 2: Use Regular Expressions (More Flexible)
If the text has minor variations (like extra spaces or inconsistent spacing around labels), regex is a great tool to target the content precisely:
import re # Get the full text content full_text = first.find('p', {'class': ''}).get_text(strip=True) # Match director content: everything after "导演:" up to "演员:" or the end director_match = re.search(r'导演:\s*(.*?)\s*(演员:|$)', full_text) director = director_match.group(1) if director_match else "未找到导演信息" # Match actor content: everything after "演员:" to the end of the text actor_match = re.search(r'演员:\s*(.*)', full_text) actors = actor_match.group(1) if actor_match else "未找到演员信息" print(f"导演: {director}") print(f"演员: {actors}")
This method handles cases where there might be extra spaces or unexpected formatting around the labels.
Method 3: Handle Line Breaks (For <br> Separated Content)
If the director and actor info are on separate lines using <br> tags, split the text by newlines to process each line individually:
# Get text with line breaks preserved (using '\n' as the separator) full_text = first.find('p', {'class': ''}).get_text('\n', strip=True) lines = full_text.split('\n') director = "" actors = "" # Loop through each line to find the relevant sections for line in lines: if '导演:' in line: director = line.split('导演:')[1].strip() elif '演员:' in line: actors = line.split('演员:')[1].strip() print(f"导演: {director}") print(f"演员: {actors}")
This works perfectly when the content is split across lines in the HTML.
Quick Pro Tip
Always print out full_text first to see the exact structure of the content you're working with—this will help you pick the method that fits best. If your labels are in English (like "Director:" or "Actors:"), just swap those strings in the code!
内容的提问来源于stack exchange,提问作者Dikshit Bhanushali

