网页抓取:如何从Box Office页面相同HTML标签提取不同值?
提取Box Office Mojo页面同标签下不同数值的方法
针对你遇到的属性名和对应值使用相同HTML标签的情况,可通过以下几种方式精准提取数据:
1. 基于元素位置成对提取(BeautifulSoup)
页面中属性名和对应值通常连续成对出现,可将所有目标标签元素按顺序每两个分为一组,分别作为键和值:
from bs4 import BeautifulSoup import requests url = "https://www.boxofficemojo.com/title/tt1205489/?ref_=bo_cso_table_200" response = requests.get(url, headers={"User-Agent": "Mozilla/5.0"}) soup = BeautifulSoup(response.text, 'html.parser') # 替换为页面中实际的目标标签及类名 target_elements = soup.find_all('span', class_='a-size-medium') # 按成对关系提取数据 result = {} for idx in range(0, len(target_elements), 2): if idx + 1 < len(target_elements): attr_name = target_elements[idx].get_text(strip=True) attr_value = target_elements[idx+1].get_text(strip=True) result[attr_name] = attr_value # 示例:获取Running Time的值 print(result.get("Running Time"))
2. 基于父容器分组提取
观察页面结构,每个属性-值对通常会被一个父容器(如div、tr)包裹,先定位父容器,再在容器内提取对应元素:
# 替换为实际的父容器标签及类名 containers = soup.find_all('div', class_='a-section a-spacing-small') result = {} for container in containers: items = container.find_all('span') # 替换为实际的相同标签 if len(items) >= 2: attr_name = items[0].get_text(strip=True) attr_value = items[1].get_text(strip=True) result[attr_name] = attr_value
3. 利用XPath定位(lxml/Scrapy)
通过XPath的位置索引或轴关系,直接区分属性名和对应值:
from lxml import html tree = html.fromstring(response.content) # 提取所有属性名 attr_names = tree.xpath("//div[@class='a-section a-spacing-small']/span[1]/text()") # 提取所有对应值 attr_values = tree.xpath("//div[@class='a-section a-spacing-small']/span[2]/text()") result = dict(zip(attr_names, attr_values)) print(result.get("Running Time"))
注意事项
- 需先用浏览器开发者工具(F12)确认页面真实的HTML结构,替换代码中的标签、类名等定位符
- 提取文本时使用
strip=True去除多余的空格、换行符 - 避免频繁请求页面,可添加请求头、设置请求延迟以规避反爬限制
内容的提问来源于stack exchange,提问作者Khaled Hamdy
相关产品推荐
相关产品推荐

