Selenium爬取电影信息:如何在Pandas中为每行添加对应电影类型
问题解决:为Pandas DataFrame添加对应电影的类型列
需求说明
现有包含电影链接的Pandas DataFrame,需要遍历每个链接获取对应电影的类型(数量不固定,可能多个),并在DataFrame中新增一列存储对应行的类型列表。
原始问题
当前代码运行后得到的是所有电影类型的扁平列表,无法对应到每个电影的行。
修改后的代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd data = {"link":["http://www.boxofficemojo.com/movies/?id=ateam.htm", "http://www.boxofficemojo.com/movies/?id=acod.htm","http://www.boxofficemojo.com/movies/?id=ai.htm", "http://www.boxofficemojo.com/movies/?id=axl.htm","http://www.boxofficemojo.com/movies/?id=aaa.htm"]} dataframe = pd.DataFrame(data) # 初始化存储每个电影类型的列表(每个元素是对应电影的类型列表) genre_list = [] # 只初始化一次driver,提升效率 driver = webdriver.Chrome("C:\SeleniumDrivers\chromedriver.exe") for link in dataframe['link']: driver.get(link) # 获取类型元素的HTML内容 tag = WebDriverWait(driver, 20).until(EC.visibility_of_element_located((By.XPATH, "//span[text()='Genres']//following::span[1]"))).get_attribute("innerHTML") # 处理当前电影的类型 current_genres = [] for item in tag.split("\n"): cleaned_item = item.strip() if len(cleaned_item) > 1: # 过滤空字符串和无效内容 current_genres.append(cleaned_item) # 将当前电影的类型列表加入总列表 genre_list.append(current_genres) # 关闭driver driver.quit() # 为DataFrame新增列 dataframe['genres'] = genre_list print(dataframe)
关键修改点
- 存储结构调整:不再将所有类型扁平追加到同一个列表,而是为每个电影创建临时列表
current_genres,收集完当前电影的所有类型后,将这个列表加入总列表genre_list,确保每个元素对应一行电影的类型集合。 - Driver优化:只初始化一次Chrome Driver,循环结束后再关闭,避免频繁创建销毁浏览器实例,提升运行效率。
- 简化类型处理逻辑:去掉冗余的嵌套循环,直接拆分后清洗内容,过滤无效项。
效果示例
最终DataFrame的genres列会存储对应电影的类型列表,例如:
| link | genres |
|---|---|
| http://www.boxofficemojo.com/movies/?id=ateam.htm | ['Action', 'Adventure', 'Thriller'] |
| ... | ... |
内容的提问来源于stack exchange,提问作者Almosino
相关产品推荐
相关产品推荐

