如何用BeautifulSoup从网页代码中提取特定词组?附已获Script代码
用BeautifulSoup提取网页中特定词组的方法
嘿,我来帮你搞定这个问题!从你给出的内容来看,你已经拿到了包含LifeTimeStats数据的script标签,接下来只需要几步就能提取到你要的特定词组:
步骤1:定位并提取script标签的文本内容
首先,你需要从BeautifulSoup对象中精准定位到包含LifeTimeStats的script标签,然后提取它的文本内容:
# 假设你的BeautifulSoup对象名为soup script_tag = soup.find('script', text=lambda text: text and 'LifeTimeStats' in text) if script_tag: script_content = script_tag.string else: print("未找到包含LifeTimeStats的script标签")
步骤2:分离并解析JSON数据
script标签里的内容是JS变量定义,我们需要把里面的JSON部分单独截取出来,再转换成Python可操作的数据结构:
import json # 截取JSON字符串(去掉开头的"var LifeTimeStats = "和结尾的分号) json_str = script_content.split('var LifeTimeStats = ')[1].rstrip(';') # 解析JSON为Python列表 life_time_stats = json.loads(json_str)
步骤3:提取特定词组
现在life_time_stats是一个包含字典的列表,你可以遍历它来提取目标Key对应的Value:
方式1:遍历查找单个目标
# 比如提取"Wins"(获胜次数)的值 for stat in life_time_stats: if stat['Key'] == 'Wins': print(f"获胜次数: {stat['Value']}") break
方式2:封装成函数批量获取
如果需要多次提取不同的Key,可以写个小函数提高效率:
def get_stat(stats_list, target_key): for stat in stats_list: if stat['Key'] == target_key: return stat['Value'] return "未找到该统计项" # 示例:获取Kills(击杀数)和Score(总分) kills = get_stat(life_time_stats, 'Kills') score = get_stat(life_time_stats, 'Score') print(f"击杀数: {kills}, 总分: {score}")
额外提示:异常处理
为了避免因为网页结构变化导致代码报错,建议加上简单的异常处理:
try: script_tag = soup.find('script', text=lambda text: text and 'LifeTimeStats' in text) if not script_tag: raise ValueError("未找到目标script标签") script_content = script_tag.string json_str = script_content.split('var LifeTimeStats = ')[1].rstrip(';') life_time_stats = json.loads(json_str) wins = get_stat(life_time_stats, 'Wins') print(f"获胜次数: {wins}") except (ValueError, IndexError, json.JSONDecodeError) as e: print(f"提取失败: {str(e)}")
内容的提问来源于stack exchange,提问作者Sh0x
相关产品推荐
相关产品推荐

