You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取企业语言占比失败:数据来自script标签问题

Fixing Language Usage Data Extraction from Zippia Pages

I see the issue here—your current code is targeting a div with class companyEducationDegrees, but that's not where the language usage data lives. Zippia loads this kind of demographic data into JavaScript objects inside <script> tags instead of rendering it directly in the HTML structure. Let's adjust your approach to pull that data correctly.

Step-by-Step Solution

1. Identify the Right Script Tag

First, we need to locate the <script> tag that contains the language distribution data. Zippia stores most of its dynamic data in a global JavaScript variable (usually something like window.__INITIAL_STATE__), which holds a JSON object with all the company details.

2. Extract and Parse the JSON Data

Once we find the correct script, we'll use a regex to pull out the JSON content from the script text, then parse it into a Python dictionary to access the language data.

Full Working Code

import requests
from bs4 import BeautifulSoup
import re
import json

webpage = "https://www.zippia.com/amazon-com-careers-487/"
page = requests.get(webpage)
soup = BeautifulSoup(page.content, 'lxml')

# Find the script tag containing the initial state data
script_tags = soup.find_all('script')
target_script = None
for script in script_tags:
    if 'languageDistribution' in script.text:
        target_script = script.text
        break

if target_script:
    # Extract the JSON part using regex (matches the object after window.__INITIAL_STATE__ = )
    json_match = re.search(r'window\.__INITIAL_STATE__ = (.*?);', target_script)
    if json_match:
        data = json.loads(json_match.group(1))
        
        # Navigate to the language distribution data
        language_data = data.get('company', {}).get('demographics', {}).get('languageDistribution', [])
        
        # Format and print the results
        if language_data:
            print("Language Usage Breakdown:")
            for lang in language_data:
                print(f"{lang['name']} {lang['percentage']}%")
        else:
            print("No language distribution data found.")
    else:
        print("Could not extract JSON data from script tag.")
else:
    print("No script tag containing language data found.")

What This Code Does

  • Locates the correct script: We loop through all script tags and pick the one that mentions languageDistribution (a dead giveaway for our target data).
  • Extracts JSON: The regex window\.__INITIAL_STATE__ = (.*?); grabs the JSON object stored in the global variable, ignoring the surrounding JavaScript syntax.
  • Parses and accesses data: We navigate the parsed JSON to get to the languageDistribution array, then loop through each entry to print the language name and percentage just like your desired output (Spanish 61.1%, French 9.7%, etc).

Notes

  • Zippia might change their variable names or data structure over time, but as of now, this works for their company pages. If it breaks later, check the script tags again to find where the demographic data is stored.
  • Make sure you have all required libraries installed (requests, beautifulsoup4, lxml)—you can install them with pip install requests beautifulsoup4 lxml if needed.

内容的提问来源于stack exchange,提问作者hard1009

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 09:03:16