使用BeautifulSoup提取企业语言占比失败:数据来自script标签问题
I see the issue here—your current code is targeting a div with class companyEducationDegrees, but that's not where the language usage data lives. Zippia loads this kind of demographic data into JavaScript objects inside <script> tags instead of rendering it directly in the HTML structure. Let's adjust your approach to pull that data correctly.
Step-by-Step Solution
1. Identify the Right Script Tag
First, we need to locate the <script> tag that contains the language distribution data. Zippia stores most of its dynamic data in a global JavaScript variable (usually something like window.__INITIAL_STATE__), which holds a JSON object with all the company details.
2. Extract and Parse the JSON Data
Once we find the correct script, we'll use a regex to pull out the JSON content from the script text, then parse it into a Python dictionary to access the language data.
Full Working Code
import requests from bs4 import BeautifulSoup import re import json webpage = "https://www.zippia.com/amazon-com-careers-487/" page = requests.get(webpage) soup = BeautifulSoup(page.content, 'lxml') # Find the script tag containing the initial state data script_tags = soup.find_all('script') target_script = None for script in script_tags: if 'languageDistribution' in script.text: target_script = script.text break if target_script: # Extract the JSON part using regex (matches the object after window.__INITIAL_STATE__ = ) json_match = re.search(r'window\.__INITIAL_STATE__ = (.*?);', target_script) if json_match: data = json.loads(json_match.group(1)) # Navigate to the language distribution data language_data = data.get('company', {}).get('demographics', {}).get('languageDistribution', []) # Format and print the results if language_data: print("Language Usage Breakdown:") for lang in language_data: print(f"{lang['name']} {lang['percentage']}%") else: print("No language distribution data found.") else: print("Could not extract JSON data from script tag.") else: print("No script tag containing language data found.")
What This Code Does
- Locates the correct script: We loop through all script tags and pick the one that mentions
languageDistribution(a dead giveaway for our target data). - Extracts JSON: The regex
window\.__INITIAL_STATE__ = (.*?);grabs the JSON object stored in the global variable, ignoring the surrounding JavaScript syntax. - Parses and accesses data: We navigate the parsed JSON to get to the
languageDistributionarray, then loop through each entry to print the language name and percentage just like your desired output (Spanish 61.1%, French 9.7%, etc).
Notes
- Zippia might change their variable names or data structure over time, but as of now, this works for their company pages. If it breaks later, check the script tags again to find where the demographic data is stored.
- Make sure you have all required libraries installed (
requests,beautifulsoup4,lxml)—you can install them withpip install requests beautifulsoup4 lxmlif needed.
内容的提问来源于stack exchange,提问作者hard1009

