LinkedIn网页抓取技术求助:API令牌整合实现公司主页HTML自动抓取
Hey there! I totally get the frustration of hitting access denied errors when trying to scrape LinkedIn pages directly—LinkedIn has pretty strict anti-scraping measures, and scraping their raw HTML violates their Terms of Service anyway. Using their official API is the right way to go, so let's walk through how to integrate your token and build that "input company name, get data" workflow step by step.
First: Why HTML Scraping Isn't Working
LinkedIn actively blocks automated scraping tools like BeautifulSoup by detecting non-human traffic patterns, so even if you tweak headers, you'll likely keep hitting blocks. Their API is designed for legitimate data access, so we'll focus on that instead.
Step 1: Integrate Your API Token into Requests
Your API token needs to be included in the request headers for every API call to authenticate yourself. Here's how to set that up:
import requests import json # Replace with your actual API token API_TOKEN = "your_linkedin_api_token_here" REQUEST_HEADERS = { "Authorization": f"Bearer {API_TOKEN}", "Content-Type": "application/json" }
Step 2: Search for a Company by Name to Get Its ID
LinkedIn's API requires a unique company ID to pull detailed data, so first we'll write a function to search for a company by name and grab its ID (we'll use the first matching result for simplicity):
def get_company_id(company_name): search_endpoint = f"https://api.linkedin.com/v2/companies?q=name&keywords={company_name}&count=1" response = requests.get(search_endpoint, headers=REQUEST_HEADERS) if response.status_code != 200: print(f"Search failed: {response.json().get('message', 'Unknown error')}") return None search_results = response.json() if not search_results.get("elements"): print(f"No company found matching '{company_name}'") return None return search_results["elements"][0]["id"]
Step 3: Fetch Detailed Company Data Using the ID
Once we have the company ID, we can pull the specific data points you need (like description, website, employee count, etc.) using the Companies API endpoint:
def get_company_details(company_id): # Use the 'projection' parameter to specify exactly which fields you want details_endpoint = f"https://api.linkedin.com/v2/companies/{company_id}?projection=(id,name,description,industry,website,employeeCountRange,foundedDate,location)" response = requests.get(details_endpoint, headers=REQUEST_HEADERS) if response.status_code != 200: print(f"Failed to fetch details: {response.json().get('message', 'Unknown error')}") return None return response.json()
Step 4: Build the Full "Input Company Name" Workflow
Now tie it all together into a single function that handles everything automatically:
def fetch_linkedin_company_data(company_name): company_id = get_company_id(company_name) if not company_id: return None company_data = get_company_details(company_id) if company_data: print("Successfully fetched company data:") print(json.dumps(company_data, indent=2)) return company_data # Example usage fetch_linkedin_company_data("Google")
Key Things to Keep in Mind
- Permissions: Make sure your API token has the required scopes (like
r_organization_socialorr_company_admin) to access the data you want. Check LinkedIn's API docs for exact scope requirements. - Rate Limits: LinkedIn enforces rate limits on API calls—keep an eye on the
X-RateLimit-Remainingheader in responses to avoid getting blocked temporarily. - Projection Customization: Adjust the
projectionparameter in the details endpoint to add or remove fields based on your needs (e.g., addlogoV2if you want the company logo URL).
内容的提问来源于stack exchange,提问作者Gaurav Chavan

