使用rvest爬取TripAdvisor餐厅经纬度数据失败求解决
Hey there! Sorry to hear you're hitting a wall grabbing lat/long data from TripAdvisor with rvest—let's break this down and get you those coordinates. Since you can see the lat/long in the page source, the issue is almost certainly with how you're targeting that data in your code. Here are the most common fixes and approaches:
1. First, Confirm Where the Lat/Long Lives in the Page
Fire up your browser's DevTools (hit F12), search for "lat" or "long" in the page source, and note exactly where the coordinates are stored. TripAdvisor typically uses one of these three spots:
- Meta tags: Look for
<meta property="geo:lat" content="XXX.XXX">or similar - Data attributes: A div/element with
data-latanddata-longproperties - Embedded JSON: A
<script>tag holding a JSON object (often in something likewindow.__INITIAL_STATE__)
2. Target Meta Tags (Simple Case)
If your coordinates are in meta tags, use this straightforward code:
library(rvest) # Replace with your target TripAdvisor URL target_url <- "https://www.tripadvisor.com/..." page <- read_html(target_url) # Grab latitude and longitude latitude <- page %>% html_element(xpath = '//meta[@property="geo:lat"]') %>% html_attr("content") longitude <- page %>% html_element(xpath = '//meta[@property="geo:long"]') %>% html_attr("content") # Print results cat("Latitude:", latitude, "\nLongitude:", longitude)
3. Target Data Attributes (If Coordinates Are in Element Properties)
If you see something like <div class="some-class" data-lat="XXX.XXX" data-long="YYY.YYY">, use this code:
library(rvest) target_url <- "https://www.tripadvisor.com/..." page <- read_html(target_url) latitude <- page %>% html_element(xpath = '//div[@data-lat]') %>% html_attr("data-lat") longitude <- page %>% html_element(xpath = '//div[@data-long]') %>% html_attr("data-long") cat("Latitude:", latitude, "\nLongitude:", longitude)
4. Extract from Embedded JSON (Most Common for Modern TripAdvisor Pages)
TripAdvisor often loads core data into a JSON object in a script tag. Here's how to pull that out:
library(rvest) library(jsonlite) target_url <- "https://www.tripadvisor.com/..." page <- read_html(target_url) # Find the script tag with lat/long data (adjust the grep pattern if needed) script_text <- page %>% html_elements("script") %>% html_text() %>% grep("lat|long|coordinates", ., value = TRUE) %>% .[1] # Grab the first matching script (tweak the index if necessary) # Clean the script to get pure JSON (adjust the regex to match the actual variable name) json_raw <- gsub("^window\\.__INITIAL_STATE__ = |;$", "", script_text) parsed_json <- fromJSON(json_raw) # Navigate the JSON structure to find lat/long (adjust this path to match your page!) # Example path—yours might be parsed_json$location$coordinates or similar latitude <- parsed_json$entity$location$lat longitude <- parsed_json$entity$location$lng cat("Latitude:", latitude, "\nLongitude:", longitude)
Quick Debug Tip:
Run str(parsed_json) in your console to explore the full structure of the JSON data—this will help you find the exact path to your lat/long values.
5. If You Still Get NA
- Use
html_elements()(plural) instead ofhtml_element()to see if there are multiple matching elements—you might be targeting the wrong one. - Right-click the element with lat/long in DevTools, choose "Copy > Copy XPath" or "Copy > Copy selector", and paste that directly into your code to ensure your selector is accurate.
内容的提问来源于stack exchange,提问作者G H

