如何用Selenium Python获取Twitter的年-月-日格式时间戳?
Hey there! I’ve dealt with this exact frustration before—those relative timestamps like "4h" or "Mar 7" are totally useless when you need structured date data. Luckily, Twitter actually hides the full, precise timestamp right in the page code, so we don’t have to guess or calculate it manually. Here are the best ways to pull that YYYY-MM-DD format you need:
Method 1: Grab the datetime Attribute (Most Reliable)
Twitter’s timestamp elements (usually wrapped in a <time> tag) include a datetime attribute that stores the full ISO 8601 timestamp (e.g., 2024-03-07T14:22:15Z). We can extract this directly and convert it to your desired YYYY-MM-DD format with Python’s datetime module.
Example Code:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from datetime import datetime # Initialize the driver (use Chrome, Firefox, etc.) driver = webdriver.Chrome() driver.get("https://twitter.com/your_target_account/status/your_tweet_id") # Wait for the timestamp element to load (avoids race conditions) try: timestamp_element = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//time[@datetime]")) ) # Extract the full ISO datetime string iso_timestamp = timestamp_element.get_attribute("datetime") # Convert to YYYY-MM-DD format (handles UTC timezone correctly) utc_datetime = datetime.fromisoformat(iso_timestamp.replace("Z", "+00:00")) formatted_date = utc_datetime.strftime("%Y-%m-%d") print(f"Full date: {formatted_date}") finally: driver.quit()
Method 2: Handle Multiple Tweets & Dynamic Loading
If you’re scraping a timeline with multiple tweets, you’ll need to loop through all <time> elements. For infinite scroll timelines, you’ll also need to scroll to load more tweets before extracting:
Example for Timelines:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.keys import Keys import time from datetime import datetime driver = webdriver.Chrome() driver.get("https://twitter.com/your_target_account") # Scroll to load more tweets (adjust scroll count as needed) for _ in range(3): driver.find_element(By.TAG_NAME, "body").send_keys(Keys.END) time.sleep(2) # Give time for tweets to load # Get all timestamp elements timestamp_elements = driver.find_elements(By.XPATH, "//time[@datetime]") # Extract and format each date for elem in timestamp_elements: iso_timestamp = elem.get_attribute("datetime") utc_datetime = datetime.fromisoformat(iso_timestamp.replace("Z", "+00:00")) formatted_date = utc_datetime.strftime("%Y-%m-%d") print(formatted_date) driver.quit()
Why Avoid Parsing Relative Timestamps?
Trying to convert "4h" or "Mar 7" to a full date is error-prone:
- You have to infer the current year for "Mar 7" (what if it’s January and the tweet is from last March?)
- Time zones and daylight saving can mess up relative time calculations
- Twitter’s relative time updates dynamically (a "4h" tweet becomes "5h" an hour later)
Sticking to the datetime attribute gives you a static, accurate timestamp every time.
内容的提问来源于stack exchange,提问作者Jake Park

