如何使用REGEX去除无关字符,将亚马逊长链接转为产品短链接?
Alright, let's figure out how to trim those bloated Amazon product links down to their clean, core version using regex. Here's the most reliable approach I've used for this exact task:
Core Idea
Amazon's basic product link follows this structure: https://<domain>/gp/product/<ASIN>/ (or sometimes https://<domain>/dp/<ASIN>/). The ASIN is Amazon's 10-character alphanumeric product identifier—this is the only part we need to keep alongside the base domain and path. All the extra query parameters (the stuff after ? or /ref=...) are irrelevant for directly accessing the product page.
Regex Pattern & Explanation
We'll use a regex with capture groups to extract the essential parts of the link:
^(https?://[^/]+/(?:gp/product/|dp/))([A-Za-z0-9]{10})
Let's break this down:
^: Ensures we start matching from the very beginning of the link (avoids accidental matches in the middle of text)(https?://[^/]+/(?:gp/product/|dp/)): Captures the base URL segment, including:https?://: Matches bothhttpandhttpsprotocols[^/]+: Grabs the full domain (e.g.,www.amazon.de,www.amazon.com)(?:gp/product/|dp/): Matches either the/gp/product/or/dp/path prefix (non-capturing group since we don't need to separate these)
([A-Za-z0-9]{10}): Captures the 10-character ASIN code (alphanumeric, case-insensitive)
Implementation Examples
Python
import re long_amazon_link = "https://www.amazon.de/gp/product/ADKLHJADK/ref=as_li_ss_tl?ie=UTF8&pd_rd_i=B01J7LLL9Q&pd_rd_r=a8c7bb4b-49da-11e8-ad28-014ae5dc2f42&pd_rd_w=9QOk2&pd_rd_wg=zc1s7&pf_rd_m=A3JWKAKR8XB7XF&pf_rd_s=&pf_rd_r=VF3C7MDNZ741H8S13AYV&pf_rd_t=36701&pf_rd_p=1c175abe-9bc7-490b-bbe1-2caf7e752c98&pf_rd_i=desktop&linkCode=ll1" regex_pattern = r"^(https?://[^/]+/(?:gp/product/|dp/))([A-Za-z0-9]{10})" match_result = re.match(regex_pattern, long_amazon_link) if match_result: clean_link = f"{match_result.group(1)}{match_result.group(2)}/" print(clean_link) # Output: https://www.amazon.de/gp/product/ADKLHJADK/
JavaScript
const longAmazonLink = "https://www.amazon.de/gp/product/ADKLHJADK/ref=as_li_ss_tl?ie=UTF8&pd_rd_i=B01J7LLL9Q&pd_rd_r=a8c7bb4b-49da-11e8-ad28-014ae5dc2f42&pd_rd_w=9QOk2&pd_rd_wg=zc1s7&pf_rd_m=A3JWKAKR8XB7XF&pf_rd_s=&pf_rd_r=VF3C7MDNZ741H8S13AYV&pf_rd_t=36701&pf_rd_p=1c175abe-9bc7-490b-bbe1-2caf7e752c98&pf_rd_i=desktop&linkCode=ll1"; const regexPattern = /^(https?:\/\/[^\/]+\/(?:gp\/product\/|dp\/))([A-Za-z0-9]{10})/; const matchResult = longAmazonLink.match(regexPattern); if (matchResult) { const cleanLink = `${matchResult[1]}${matchResult[2]}/`; console.log(cleanLink); // Output: https://www.amazon.de/gp/product/ADKLHJADK/ }
Key Notes
- Compatibility: This regex works for all Amazon regional domains (
.de,.com,.co.uk, etc.) and both common product link formats (/gp/product/and/dp/). - Reliability: Since ASINs are always exactly 10 alphanumeric characters, this avoids false positives or incorrect truncation.
- Case Insensitivity: If you're working with links that have lowercase ASINs, you can add the case-insensitive flag (
re.IGNORECASEin Python,/iin JavaScript) to the regex for extra robustness.
内容的提问来源于stack exchange,提问作者Jopy Downstair

