如何从URL中提取可变的10字符指定字符串?
Hey there! Great question—this is a super common task when working with e-commerce URLs like Amazon’s. While there’s no universal "out-of-the-box" function that works across all languages, you can easily build your own reusable function with a few straightforward approaches. Let me walk you through them:
1. String Splitting (Simple & Direct)
This method leverages the fixed structure of Amazon’s product URLs: your target string always comes right after /gp/product/ and ends at the next /. We can use basic string operations to find those positions and extract the substring.
Example in Python:
def extract_product_id(url): # Define the marker that precedes our target string marker = "/gp/product/" # Find the start index right after the marker start_idx = url.find(marker) + len(marker) # Find the next slash after the start index to get the end position end_idx = url.find("/", start_idx) # Return the substring between the two indices return url[start_idx:end_idx] # Test with your sample URL sample_url = "https://www.amazon.de/gp/product/SOMETEXTHERE/ref=oh_aui_detailpage_o00_s00?ie=UTF8&psc=1" print(extract_product_id(sample_url)) # Output: SOMETEXTHERE
Example in JavaScript:
function extractProductId(url) { const marker = "/gp/product/"; const startIdx = url.indexOf(marker) + marker.length; const endIdx = url.indexOf("/", startIdx); return url.slice(startIdx, endIdx); } // Test with your sample URL const sampleUrl = "https://www.amazon.de/gp/product/SOMETEXTHERE/ref=oh_aui_detailpage_o00_s00?ie=UTF8&psc=1"; console.log(extractProductId(sampleUrl)); // Output: SOMETEXTHERE
2. Regular Expressions (Flexible for Edge Cases)
If you need to handle slight variations in the URL structure, regex is a great option. We can write a pattern that matches the /gp/product/ segment and captures everything until the next slash.
Example in Python:
import re def extract_product_id_regex(url): # The regex pattern captures any characters that aren't a slash after "/gp/product/" match = re.search(r'/gp/product/([^/]+)', url) # Return the captured group if a match is found, else None return match.group(1) if match else None print(extract_product_id_regex(sample_url)) # Output: SOMETEXTHERE
Example in JavaScript:
function extractProductIdRegex(url) { const match = url.match(/\/gp\/product\/([^/]+)/); return match ? match[1] : null; } console.log(extractProductIdRegex(sampleUrl)); // Output: SOMETEXTHERE
3. URL Parsing Libraries (Most Robust)
For a more standardized approach, use your language’s built-in URL parsing tools. These handle edge cases like query parameters or varying domain extensions automatically.
Example in Python:
from urllib.parse import urlparse def extract_product_id_with_parser(url): parsed_url = urlparse(url) # Split the path into segments (e.g., ["", "gp", "product", "SOMETEXTHERE", "ref=..."]) path_segments = parsed_url.path.split("/") try: # Find the index of "product" and take the next segment product_index = path_segments.index("product") return path_segments[product_index + 1] except ValueError: # Return None if "product" isn't found in the path return None print(extract_product_id_with_parser(sample_url)) # Output: SOMETEXTHERE
Example in JavaScript:
function extractProductIdWithURL(url) { const urlObj = new URL(url); const pathSegments = urlObj.pathname.split("/"); const productIndex = pathSegments.indexOf("product"); return productIndex !== -1 ? pathSegments[productIndex + 1] : null; } console.log(extractProductIdWithURL(sampleUrl)); // Output: SOMETEXTHERE
All of these approaches will reliably pull that variable string (whether it’s exactly 10 characters or another length) from the URL structure. Just pick the one that fits your programming language and use case best, wrap it into a function, and you’re good to go!
内容的提问来源于stack exchange,提问作者MrMinemeet

