You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过spaCy结合正则自定义URL的词元(Lemma)与标记(Tag)?

Customizing URL Token Processing in spaCy

Great question! To adjust how spaCy handles URLs—setting the lemma to the domain and tagging it as URL—you can create a custom pipeline component that uses regex to detect URLs, extract their domains, and modify the token attributes. Here's a step-by-step implementation:

Step 1: Define Regex Patterns

First, we need two regex patterns: one to identify full URLs, and another to extract the domain from those URLs.

import re
import spacy

# Pattern to match full URLs (starts with http/https, followed by non-whitespace characters)
URL_PATTERN = re.compile(r'https?://\S+')
# Pattern to extract the domain (captures everything after http(s):// and optional www.)
DOMAIN_PATTERN = re.compile(r'https?://(?:www\.)?([a-zA-Z0-9.-]+)')

Step 2: Create the Custom Pipeline Component

This component will run after spaCy's tagger (so we can override the default tagging) and modify URL tokens:

def process_urls(doc):
    for token in doc:
        # Check if the token is a full URL
        if URL_PATTERN.fullmatch(token.text):
            # Extract the domain from the URL
            domain_match = DOMAIN_PATTERN.search(token.text)
            if domain_match:
                # Set lemma to the extracted domain
                token.lemma_ = domain_match.group(1)
            # Override the tag to 'URL'
            token.tag_ = 'URL'
            # Optional: Set POS tag to 'SYM' (symbol) for consistency
            token.pos_ = 'SYM'
    return doc

Step 3: Add the Component to spaCy's Pipeline

We'll add our custom component right after the tagger to ensure we overwrite the default tags and lemmas:

# Load the default English model
nlp = spacy.load('en')

# Add the custom component after the tagger
nlp.add_pipe(process_urls, after='tagger')

Step 4: Test the Implementation

Let's run it on your example text:

msg = 'I got this URL https://stackoverflow.com/questions/47637005/handmade-estimator-modifies-parameters-in-init/47637293?noredirect=1#comment82268544_47637293 freed'
doc = nlp(msg)

# Print the token details
print([(t.text, t.lemma_, t.pos_, t.tag_, t.dep_) for t in doc])

Expected Output:

[
    ('I', '-PRON-', 'PRON', 'PRP', 'nsubj'),
    ('got', 'get', 'VERB', 'VBD', 'ROOT'),
    ('this', 'this', 'DET', 'DT', 'det'),
    ('URL', 'url', 'NOUN', 'NN', 'compound'),
    ('https://stackoverflow.com/questions/47637005/handmade-estimator-modifies-parameters-in-init/47637293?noredirect=1#comment82268544_47637293', 'stackoverflow.com', 'SYM', 'URL', 'nsubj'),
    ('freed', 'free', 'VERB', 'VBN', 'ccomp')
]

Notes:

  • Tokenization: SpaCy's default tokenizer usually keeps URLs as single tokens, but if you run into cases where URLs are split, you can add custom tokenizer rules to prevent this (e.g., adding http:// as a prefix that shouldn't be split).
  • Custom Tags: URL isn't a standard spaCy POS tag, but spaCy allows you to set custom tags as needed. If you want to use a standard tag, SYM (symbol) is a good fit for URLs.
  • Fallback: If the domain can't be extracted (e.g., malformed URL), the lemma will remain the original URL text—you can adjust this to a default value if needed.

内容的提问来源于stack exchange,提问作者Dror

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:21:31