You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多词键字典匹配推特公司名返回股票代码问题求助

Solution for Matching Both Single-Word and Multi-Word Company Names in Tweets

Got it, let's work through this problem. The core pain point here is that splitting your tweet content for single-word matching breaks your ability to catch multi-word company names like CHINA UNICOM or EXPRESS SCRIPTS. Luckily, there are a few straightforward ways to fix this without ditching your single-word matching logic. Here are my top recommendations:

1. Prioritize Multi-Word Matches First (Then Handle Single Words)

This approach lets you keep your single-word workflow intact, but we'll add a pre-step to catch multi-word phrases before splitting the text.

How it works:

  • First, scan the full tweet text for any multi-word company names from your dictionary.
  • Once a multi-word match is found, record its stock code and remove that phrase from the text (to avoid duplicate or partial matches later).
  • Finally, split the remaining text and process it for single-word matches like you were doing before.

Example Code (Python):

# Your company dictionary (mix of multi-word and single-word entries)
company_dict = {
    "CHINA UNICOM": "0762.HK",
    "EXPRESS SCRIPTS": "ESRX",
    "APPLE": "AAPL",
    "GOOGLE": "GOOGL"
}

def match_company_codes(tweet_text):
    matched_codes = []
    processed_text = tweet_text.upper()  # Normalize to uppercase for consistent matching

    # Sort multi-word entries by length (longest first) to avoid partial matches
    sorted_companies = sorted(company_dict.keys(), key=lambda x: len(x.split()), reverse=True)

    # First pass: catch multi-word companies
    for company in sorted_companies:
        if company in processed_text:
            matched_codes.append(company_dict[company])
            # Remove the matched phrase to prevent re-matching
            processed_text = processed_text.replace(company, "")

    # Second pass: handle single-word matches on the remaining text
    for word in processed_text.split():
        if word in company_dict:
            matched_codes.append(company_dict[word])

    # Return unique codes (in case of duplicates)
    return list(set(matched_codes))

2. Use Regular Expressions with Word Boundaries

If you don't strictly need to split the text upfront, regex is a clean way to match both single and multi-word entries in one pass. The \b word boundary ensures you only match full phrases/words, no partial hits.

How it works:

  • Build a regex pattern that includes all your company names (both single and multi-word), separated by |.
  • Use word boundaries (\b) around each entry to ensure you're matching complete terms.
  • Scan the tweet text with this regex to pull out all matches, then map them to their stock codes.

Example Code (Python):

import re

company_dict = {
    "CHINA UNICOM": "0762.HK",
    "EXPRESS SCRIPTS": "ESRX",
    "APPLE": "AAPL"
}

# Build regex pattern: escape special chars, add word boundaries, case-insensitive
company_pattern = r'\b(' + '|'.join(re.escape(company) for company in company_dict.keys()) + r')\b'
regex_matcher = re.compile(company_pattern, re.IGNORECASE)

def match_company_codes(tweet_text):
    # Find all matching company names in the tweet
    matched_companies = regex_matcher.findall(tweet_text)
    # Map matches to their stock codes (normalize to uppercase for dict lookup)
    matched_codes = [company_dict[match.upper()] for match in matched_companies]
    # Return unique codes
    return list(set(matched_codes))

3. Hybrid Approach: Check Adjacent Words If You Must Split Text

If you absolutely have to split the tweet text (for other processing needs), you can adjust your workflow to check if consecutive words form a multi-word company name.

How it works:

  • Split the text into words like before, but iterate through the list while checking if groups of adjacent words match any multi-word entries in your dictionary.
  • When a multi-word match is found, skip ahead in the word list to avoid reprocessing those words.
  • For non-matching positions, check for single-word matches as usual.

Example Code (Python):

company_dict = {
    "CHINA UNICOM": "0762.HK",
    "EXPRESS SCRIPTS": "ESRX",
    "APPLE": "AAPL"
}

# Separate multi-word and single-word entries for easier handling
multi_word_firms = {tuple(comp.split()): code for comp, code in company_dict.items() if len(comp.split()) > 1}
single_word_firms = {comp: code for comp, code in company_dict.items() if len(comp.split()) == 1}

def match_company_codes(tweet_text):
    matched_codes = []
    words = tweet_text.upper().split()
    word_count = len(words)
    index = 0

    while index < word_count:
        matched_multi = False
        # Check for the longest possible multi-word match first
        max_multi_length = max(len(phrase) for phrase in multi_word_firms.keys()) if multi_word_firms else 1
        # Iterate from longest possible down to 2 words
        for length in range(min(max_multi_length, word_count - index), 1, -1):
            current_phrase = tuple(words[index:index+length])
            if current_phrase in multi_word_firms:
                matched_codes.append(multi_word_firms[current_phrase])
                index += length
                matched_multi = True
                break
        # If no multi-word match, check single word
        if not matched_multi:
            if words[index] in single_word_firms:
                matched_codes.append(single_word_firms[words[index]])
            index += 1

    return list(set(matched_codes))

All three methods let you keep supporting single-word matches while successfully catching those tricky multi-word company names. The regex approach is my go-to if you don't have a hard requirement to split the text upfront—it's concise and avoids extra text manipulation steps.

内容的提问来源于stack exchange,提问作者Sahil Desai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:52:29