You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中匹配字符串形式的基因名并筛选数据集子集

筛选包含目标基因的DataFrame行

Got it, let's work through this problem. You need to filter rows in your second DataFrame where the gene column (with semicolon-separated values) contains any gene from your hits list—makes total sense. Here are two reliable approaches using pandas, depending on your use case:

方法1:直观的集合交集检查(适合小数据集)

First, let's set up some sample data to match your example:

import pandas as pd

# Your target gene list
hits = ["IL1", "NRC31", "AR"]

# Simulate your second DataFrame
data = {
    'id': [68, 236, 272, 324, 426033, 425700],
    'genes': [
        "NFKBIL1;NFKBIL1;ATP6V1G2;NFKBIL1;NFKBIL1;NFKBI",
        "BARHL2",
        "ARPC2;ARPC2",
        "MARCH5",
        "ABC1;IL1;XYZ2",
        "IL17D"
    ]
}
df = pd.DataFrame(data)

Then, we'll write a simple function to check if any gene in the semicolon-separated string matches our hits:

def matches_hits(gene_string, target_genes):
    # Split the string into individual genes and convert to a set for fast lookups
    row_genes = set(gene_string.split(';'))
    # Check if there's any overlap with our target hits
    return bool(row_genes.intersection(target_genes))

# Apply the function to filter rows
filtered_df = df[df['genes'].apply(lambda x: matches_hits(x, hits))]

This works by splitting each cell's gene string into a set, then checking for any intersection with the hits set. If there's at least one match, the row is kept.

方法2:向量化正则匹配(适合大数据集)

For larger datasets, using pandas' vectorized string operations will be way faster than apply. We'll build a regex pattern that matches exact gene names (to avoid partial matches like confusing IL1 with IL17D):

import re

# Escape any special characters in gene names (in case you have genes with +, ., etc.)
escaped_hits = [re.escape(gene) for gene in hits]
# Build a pattern that matches a gene either at the start, after a semicolon,
# and followed by a semicolon or end of string
pattern = r'(?:^|;)({})(?:;|$)'.format('|'.join(escaped_hits))

# Use str.contains to filter rows
filtered_df = df[df['genes'].str.contains(pattern, regex=True)]

处理大小写敏感的情况

If your gene names might have inconsistent capitalization (e.g., il1 vs IL1), adjust the code to normalize case:

  • For method 1:
    def matches_hits(gene_string, target_genes):
        row_genes = set(g.lower() for g in gene_string.split(';'))
        target_set = set(g.lower() for g in target_genes)
        return bool(row_genes.intersection(target_set))
    
  • For method 2:
    filtered_df = df[df['genes'].str.lower().str.contains(pattern.lower(), regex=True)]
    

结果示例

Running either method on our sample data will return this filtered DataFrame:

idgenes
426033ABC1;IL1;XYZ2

That's exactly the row containing IL1 from our hits list!

内容的提问来源于stack exchange,提问作者Calen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:28:20