CellMarker

lamindb provides access to the following public cell marker ontologies through bionty:

  1. CellMarker

Here we show how to access and search cell marker ontologies to standardize new data.

import bionty as bt
import pandas as pd

PublicOntology objects

Let us create a public ontology accessor with public(), which chooses a default public ontology source from Source. It’s a PublicOntology object, which you can think about as a public registry:

public = bt.CellMarker.public(organism="human")
public
 connected lamindb: testuser1/test-public-ontologies
PublicOntology
Entity: CellMarker
Organism: human
Source: cellmarker, 2.0
#terms: 15466

As for registries, you can export the ontology as a DataFrame:

df = public.df()
df.head()
name synonyms gene_symbol ncbi_gene_id uniprotkb_id
0 A1BG A1BG 1 P04217
1 A2M A2M 3494 None
2 A2ML1 A2ML1 144568 A8K2U0
3 A4GALT A4GALT 53947 A0A0S2Z5J1
4 AADAC AADAC 13 P22760

Unlike registries, you can also export it as a Pronto object via public.ontology.

Look up terms

As for registries, terms can be looked up with auto-complete:

lookup = public.lookup()

The . accessor provides normalized terms (lower case, only contains alphanumeric characters and underscores):

lookup.immp1l
CellMarker(name='IMMP1L', synonyms='', gene_symbol='IMMP1L', ncbi_gene_id='196294', uniprotkb_id='Q96LU5')

To look up the exact original strings, convert the lookup object to dict and use the [] accessor:

lookup_dict = lookup.dict()
lookup_dict["IMMP1L"]
CellMarker(name='IMMP1L', synonyms='', gene_symbol='IMMP1L', ncbi_gene_id='196294', uniprotkb_id='Q96LU5')

Search terms

Search behaves in the same way as it does for registries:

public.search("CD4").head(5)
name synonyms gene_symbol ncbi_gene_id uniprotkb_id
1900 Cd4 CD4 920 B4DT49
1901 CD40 CD40 958 A0A0S2Z3C7
1905 CD40LG CD40LG 959 P29965
1908 CD46 CD46 4179 P15529
1907 Cd44 CD44 960 P16070

Search another field (default is .name):

public.search("CD4", field=public.gene_symbol).head(1)
name synonyms gene_symbol ncbi_gene_id uniprotkb_id
1900 Cd4 CD4 920 B4DT49

Standardize cell marker identifiers

Let us generate a DataFrame that stores a number of cell markers identifiers, some of which corrupted:

markers = pd.DataFrame(
    index=[
        "KI67",
        "CCR7",
        "CD14",
        "CD8",
        "CD45RA",
        "CD4",
        "CD3",
        "CD127a",
        "PD1",
        "Invalid-1",
        "Invalid-2",
        "CD66b",
        "Siglec8",
        "Time",
    ]
)

Now let’s check which cell markers can be found in the reference:

public.inspect(markers.index, public.name);
! 8 unique terms (57.10%) are not validated for name: 'KI67', 'CCR7', 'CD14', 'CD4', 'CD127a', 'Invalid-1', 'Invalid-2', 'Time'
   detected 4 unique terms with inconsistent casing/synonyms: KI67, CCR7, CD14, CD4
→  standardize terms via .standardize()

Logging suggests to map synonyms:

synonyms_mapper = public.standardize(markers.index, return_mapper=True)
synonyms_mapper
{'KI67': 'Ki67', 'CCR7': 'Ccr7', 'CD14': 'Cd14', 'CD4': 'Cd4'}

Let’s replace the synonyms with standardized names in the DataFrame:

markers.rename(index=synonyms_mapper, inplace=True)

The Time, Invalid-1 and Invalid-2 are non-marker channels which won’t be curated by cell marker:

public.inspect(markers.index, public.name);
! 4 unique terms (28.60%) are not validated for name: 'CD127a', 'Invalid-1', 'Invalid-2', 'Time'

We don’t find CD127a, let’s check in the lookup with auto-completion:

lookup = public.lookup()
lookup.cd127
CellMarker(name='CD127', synonyms='', gene_symbol='IL7R', ncbi_gene_id='3575', uniprotkb_id='P16871', _5='cd127')

It should be cd127, we had a typo there with cd127a:

curated_df = markers.rename(index={"CD127a": lookup.cd127.name})

Optionally, search:

public.search("CD127a").head()
name synonyms gene_symbol ncbi_gene_id uniprotkb_id __agg__

Now we see that all cell marker candidates validate:

public.validate(curated_df.index, public.name);
! 3 unique terms (21.40%) are not validated: 'Invalid-1', 'Invalid-2', 'Time'

Ontology source versions

For any given entity, we can choose from a number of versions:

bt.Source.filter(entity="bionty.CellMarker").df()
Hide code cell output
uid entity organism name in_db currently_used description url md5 source_website space_id dataframe_artifact_id version run_id created_at created_by_id _aux _branch_code
id
28 3kDh bionty.CellMarker human cellmarker False True CellMarker s3://bionty-assets/human_cellmarker_2.0_CellMa... d565d4a542a5c7e7a06255975358e4f4 http://bio-bigdata.hrbmu.edu.cn/CellMarker 1 None 2.0 None 2025-01-20 07:34:38.666000+00:00 1 None 1
29 7bV5 bionty.CellMarker mouse cellmarker False True CellMarker s3://bionty-assets/mouse_cellmarker_2.0_CellMa... 189586732c63be949e40dfa6a3636105 http://bio-bigdata.hrbmu.edu.cn/CellMarker 1 None 2.0 None 2025-01-20 07:34:38.666000+00:00 1 None 1
# only lists the sources that are currently used
bt.Source.filter(entity="bionty.CellMarker", currently_used=True).df()
uid entity organism name in_db currently_used description url md5 source_website space_id dataframe_artifact_id version run_id created_at created_by_id _aux _branch_code
id
28 3kDh bionty.CellMarker human cellmarker False True CellMarker s3://bionty-assets/human_cellmarker_2.0_CellMa... d565d4a542a5c7e7a06255975358e4f4 http://bio-bigdata.hrbmu.edu.cn/CellMarker 1 None 2.0 None 2025-01-20 07:34:38.666000+00:00 1 None 1
29 7bV5 bionty.CellMarker mouse cellmarker False True CellMarker s3://bionty-assets/mouse_cellmarker_2.0_CellMa... 189586732c63be949e40dfa6a3636105 http://bio-bigdata.hrbmu.edu.cn/CellMarker 1 None 2.0 None 2025-01-20 07:34:38.666000+00:00 1 None 1

When instantiating a Bionty object, we can choose a source or version:

source = bt.Source.get(name="cellmarker", version="2.0", organism="human")
public = bt.CellMarker.public(source=source)
public
PublicOntology
Entity: CellMarker
Organism: human
Source: cellmarker, 2.0
#terms: 15466

The currently used ontologies can be displayed using:

bt.Source.filter(currently_used=True).df()
Hide code cell output
uid entity organism name in_db currently_used description url md5 source_website space_id dataframe_artifact_id version run_id created_at created_by_id _aux _branch_code
id
1 33TU bionty.Organism vertebrates ensembl False True Ensembl https://ftp.ensembl.org/pub/release-112/specie... 0ec37e77f4bc2d0b0b47c6c62b9f122d https://www.ensembl.org 1 None release-112 None 2025-01-20 07:34:38.666000+00:00 1 None 1
6 6bbV bionty.Organism bacteria ensembl False True Ensembl https://ftp.ensemblgenomes.ebi.ac.uk/pub/bacte... ee28510ed5586ea7ab4495717c96efc8 https://www.ensembl.org 1 None release-57 None 2025-01-20 07:34:38.666000+00:00 1 None 1
7 6s9n bionty.Organism fungi ensembl False True Ensembl http://ftp.ensemblgenomes.org/pub/fungi/releas... dbcde58f4396ab8b2480f7fe9f83df8a https://www.ensembl.org 1 None release-57 None 2025-01-20 07:34:38.666000+00:00 1 None 1
8 2PmT bionty.Organism metazoa ensembl False True Ensembl http://ftp.ensemblgenomes.org/pub/metazoa/rele... 424636a574fec078a61cbdddb05f9132 https://www.ensembl.org 1 None release-57 None 2025-01-20 07:34:38.666000+00:00 1 None 1
9 7GPH bionty.Organism plants ensembl False True Ensembl https://ftp.ensemblgenomes.ebi.ac.uk/pub/plant... eadaa1f3e527e4c3940c90c7fa5c8bf4 https://www.ensembl.org 1 None release-57 None 2025-01-20 07:34:38.666000+00:00 1 None 1
10 4tsk bionty.Organism all ncbitaxon False True NCBItaxon Ontology s3://bionty-assets/df_all__ncbitaxon__2023-06-... 00d97ba65627f1cd65636d2df22ea76c https://github.com/obophenotype/ncbitaxon 1 None 2023-06-20 None 2025-01-20 07:34:38.666000+00:00 1 None 1
11 4UGN bionty.Gene human ensembl False True Ensembl s3://bionty-assets/df_human__ensembl__release-... 4ccda4d88720a326737376c534e8446b https://www.ensembl.org 1 None release-112 None 2025-01-20 07:34:38.666000+00:00 1 None 1
15 4r4f bionty.Gene mouse ensembl False True Ensembl s3://bionty-assets/df_mouse__ensembl__release-... 519cf7b8acc3c948274f66f3155a3210 https://www.ensembl.org 1 None release-112 None 2025-01-20 07:34:38.666000+00:00 1 None 1
19 4RPA bionty.Gene saccharomyces cerevisiae ensembl False True Ensembl s3://bionty-assets/df_saccharomyces cerevisiae... 11775126b101233525a0a9e2dd64edae https://www.ensembl.org 1 None release-112 None 2025-01-20 07:34:38.666000+00:00 1 None 1
22 3EYy bionty.Protein human uniprot False True Uniprot s3://bionty-assets/df_human__uniprot__2024-03_... b5b9e7645065b4b3187114f07e3f402f https://www.uniprot.org 1 None 2024-03 None 2025-01-20 07:34:38.666000+00:00 1 None 1
25 01RW bionty.Protein mouse uniprot False True Uniprot s3://bionty-assets/df_mouse__uniprot__2024-03_... b1b6a196eb853088d36198d8e3749ec4 https://www.uniprot.org 1 None 2024-03 None 2025-01-20 07:34:38.666000+00:00 1 None 1
28 3kDh bionty.CellMarker human cellmarker False True CellMarker s3://bionty-assets/human_cellmarker_2.0_CellMa... d565d4a542a5c7e7a06255975358e4f4 http://bio-bigdata.hrbmu.edu.cn/CellMarker 1 None 2.0 None 2025-01-20 07:34:38.666000+00:00 1 None 1
29 7bV5 bionty.CellMarker mouse cellmarker False True CellMarker s3://bionty-assets/mouse_cellmarker_2.0_CellMa... 189586732c63be949e40dfa6a3636105 http://bio-bigdata.hrbmu.edu.cn/CellMarker 1 None 2.0 None 2025-01-20 07:34:38.666000+00:00 1 None 1
30 6LyR bionty.CellLine all clo False True Cell Line Ontology https://data.bioontology.org/ontologies/CLO/su... ea58a1010b7e745702a8397a526b3a33 https://bioportal.bioontology.org/ontologies/CLO 1 None 2022-03-21 None 2025-01-20 07:34:38.666000+00:00 1 None 1
32 3Uw2 bionty.CellType all cl False True Cell Ontology http://purl.obolibrary.org/obo/cl/releases/202... https://obophenotype.github.io/cell-ontology 1 None 2024-08-16 None 2025-01-20 07:34:38.666000+00:00 1 None 1
41 MUtA bionty.Tissue all uberon False True Uberon multi-species anatomy ontology http://purl.obolibrary.org/obo/uberon/releases... http://obophenotype.github.io/uberon 1 None 2024-08-07 None 2025-01-20 07:34:38.666000+00:00 1 None 1
50 4a3e bionty.Disease all mondo False True Mondo Disease Ontology http://purl.obolibrary.org/obo/mondo/releases/... https://mondo.monarchinitiative.org 1 None 2024-08-06 None 2025-01-20 07:34:38.666000+00:00 1 None 1
59 4ksw bionty.Disease human doid False True Human Disease Ontology http://purl.obolibrary.org/obo/doid/releases/2... bbefd72247d638edfcd31ec699947407 https://disease-ontology.org 1 None 2024-05-29 None 2025-01-20 07:34:38.670000+00:00 1 None 1
67 2a1H bionty.ExperimentalFactor all efo False True The Experimental Factor Ontology http://www.ebi.ac.uk/efo/releases/v3.70.0/efo.owl https://bioportal.bioontology.org/ontologies/EFO 1 None 3.70.0 None 2025-01-20 07:34:38.670000+00:00 1 None 1
74 48fB bionty.Phenotype human hp False True Human Phenotype Ontology https://github.com/obophenotype/human-phenotyp... e0f2e534eb2ad44a4d45573ef27b508f https://hpo.jax.org 1 None 2024-04-26 None 2025-01-20 07:34:38.670000+00:00 1 None 1
79 4t7Q bionty.Phenotype mammalian mp False True Mammalian Phenotype Ontology https://github.com/mgijax/mammalian-phenotype-... 795d8378fe48ec13b41d01a86dd1c86c https://github.com/mgijax/mammalian-phenotype-... 1 None 2024-06-18 None 2025-01-20 07:34:38.670000+00:00 1 None 1
82 sqPX bionty.Phenotype zebrafish zp False True Zebrafish Phenotype Ontology https://github.com/obophenotype/zebrafish-phen... 2231ebaa95becf8ff34a33c95a8d4350 https://github.com/obophenotype/zebrafish-phen... 1 None 2024-04-18 None 2025-01-20 07:34:38.670000+00:00 1 None 1
86 6S4q bionty.Phenotype all pato False True Phenotype And Trait Ontology http://purl.obolibrary.org/obo/pato/releases/2... 6b1eaacd3d453b34375ce2e31c16328a https://github.com/pato-ontology/pato 1 None 2024-03-28 None 2025-01-20 07:34:38.670000+00:00 1 None 1
88 7Ent bionty.Pathway all go False True Gene Ontology https://data.bioontology.org/ontologies/GO/sub... 7fa7ade5e3e26eab3959a7e4bc89ad4f http://geneontology.org 1 None 2024-06-17 None 2025-01-20 07:34:38.670000+00:00 1 None 1
93 3rm9 BFXPipeline all lamin False True Bioinformatics Pipeline s3://bionty-assets/df_all__lamin__1.0.0__BFXpi... https://lamin.ai 1 None 1.0.0 None 2025-01-20 07:34:38.670000+00:00 1 None 1
94 ugaI Drug all dron False True Drug Ontology https://data.bioontology.org/ontologies/DRON/s... https://bioportal.bioontology.org/ontologies/DRON 1 None 2024-08-05 None 2025-01-20 07:34:38.670000+00:00 1 None 1
98 1GbF bionty.DevelopmentalStage human hsapdv False True Human Developmental Stages https://github.com/obophenotype/developmental-... https://github.com/obophenotype/developmental-... 1 None 2024-05-28 None 2025-01-20 07:34:38.670000+00:00 1 None 1
100 10va bionty.DevelopmentalStage mouse mmusdv False True Mouse Developmental Stages https://github.com/obophenotype/developmental-... https://github.com/obophenotype/developmental-... 1 None 2024-05-28 None 2025-01-20 07:34:38.670000+00:00 1 None 1
102 MJRq bionty.Ethnicity human hancestro False True Human Ancestry Ontology https://github.com/EBISPOT/hancestro/raw/3.0/h... 76dd9efda9c2abd4bc32fc57c0b755dd https://github.com/EBISPOT/hancestro 1 None 3.0 None 2025-01-20 07:34:38.670000+00:00 1 None 1
103 5JnV BioSample all ncbi False True NCBI BioSample attributes s3://bionty-assets/df_all__ncbi__2023-09__BioS... 918db9bd1734b97c596c67d9654a4126 https://www.ncbi.nlm.nih.gov/biosample/docs/at... 1 None 2023-09 None 2025-01-20 07:34:38.670000+00:00 1 None 1